Shapefile Field Name Truncation¶
A Shapefile dataset with attribute field names that exceed the DBF 10-character limit. When written to Shapefile format, field names are truncated or renamed, potentially causing data loss or field name collisions. This case exposes workflows that don't handle Shapefile field name limitations.
| Property | Value |
|---|---|
| Case ID | shapefile_field_truncation |
| Category | vector |
| Format | Shapefile |
| Geometry type | Point |
| CRS | EPSG:4326 |
| Location | Central Europe (synthetic) — 10.00°E, 50.00°N → 12.00°E, 52.00°N |
| Test tier | unit |
| Size class | tiny |
| Storage class | bundled |
| Redistributable | yes |
| Loader | geopandas |
| Status | validated |
Use this case¶
import pytest
@pytest.mark.geocase_case("shapefile_field_truncation")
def test_shapefile_field_truncation(geocase_case) -> None:
data = geocase_case.load()
assert data is not None
Use GeoCase in your tests¶
Install the complete set of vector, raster, and NetCDF dependencies:
What this case checks¶
Expose workflows that assume arbitrary field name lengths or fail to detect/handle Shapefile's 10-character field name limit. Detect cases where truncation causes field name collisions.
Risk types covered¶
attribute/field_name_truncationattribute/lossattribute/schema_mismatchformat/limitation
Expected behavior¶
| Assertion | Expected |
|---|---|
expect_loadable |
yes |
expect_valid_geometry |
yes |
expect_crs |
yes |
expected_epsg |
4326 |
expected_geometry_types |
Point |
Notes¶
Purpose¶
This case tests Shapefile's 10-character field name limitation. The DBF format used by Shapefiles restricts field names to 10 characters, which causes truncation when longer names are used.
Problem Demonstrated¶
When writing data with long field names to Shapefile format:
- Truncation: Field names longer than 10 characters are silently truncated
- Collision: Truncation can cause multiple fields to have the same name
- Data Loss: Some drivers may drop or rename colliding fields
Original vs Truncated Field Names¶
| Original Name | Truncated | Notes |
|---|---|---|
temperature_celsius |
temperatur |
Truncated at 10 chars |
temperature_fahrenheit |
temper_1 |
Renamed to avoid collision |
precipitation_mm |
precipitat |
Truncated at 10 chars |
wind_speed_knots |
wind_speed |
Exactly 10 chars, no truncation |
Expected Behavior¶
- Loaders should successfully read the file
- Attribute inspection should reveal truncated/renamed field names
- Roundtrip tests should detect field name changes
- Schema validation should flag the truncation if comparing against original schema
Test Scenarios¶
- Load and inspect: Verify file loads and geometry is valid
- Field name check: Compare loaded field names against expected truncated names
- Roundtrip warning: If re-exporting to Shapefile, verify no additional truncation occurs
- Cross-format comparison: Compare against GeoJSON version with original field names
Format-Specific Behavior¶
This edge case is unique to Shapefile format due to the DBF specification. Other formats (GeoJSON, GPKG, Parquet) do not have this limitation.
Required capabilities¶
loadattribute-inspectionschema-validation
Files¶
- Primary:
truncated_fields.shp - Sidecar:
truncated_fields.dbf - Sidecar:
truncated_fields.shx - Sidecar:
truncated_fields.prj - Notes:
notes.md
Source and license¶
- Source: geocase-curated
- License: MIT
Tags¶
attributes encoding field_truncation format_specific schema shapefile vector
Related cases¶
- Format-Limited KML Case --
format_limited_kml_case - Parquet Mixed Schema Attributes --
parquet_mixed_schema_attributes - Shapefile Legacy DBF Encoding --
shapefile_encoding_legacy - Mixed Encoding Attributes --
mixed_encoding_attributes - Shapefile Ring Orientation Reversal --
shapefile_ring_orientation