Shapefile Legacy DBF Encoding¶
A Shapefile dataset with attribute text encoded using a legacy code page (Windows-1252/CP1252) instead of UTF-8. Many older Shapefiles and GIS tools use code page-based encoding, which can cause mojibake when read with UTF-8 assumptions. This case exposes encoding detection failures and character corruption in attribute parsing.
| Property | Value |
|---|---|
| Case ID | shapefile_encoding_legacy |
| Category | vector |
| Format | Shapefile |
| Geometry type | Point |
| CRS | EPSG:4326 |
| Location | West Africa, off Senegal — 46.63°W, 23.55°S → 13.00°E, 55.60°N |
| Test tier | unit |
| Size class | tiny |
| Storage class | bundled |
| Redistributable | yes |
| Loader | geopandas |
| Status | validated |
Use this case¶
import pytest
@pytest.mark.geocase_case("shapefile_encoding_legacy")
def test_shapefile_encoding_legacy(geocase_case) -> None:
data = geocase_case.load()
assert data is not None
Use GeoCase in your tests¶
Install the complete set of vector, raster, and NetCDF dependencies:
What this case checks¶
Expose loaders that assume UTF-8 encoding for Shapefile attributes. Detect mojibake or character corruption when reading legacy code page encoded text. Test encoding detection and .cpg file handling.
Risk types covered¶
attribute/corruptionattribute/encoding_errorattribute/mojibakeformat/limitation
Expected behavior¶
| Assertion | Expected |
|---|---|
expect_loadable |
yes |
expect_valid_geometry |
yes |
expect_crs |
yes |
expected_epsg |
4326 |
expected_geometry_types |
Point |
Notes¶
Purpose¶
This case tests handling of legacy code page encoding in Shapefile DBF files. Many older Shapefiles use Windows-1252 (CP1252) or other code pages instead of UTF-8 for text attributes.
Problem Demonstrated¶
Shapefile encoding handling is complex:
- No standard encoding: DBF files don't have a built-in encoding declaration
- Code page files: The optional
.cpgsidecar file specifies encoding - Legacy defaults: Many tools assume a system default (often Windows-1252)
- Mojibake risk: UTF-8 assumption on CP1252 data corrupts characters
Test Data¶
City names with special characters encoded in Windows-1252:
| City | Characters | Encoding Challenge |
|---|---|---|
| Zürich | ü (0xFC) | German umlaut |
| Köln | ö (0xF6) | German umlaut |
| Malmö | ö (0xF6) | Swedish character |
| São Paulo | ã (0xE3) | Portuguese tilde |
Expected Behavior¶
With .cpg file present (this case)¶
- Loaders should detect Windows-1252 encoding from
.cpgfile - Characters should render correctly without mojibake
- Attribute values should match expected strings
Without .cpg file¶
- Behavior varies by loader and platform
- May produce mojibake (e.g., "Zürich" instead of "Zürich")
- Some loaders attempt encoding detection heuristics
Format-Specific Behavior¶
This edge case is specific to Shapefile/DBF format:
- GeoJSON mandates UTF-8
- GeoPackage uses SQLite's UTF-8 text handling
- Parquet/Arrow use UTF-8 by default
Shapefile's lack of encoding standardization makes it a common source of internationalization bugs in geospatial workflows.
Required capabilities¶
loadattribute-inspectionencoding-detection
Files¶
- Primary:
legacy_encoding.shp - Sidecar:
legacy_encoding.dbf - Sidecar:
legacy_encoding.shx - Sidecar:
legacy_encoding.prj - Sidecar:
legacy_encoding.cpg - Notes:
notes.md
Source and license¶
- Source: geocase-curated
- License: MIT
Tags¶
attributes encoding format_specific i18n legacy shapefile text vector
Related cases¶
- Mixed Encoding Attributes --
mixed_encoding_attributes - Shapefile Field Name Truncation --
shapefile_field_truncation - Format-Limited KML Case --
format_limited_kml_case - Shapefile Ring Orientation Reversal --
shapefile_ring_orientation - Empty Geometry in GeoPackage --
empty_geometry_gpkg