Skip to content

Shapefile Legacy DBF Encoding

vectorShapefilePointEPSG:4326tinybundled

A Shapefile dataset with attribute text encoded using a legacy code page (Windows-1252/CP1252) instead of UTF-8. Many older Shapefiles and GIS tools use code page-based encoding, which can cause mojibake when read with UTF-8 assumptions. This case exposes encoding detection failures and character corruption in attribute parsing.

Point geometry of shapefile_encoding_legacy, rendered from the case's data
Point geometry, rendered from the case's actual geometry. Scale is normalized to the viewport and is not comparable between cases.
Property Value
Case ID shapefile_encoding_legacy
Category vector
Format Shapefile
Geometry type Point
CRS EPSG:4326
Location West Africa, off Senegal — 46.63°W, 23.55°S → 13.00°E, 55.60°N
Test tier unit
Size class tiny
Storage class bundled
Redistributable yes
Loader geopandas
Status validated

Use this case

import pytest


@pytest.mark.geocase_case("shapefile_encoding_legacy")
def test_shapefile_encoding_legacy(geocase_case) -> None:
    data = geocase_case.load()
    assert data is not None

Use GeoCase in your tests

Install the complete set of vector, raster, and NetCDF dependencies:

pip install "geocase[all]"

View GeoCase on PyPI.

What this case checks

Expose loaders that assume UTF-8 encoding for Shapefile attributes. Detect mojibake or character corruption when reading legacy code page encoded text. Test encoding detection and .cpg file handling.

Risk types covered

Expected behavior

Assertion Expected
expect_loadable yes
expect_valid_geometry yes
expect_crs yes
expected_epsg 4326
expected_geometry_types Point

Notes

Purpose

This case tests handling of legacy code page encoding in Shapefile DBF files. Many older Shapefiles use Windows-1252 (CP1252) or other code pages instead of UTF-8 for text attributes.

Problem Demonstrated

Shapefile encoding handling is complex:

  1. No standard encoding: DBF files don't have a built-in encoding declaration
  2. Code page files: The optional .cpg sidecar file specifies encoding
  3. Legacy defaults: Many tools assume a system default (often Windows-1252)
  4. Mojibake risk: UTF-8 assumption on CP1252 data corrupts characters

Test Data

City names with special characters encoded in Windows-1252:

City Characters Encoding Challenge
Zürich ü (0xFC) German umlaut
Köln ö (0xF6) German umlaut
Malmö ö (0xF6) Swedish character
São Paulo ã (0xE3) Portuguese tilde

Expected Behavior

With .cpg file present (this case)

  • Loaders should detect Windows-1252 encoding from .cpg file
  • Characters should render correctly without mojibake
  • Attribute values should match expected strings

Without .cpg file

  • Behavior varies by loader and platform
  • May produce mojibake (e.g., "Zürich" instead of "Zürich")
  • Some loaders attempt encoding detection heuristics

Format-Specific Behavior

This edge case is specific to Shapefile/DBF format:

  • GeoJSON mandates UTF-8
  • GeoPackage uses SQLite's UTF-8 text handling
  • Parquet/Arrow use UTF-8 by default

Shapefile's lack of encoding standardization makes it a common source of internationalization bugs in geospatial workflows.

Required capabilities

  • load
  • attribute-inspection
  • encoding-detection

Files

Browse this case on GitHub

Source and license

  • Source: geocase-curated
  • License: MIT

Tags

attributes encoding format_specific i18n legacy shapefile text vector