Skip to content

First NULL After 10,000 Non-NULL Values (GeoPackage)

vectorGPKGPointEPSG:4326smallbundled

A GeoPackage of 10,001 points whose integer measure column is non-NULL for the first 10,000 rows and NULL for the last. A reader that infers a column type from a head sample types it as an integer; a reader that scans the whole column has to widen it to accommodate the missing value. pyogrio returns int64 from a partial read and float64 from the full read of this same column -- a dtype that depends on how much of the file was read.

Point geometry of null_after_batch_boundary_gpkg, rendered from the case's data
Point geometry, rendered from the case's actual geometry. Scale is normalized to the viewport and is not comparable between cases.
Property Value
Case ID null_after_batch_boundary_gpkg
Category vector
Format GPKG
Geometry type Point
CRS EPSG:4326
Location Central Europe (synthetic lattice) — 10.00°E, 50.00°N → 10.49°E, 50.50°N
Test tier unit
Size class small
Storage class bundled
Redistributable yes
Loader geopandas
Status validated

Use this case

import pytest


@pytest.mark.geocase_case("null_after_batch_boundary_gpkg")
def test_null_after_batch_boundary_gpkg(geocase_case) -> None:
    data = geocase_case.load()
    assert data is not None

Use GeoCase in your tests

Install the complete set of vector, raster, and NetCDF dependencies:

pip install "geocase[all]"

View GeoCase on PyPI.

What this case checks

Expose consumers whose column dtype depends on how many rows they read. A schema inferred from the first batch says integer; the full column is nullable, so the value read back at row 10,000 is either a float, a sentinel, or a silent zero depending on the reader.

Risk types covered

Expected behavior

Assertion Expected
expect_loadable yes
expect_valid_geometry yes
expect_crs yes
expected_epsg 4326
expected_geometry_types Point

Notes

10,001 points. The measure column holds an integer for the first 10,000 rows and NULL for the last.

What it discriminates

A reader that types a column from a head sample sees 10,000 integers and says integer. A reader that scans the column has to accommodate one missing value, which in pandas means widening to float — so the same column of the same file comes back with a different dtype depending on how much was read.

Observed with pyogrio 0.12.1 / GDAL 3.12.2 on both code paths:

read dtype
read_dataframe(path, max_features=100) int64
read_dataframe(path) float64
read_dataframe(path, max_features=100, use_arrow=True) int64
read_dataframe(path, use_arrow=True) float64

This is documented pyogrio behaviour rather than a bug — but a consumer that builds a schema from a sample and then reads the rest against it gets a type mismatch it did not cause and cannot see in a small fixture. That is the failure mode the case exists to make reachable.

Why Int64 on the write side

The column is built as pandas' nullable Int64, not int64 or float64. A numpy int column cannot hold the NULL at all, and a float column would put the widening in the fixture — where it is the thing under test, not a property of the data. Int64 writes a genuine SQL NULL into an INTEGER column, which is what a real dataset looks like.

Generated, not committed

Built by scripts/generate_vector_fixtures.py (_large_specs), under the --check regeneration gate. The feature count comes from params.expected_feature_count in case.yaml, so the generator and the content gate cannot disagree about the size.

Written with SPATIAL_INDEX=NO; see the sibling case's notes for why.

Required capabilities

  • load
  • null-handling
  • attribute-check

Files

Browse this case on GitHub

Source and license

  • Source: geocase-curated
  • License: MIT

Tags

batch_boundary large null_handling point procedural vector