First NULL After 10,000 Non-NULL Values (GeoPackage)¶
A GeoPackage of 10,001 points whose integer measure column is non-NULL for the first 10,000 rows and NULL for the last. A reader that infers a column type from a head sample types it as an integer; a reader that scans the whole column has to widen it to accommodate the missing value. pyogrio returns int64 from a partial read and float64 from the full read of this same column -- a dtype that depends on how much of the file was read.
| Property | Value |
|---|---|
| Case ID | null_after_batch_boundary_gpkg |
| Category | vector |
| Format | GPKG |
| Geometry type | Point |
| CRS | EPSG:4326 |
| Location | Central Europe (synthetic lattice) — 10.00°E, 50.00°N → 10.49°E, 50.50°N |
| Test tier | unit |
| Size class | small |
| Storage class | bundled |
| Redistributable | yes |
| Loader | geopandas |
| Status | validated |
Use this case¶
import pytest
@pytest.mark.geocase_case("null_after_batch_boundary_gpkg")
def test_null_after_batch_boundary_gpkg(geocase_case) -> None:
data = geocase_case.load()
assert data is not None
Use GeoCase in your tests¶
Install the complete set of vector, raster, and NetCDF dependencies:
What this case checks¶
Expose consumers whose column dtype depends on how many rows they read. A schema inferred from the first batch says integer; the full column is nullable, so the value read back at row 10,000 is either a float, a sentinel, or a silent zero depending on the reader.
Risk types covered¶
Expected behavior¶
| Assertion | Expected |
|---|---|
expect_loadable |
yes |
expect_valid_geometry |
yes |
expect_crs |
yes |
expected_epsg |
4326 |
expected_geometry_types |
Point |
Notes¶
10,001 points. The measure column holds an integer for the first 10,000 rows
and NULL for the last.
What it discriminates¶
A reader that types a column from a head sample sees 10,000 integers and says integer. A reader that scans the column has to accommodate one missing value, which in pandas means widening to float — so the same column of the same file comes back with a different dtype depending on how much was read.
Observed with pyogrio 0.12.1 / GDAL 3.12.2 on both code paths:
| read | dtype |
|---|---|
read_dataframe(path, max_features=100) |
int64 |
read_dataframe(path) |
float64 |
read_dataframe(path, max_features=100, use_arrow=True) |
int64 |
read_dataframe(path, use_arrow=True) |
float64 |
This is documented pyogrio behaviour rather than a bug — but a consumer that builds a schema from a sample and then reads the rest against it gets a type mismatch it did not cause and cannot see in a small fixture. That is the failure mode the case exists to make reachable.
Why Int64 on the write side¶
The column is built as pandas' nullable Int64, not int64 or float64.
A numpy int column cannot hold the NULL at all, and a float column would put
the widening in the fixture — where it is the thing under test, not a
property of the data. Int64 writes a genuine SQL NULL into an INTEGER column,
which is what a real dataset looks like.
Generated, not committed¶
Built by scripts/generate_vector_fixtures.py (_large_specs), under the
--check regeneration gate. The feature count comes from
params.expected_feature_count in case.yaml, so the generator and the
content gate cannot disagree about the size.
Written with SPATIAL_INDEX=NO; see the sibling case's notes for why.
Required capabilities¶
loadnull-handlingattribute-check
Files¶
- Primary:
null_after_boundary.gpkg - Notes:
notes.md
Source and license¶
- Source: geocase-curated
- License: MIT
Tags¶
batch_boundary large null_handling point procedural vector
Related cases¶
- UTC Offset Changes at Row 10,000 (GeoPackage) --
mixed_timezone_after_batch_gpkg - Invalid Geometry at Feature 9,999 (GeoPackage) --
invalid_geometry_at_scale_gpkg - Empty Geometry in GeoPackage --
empty_geometry_gpkg - Dateline Chain Cluster --
dateline_chain_cluster - Dateline Points Pair --
dateline_points_pair