Skip to content

Checking Best Practices

The check commands validate GeoParquet files against best practices.

Run All Checks

gpio check all myfile.parquet
import geoparquet_io as gpio

table = gpio.read('myfile.parquet')
result = table.check()

if result.passed():
    print("All checks passed!")
else:
    for failure in result.failures():
        print(f"Failed: {failure}")

# Get full results as dictionary
details = result.to_dict()

Runs all validation checks:

  • Spatial ordering
  • Compression settings
  • Bbox structure and metadata
  • Row group optimization
  • Bloom filter detection
  • GeoParquet v2.0 upgrade recommendation (for v1.1 files)
  • Spec validation

Individual Checks

Spatial Ordering

gpio check spatial myfile.parquet

result = table.check_spatial()
print(f"Spatially ordered: {result.passed()}")

Checks if data is spatially ordered. Spatially ordered data improves:

  • Query performance (10-100x faster for spatial queries)
  • Compression ratios
  • Cloud access patterns

Method Selection:

  • GeoParquet 2.0+ files (with bbox column): Uses fast bbox-stats method by analyzing row group metadata (~10-100x faster)
  • GeoParquet 1.x files (no bbox column): Falls back to sampling method which analyzes actual geometry data

For faster spatial order checks

Add a bbox column to your file with gpio add bbox to enable the fast bbox-stats method.

How it works:

  • Bbox-stats method: Checks if consecutive row groups have overlapping bounding boxes. Non-overlapping row groups indicate good spatial ordering. Passes if < 30% of row group pairs overlap.
  • Sampling method: Compares average distance between consecutive features vs random feature pairs. Lower ratio indicates better spatial clustering. Passes if ratio < 0.5.

Compression

gpio check compression myfile.parquet

result = table.check_compression()
print(f"Compression optimal: {result.passed()}")

Validates geometry column compression settings.

Bbox Structure

gpio check bbox myfile.parquet

result = table.check_bbox()
if not result.passed():
    # Add bbox if missing
    table = table.add_bbox().add_bbox_metadata()

Verifies:

  • Bbox column structure
  • GeoParquet metadata version
  • Bbox covering metadata

A bbox column is optional in GeoParquet 2.0, but it must be declared, and the declaration must be true:

  • A bbox column that no covering entry points at costs file size and cannot be used by any reader, so check bbox fails the file and suggests gpio add bbox-metadata (or --fix to drop the column). --fix only ever removes an undeclared column — a covering the file legitimately declares is never deleted.
  • A covering that names a column the file does not contain is treated as undeclared rather than accepted. Such a covering makes readers prune away rows that genuinely match, so it is worse than none at all.

check spec runs its four covering checks — paths well-formed, column present, struct shape, field types — at GeoParquet 1.1 and 2.0. They were previously gated to 1.1 only, so an identical broken covering failed at 1.1 and passed at 2.0.

Row Groups

gpio check row-group myfile.parquet

result = table.check_row_groups()
for rec in result.recommendations():
    print(rec)

Checks row group size optimization for cloud-native access.

Spatial filter pushdown and row group sizing

For GeoParquet 2.0 or parquet-geo-only files with Hilbert sorting, row groups of 10,000-50,000 rows create tighter bounding boxes that enable more row group skipping during spatial queries.

Optimization Check

gpio check optimization myfile.parquet

result = table.check_optimization()
print(f"Score: {result.to_dict()['score']}/5")

Evaluates five factors affecting spatial query performance and returns a score from 0 to 5:

  1. Native Geo Types - Uses native Parquet geo types (GeoParquet 2.0 or parquet-geo-only)
  2. Geo Bbox Stats - Per-row-group geo bbox statistics present
  3. Spatial Sorting - Data is spatially sorted (Hilbert or similar)
  4. Row Group Size - Appropriate for file size (10k-50k rows for spatial pushdown)
  5. Compression - ZSTD compression on geometry column

Scoring levels:

  • fully_optimized (5/5) - All checks pass
  • partially_optimized (3-4/5) - Some improvements possible
  • not_optimized (0-2/5) - Significant improvements needed

Spatial Filter Pushdown Readiness

The gpio check spatial command also reports spatial filter pushdown readiness when bbox data is available:

gpio check spatial myfile.parquet

result = table.check_spatial_pushdown()
details = result.to_dict()
print(f"Skip rate: {details['estimated_skip_rate']}")

Shows:

  • Row group count and bbox coverage
  • Estimated skip rate - percentage of row groups that can be skipped for representative spatial queries
  • Avg bbox area ratio - how tight the row group bounding boxes are

Requires bbox data

Pushdown readiness requires GeoParquet 2.0 native geo stats or a bbox column. For v1.1 files, add a bbox column with gpio add bbox.

Bloom Filters

Bloom filter detection is included in gpio check all and gpio inspect meta:

# Included automatically in check all
gpio check all myfile.parquet

result = table.check_bloom_filters()
details = result.to_dict()

Reports which columns have bloom filters, coverage percentages, and total bloom filter bytes. DuckDB 1.5+ automatically writes bloom filters for low-cardinality columns.

Spec Validation

# Auto-detect version
gpio check spec data.parquet

# Validate against specific version
gpio check spec data.parquet --geoparquet-version 1.1

# JSON output for CI/CD
gpio check spec data.parquet --json

result = table.validate()
if result.passed():
    print("Valid GeoParquet!")

Validates file structure and metadata against the GeoParquet specification:

  • Supports GeoParquet 1.0, 1.1, 2.0, and Parquet native geo types
  • Auto-detects version unless --geoparquet-version is specified
  • Optional data validation against metadata claims

Metadata checks include:

  • Version sanity — an unknown geo metadata version (e.g. 99.0.0) fails validation, even in auto mode. Version/feature mismatches also fail: a file declaring version 1.0 while using GeoParquet 2.0 native Parquet geo types, or the covering key that 1.1 introduced, is flagged as inconsistent.
  • CRS structure — PROJJSON crs objects must carry the required type member, and it must be a known PROJJSON CRS type.
  • Datum-aware epoch validation — a coordinate epoch on a datum ensemble (e.g. EPSG:4326, or the OGC:CRS84 default when crs is omitted) fails; on a specific static frame (e.g. GDA2020) it warns; on a dynamic frame (e.g. ITRF) it passes. With an explicit "crs": null the datum cannot be verified, so an epoch produces a warning.
  • Malformed metadata — a geo key containing invalid JSON or a non-object value fails cleanly in auto mode instead of crashing.
  • Dimension-aware geometry typesgeometry_types entries carry Z/M suffixes ("Point Z", "LineString ZM"), and validation matches declared suffixes against the actual coordinate dimensions in both directions (declared-but-absent and present-but-undeclared).
  • GeoArrow encodings — the single-geometry type encodings (point, linestring, polygon, multipoint, multilinestring, multipolygon) are valid for GeoParquet 1.1, so files written with --geoparquet-version 1.1-geoarrow validate. The checks are layout-aware: the BYTE_ARRAY requirement applies only to WKB columns, native columns are checked for the DOUBLE coordinate group the spec requires, and the data scans read GeoArrow coordinates directly. GeoParquet 1.0 and 2.0 are WKB-only per their spec text, so a file claiming a GeoArrow encoding under either version fails.
  • Native geospatial statistics — for files using the Parquet GEOMETRY or GEOGRAPHY logical types, native_geo_stats_* reports the bounds a file declares and native_geo_stats_contains_data_* checks its geometries against them. Both use the whole file's statistics — the union over every row group, not the first one's — so a correctly written multi-row-group file passes (#721).

Exit codes:

  • 0 - All checks passed
  • 1 - One or more checks failed
  • 2 - Warnings only (all required checks passed)

Coordinate/CRS mismatch heuristic is now a warning

The heuristic that flags geographic-looking coordinates (values within ±180/±90) under a projected CRS reports a WARNING instead of a failure, so gpio check spec exits with code 2 instead of 1 for affected files. Update CI pipelines that gate only on exit code 1. The deterministic CRS-consistency check for GeoParquet 2.0 native geo statistics still fails on a real mismatch.

STAC Validation

gpio check stac output.json

from geoparquet_io import validate_stac

result = validate_stac('output.json')
if result.passed():
    print("Valid STAC!")

Validates STAC Item or Collection JSON:

  • STAC spec compliance
  • Required fields
  • Asset href resolution (local files)
  • Best practices

Options

# Verbose output with details
gpio check all myfile.parquet --verbose

# Custom sampling for spatial check
gpio check spatial myfile.parquet --random-sample-size 200 --limit-rows 1000000

# Custom sampling for spatial check
result = table.check_spatial(sample_size=200, limit_rows=1000000)

Checking Partitioned Data

When checking a directory containing partitioned data, you can control how many files are checked:

# By default, checks only the first file
gpio check all partitions/
# Output: Checking first file (of 4 total). Use --all-files or --sample-files N for more.

# Check all files in the partition
gpio check all partitions/ --all-files

# Check a sample of files (first N files)
gpio check all partitions/ --sample-files 3

--fix not available for partitions

The --fix option only works with single files. To fix issues in partitioned data, first consolidate with gpio extract, apply fixes, then re-partition if needed.

See Also