Most datasets are hive-partitioned by h0. When both sides of a join have h0, always include it in the join condition:
JOIN table2 ON table1.hX = table2.hX AND table1.h0 = table2.h0where hX is the finest resolution shared by both datasets (h8, h9, etc. — check the
schema). Omitting AND t1.h0 = t2.h0 causes DuckDB to open every partition file on S3
instead of only the matching ones (10-100x slower).
Use regions/hex/** or countries.parquet as the first CTE to establish geographic
scope before joining large thematic datasets (PADUS, carbon, wetlands, species).
WITH scope AS (
SELECT DISTINCT h8, h0
FROM read_parquet('<STAC_REGIONS_HEX_PATH>')
WHERE region = 'US-CA'
),
parks AS (
SELECT DISTINCT p.h8, p.h0
FROM scope s
JOIN read_parquet('<STAC_PADUS_HEX_PATH>') p
ON s.h8 = p.h8 AND s.h0 = p.h0
WHERE p.Des_Tp = 'NP'
)
SELECT SUM(c.carbon)/1e6
FROM parks p
JOIN read_parquet('<STAC_CARBON_HEX_PATH>') c
ON p.h8 = c.h8 AND p.h0 = c.h0Note: rook-ceph-rgw-nautiluss3.rook is an internal endpoint only accessible from k8s. Always use it — not the public endpoint — to run queries.
You must read parquet datasets from S3 using read_parquet(). There are no local tables.
Aggregate after the join, never before. When joining a small scope to a large hex dataset, do the join against the raw read_parquet(...) of the large side. Wrapping the large side in a pre-aggregation CTE (GROUP BY, MODE, SUM) before the join forces DuckDB to scan every h0 partition of it — the small scope's h0 values cannot prune through an aggregation operator.
-- Join first, aggregate after → only scope's h0 partitions opened
WITH lc_on_scope AS (
SELECT s.h8, s.h0, l.lc_class
FROM scope s
JOIN read_parquet('<large_hex>') l USING (h8, h0)
WHERE l.lc_class IS NOT NULL
)
SELECT h8, h0, MODE(lc_class) AS dominant
FROM lc_on_scope GROUP BY h8, h0;The WHERE l.lc_class IS NOT NULL here keeps a no-data cell from poisoning the aggregate — but it also changes which cells the answer describes. When that column is only partly populated, see §9 before reporting the result as a share of the whole.
To filter a global …/hex/h0=*/… dataset by a name or _cng_fid and there is no region-mask hex to join, first restrict h0 to the region, then apply the attribute filter:
- Known coordinates (island centroids, a city, bbox corners): build the
h0set withh3_latlng_to_cell(lat, lng, 0)per point andUNNEST(h3_grid_disk(h0, 1))to include neighbouring res-0 cells. For large or multi-part features, supply every known point or the bbox corners — one centroid can miss partitions the feature reaches. - Known name only: look the feature up in the dataset's GeoParquet asset (one row per feature) to read its centroid/bbox, convert to
h0, then query the hex.
WITH h0s AS (
SELECT DISTINCT d
FROM (VALUES (19.282,166.647),(16.728,-169.534),(6.383,-162.417),
(0.807,-176.617),(-0.374,-159.997)) AS i(lat,lng),
UNNEST(h3_grid_disk(h3_latlng_to_cell(lat,lng,0),1)) AS t(d)
)
SELECT SUM(h3_cell_area(h8, 'km^2')) AS area_km2
FROM (
SELECT DISTINCT h8, h0
FROM read_parquet('s3://<bucket>/<dataset>/hex/h0=*/data_00.parquet')
WHERE h0 IN (SELECT d FROM h0s) AND _cng_fid IN (201, 246, 243, 142, 90)
);For datasets already loaded in your app, the column list returned by get_schema
(or get_stac_details) is canonical: every column, with type and description.
Running follow-up queries like
SELECT column_name FROM (DESCRIBE SELECT * FROM read_parquet('…') LIMIT 1)
WHERE column_name LIKE '%corridor%'to grep for specific columns is wasteful — the same data is already in your
context. Search the existing schema response instead. DESCRIBE is appropriate
only for arbitrary parquet files not represented in any STAC collection.
DuckDB LIKE is case-sensitive by default. For fuzzy substring search on a
free-text label the user typed (site, owner, program names), normalize both
sides:
WHERE lower(site) LIKE '%' || lower('user input') || '%'For an exact match on a known key (sci_name, name_en, a country or
category code), match the stored case exactly instead of wrapping the column in
lower(). Where the file is sorted or clustered by that key, exact-case lets
DuckDB skip row groups by their stored min/max stats; lower() forces a full
column scan. Take the spelling from a get_stac_details example or a
SELECT DISTINCT probe:
WHERE sci_name = 'Rana draytonii'GeoParquet files contain a geometry column (usually geom) typed as GEOMETRY('OGC:CRS84').
This type cannot be displayed in tabular output — the server drops it automatically.
Avoid SELECT * on GeoParquet files; select only the columns you need. If you need
coordinates, cast explicitly: ST_AsText(geom) AS geom_wkt.
Site names and owner names can contain apostrophes (e.g. O'Brien Ranch). Double any
single quote inside a SQL string literal — do not use a backslash:
WHERE site = 'O''Brien Ranch' -- correct
WHERE site = 'O'Brien Ranch' -- parse errorA single NaN turns the entire aggregate into NaN — one no-data cell poisons a
whole SUM or AVG. No-data shows up as literal NaN, or as sentinel codes a
dataset reserves for "no data" (e.g. land-cover classes 0, 80, 200). Exclude both
before aggregating:
SELECT SUM(value) AS total
FROM read_parquet('<hex>')
WHERE value IS NOT NULL AND NOT isnan(value)
AND lc_class NOT IN (0, 80, 200); -- dataset's no-data sentinelsTake the sentinel codes from the dataset's STAC description. A NaN total or a
total far smaller than expected is the signature of this trap.
A vector feature's area or length is the value in its own column (acres,
length_km), summed over DISTINCT feature id — tiling replicates that value on
every cell the feature touches, so dedup by _cng_fid before summing.
COUNT(DISTINCT hN) * h3_cell_area(...) is a larger, different number: it counts
every partially-covered boundary cell as if the feature filled it. Use the hex
only to locate the feature (mask, join, overlay); take the magnitude from the
measured column. For the conserved or overlaid share of that total, weight by
the coverage fraction — area for polygons and rasters, length for lines — and keep
the weight inside the same SUM() as the measure, not derived in a subquery the
outer aggregate then ignores. (See the overlay rules in the H3 guide.)
WHERE col IS NOT NULL (or excluding sentinels) fixes the arithmetic but replaces
the population. Missing values are rarely missing at random: incomplete columns are
usually incomplete in patterns that track geography, feature type, or source
vintage, so the rows that survive are a biased sample, and any share computed from
them describes that sample, not the whole. Before using a partially-populated
column as a filter, a class selector, or a denominator:
- Quantify coverage — what fraction of rows, and of the relevant measure (area, length, count), survive the filter?
- Check it against at least one independent grouping — a partition key, region, category, or date. Uniform coverage is safe to filter on; coverage that is near 0% in some groups and near 100% in others means the filter is a group selector wearing an attribute's clothes.
- Confirm how absence is encoded.
IS NOT NULLmisses sentinel values (0,-9999,''); take them from the STAC description and exclude them before measuring coverage — a sentinel can fake coverage a column does not have. - Report it. If coverage is materially incomplete, state the covered fraction and what it covers alongside the number. If coverage is concentrated in part of the study area, the honest answer is that the breakdown is unavailable for the whole — not a number with a caveat.