H3 vs. Geohash vs. S2.
Three widely-used spatial indexes, three different tradeoffs. Here's how to pick the one that fits what you're actually doing.
The one-line versions
- Geohash — a 1970s-era hierarchical latitude/longitude index that encodes a bounding box into a short base-32 string. Simple, universally supported, uneven cell shape, awkward neighbour queries.
- S2 (Google) — a hierarchical cell index based on projecting a cube onto a sphere. Produces near-square cells, a huge resolution range, and the richest geometry library of the three. Steepest learning curve.
- H3 (Uber) — a hierarchical hexagonal index. Consistent neighbour geometry, clean rollups, and the easiest to reason about for aggregation. Approximate parent-child relationships are a feature, not a bug.
Cell shape
This is the easiest axis to compare and the one with the biggest downstream consequences.
Geohash cells are rectangles in lat/lng space, which means they're not rectangles on the ground. The aspect ratio distorts with latitude — a geohash cell in Dubai looks nothing like one in Oslo. If you plan to compare cell counts across latitudes, this is a problem.
S2 cells are approximately square on the sphere. They project through a cube and warp slightly at face edges, but to a decent approximation they're uniform in area and shape across the globe. For spatial coverage questions ("what fraction of this zone is within range of a station?") S2 is often the easiest to reason about.
H3 cells are hexagons, almost everywhere, with 12 pentagons as exceptions. Hexagons have the useful property that every neighbour sits at the same distance from the centre, which eliminates the diagonal-neighbour artifact squares have. For local-neighbourhood operations — radial smoothing, k-ring expansion, nearest-cell searches — H3 is cleanest.
Hierarchy
All three are hierarchical, but in different ways.
Geohash is a prefix hierarchy: dropping characters from the right walks you up the levels. Each additional character multiplies cell count by 32 (base-32 encoding). This means successive levels have very different granularities — a 5-character geohash is ~4.9 km, a 6-character one is ~1.2 km, a 7-character one is ~153 m. There's no finer control between those steps.
S2 is a quadtree-style hierarchy: each parent has four children. Thirty resolutions in total, so you have fine-grained control over cell size. Parent-child relationships are exact (a child cell is entirely contained in its parent).
H3 uses aperture-7 subdivision: each parent has approximately seven children. Sixteen resolutions, mid-way between S2's density and Geohash's coarseness. Parent-child relationships are approximate — a hexagon can't be perfectly subdivided into smaller hexagons, so edge points may belong to a different parent. For analytics this is almost never a problem.
Neighbour queries
Three questions that look alike but behave very differently across the three systems:
- "Give me the cells adjacent to this one."
- "Give me all cells within distance d."
- "Give me all cells intersecting this polygon."
Geohash neighbour queries are famously ugly. The eight neighbours of a geohash cell may not share a prefix: a cell near a lat/lng boundary (±90°, ±180°, or any of the base-32 encoding boundaries) has neighbours that look completely unrelated by string. Every geohash library ships a neighbors() function that handles the edge cases, but it's reinvented over and over because the base encoding doesn't help.
S2 and H3 both expose first-class neighbors() and grid_disk(k) operations with sensible semantics. H3's are especially clean because all 6 hex neighbours are equidistant; S2's work fine but you have to remember whether you want edge neighbours (4) or edge+corner (8), and the corner neighbours sit √2× further away.
Ecosystem
Geohash has the widest ecosystem by far. Every geospatial database, every analytics library, every GIS tool supports it. If you need to interop with external data, geohash is the lingua franca.
S2 has strong library support in C++, Go, Java, and Python; weaker in JavaScript and SQL contexts. Postgres has a s2geography extension; Snowflake does not natively support S2 (though you can hand-roll it).
H3 has first-party libraries in C, Python, JavaScript, Java, Go, R, and more. Native support in Snowflake, BigQuery, Databricks, DuckDB (via extension), Postgres (via extension), Elasticsearch, and most notebook-centric data stacks. If you're in a modern cloud-data workflow, H3 is the one your warehouse probably already speaks.
Precision and resolution
| Index | Levels | Finest cell | Coarsest |
|---|---|---|---|
| Geohash | 12 (effectively) | ~37 mm × 18 mm | 5 000 km |
| S2 | 30 | ~0.7 cm² | 85 M km² |
| H3 | 16 | ~0.9 m² | 4.3 M km² |
S2 wins the raw-precision contest; H3 and Geohash both top out well before sub-centimetre. For almost every real analytics problem, the chosen resolution is 6–11 anyway and the difference doesn't bite.
So which should you pick?
Default to H3 if you're doing operational analytics on moving entities (rides, deliveries, scooters, drones), building heatmaps, defining zones, or working anywhere the data-stack native (Snowflake, BigQuery, Databricks) already has H3 UDFs. The hex geometry is friendlier for nearly every downstream analysis, and the ecosystem has caught up decisively in the last few years.
Pick S2 if you need very fine resolution, exact parent-child containment, or you're inside Google's stack. Also the right choice when accurate geometric operations (coverings, intersections, areas on the sphere) are first-class in your problem — S2's geometry library is excellent.
Pick Geohash when you need maximum interop and simplicity, you're working with legacy systems that already use it, or you need the prefix-matching property for cheap bounding-box filters in a key-value store. Avoid it when neighbour queries are central to your analysis.
Can you mix them?
Yes, but with care. If you need to compare an H3-indexed dataset with a Geohash-indexed one, index both to a common resolution of a third system (lat/lng rounded to a grid, say) and go from there. Attempting to translate directly between hex and square cells will produce overlap and gaps that bias any count you compute.
Within a single pipeline, pick one and stick with it. The three systems are each good at their own thing; the glue between them is where correctness tends to go to die.