Tag

#hive

2 posts ·all posts

The partition list lives next to the data

Cycle 2 of SQE's Hive external tables replaces the per-query prefix LIST with a partition index persisted in the sidecar manifest, maintained by MSCK REPAIR TABLE and ALTER TABLE ADD/DROP PARTITION. Athena asks Glue for the partition list over an API. SQE reads one ETag-cached JSON object sitting beside the files. On a 449-partition slice of the public Bitcoin dataset a pinned query goes from 99 ms to 1.9 ms cold, and DuckDB on the same files stays at 44 ms because it re-globs every query. Partition projection is implemented too, and measurably slower than the index on the same pin.

Read

Athena's DDL, your bucket, nobody's metastore

SQE now reads Hive-style external tables over CSV, line-delimited JSON, and partitioned Parquet on any S3-compatible store, using Athena's own CREATE EXTERNAL TABLE syntax. There is no Hive Metastore, no Glue API call, and no catalog service of any kind behind the table: the definition is a JSON manifest next to the data, shaped like a Glue TableInput. This post is what emulating Athena and Glue actually requires, which parts we copied deliberately, the four Athena behaviours we refused to copy, and the credential difference that is a real change from how SQE treats Iceberg.

Read