Median latency 30.0ms over 15 timed runs
Fastest run 8.0ms best recorded
Slowest run 102.0ms worst recorded
Bytes avoided 66.7% 59.3 MB never read

Benchmark suite

benchmark data, not live telemetry

A fixed set of query shapes run against the sample dataset, chosen to exercise each stage of the scan path: a full scan with nothing to prune, partition elimination, row-group skipping on min/max statistics, and column projection. Each case records wall-clock time alongside the compressed bytes the query actually paid for.

The byte columns are the point. For every query the engine reports which files it never opened, which row-groups it skipped and why, and which columns it left on disk — measured in compressed bytes, with a regression test asserting that a full scan reports exactly 100%. Most query engines will not tell you that, and neither will Athena.

Query Shape Wall clock Bytes read Of dataset Columns
Group by partition column GROUP BY region 68.5 ms 3.1 MB 51.9% 2 of 7
Filtered aggregate WHERE + GROUP BY 127.3 ms 2.7 MB 45.1% 3 of 7
Selective row-group scan narrow WHERE on a sorted column 30.6 ms 271.0 KB 4.5% 1 of 7
Partition-pruned count WHERE on a partition column 26.7 ms 577.4 KB 9.4% 1 of 7
High-cardinality group GROUP BY pickup_hour 109.6 ms 1.8 MB 30.4% 2 of 7

Loaded from benchmark_results.json in the repository, regenerated against the sample dataset on a single machine. Timings are reported as measured.

Latency over time

last 200 successful runs

Wall-clock per run, planning included. The spikes are cold runs: row-group metadata is cached after first read, so a repeated query against the same dataset skips re-reading every footer.

Scan efficiency

read vs avoided

Column height is the compressed size of the whole dataset that run targeted. Both sides of the ratio are compressed on-disk bytes — comparing compressed reads against an uncompressed total would flatter every number here by about 1.5x.

Slowest runs

In history
Run SQL Duration Rows scanned
#1 SELECT * FROM trips LIMIT 20 102.01 ms 300000
#6 SELECT * FROM trips LIMIT 20 101.34 ms 300000
#11 SELECT * FROM trips LIMIT 20 89.94 ms 300000
#13 SELECT COUNT(*) AS total FROM trips WHERE t… 34.74 ms 300000
#3 SELECT COUNT(*) AS total FROM trips WHERE t… 31.97 ms 300000
#8 SELECT COUNT(*) AS total FROM trips WHERE t… 31.37 ms 300000
#4 SELECT region, COUNT(*) AS n, AVG(distance)… 30.89 ms 300000
#14 SELECT region, COUNT(*) AS n, AVG(distance)… 30.01 ms 300000
Latency tracks rows decoded, not bytes pruned — a query can avoid most of the dataset and still be slow if what remains is wide.

Least efficient scans

In history
Run SQL Read Of dataset
#1 SELECT * FROM trips LIMIT 20 5.9 MB 100.0%
#6 SELECT * FROM trips LIMIT 20 5.9 MB 100.0%
#11 SELECT * FROM trips LIMIT 20 5.9 MB 100.0%
#3 SELECT COUNT(*) AS total FROM trips WHERE t… 1.6 MB 26.9%
#8 SELECT COUNT(*) AS total FROM trips WHERE t… 1.6 MB 26.9%
#13 SELECT COUNT(*) AS total FROM trips WHERE t… 1.6 MB 26.9%
#4 SELECT region, COUNT(*) AS n, AVG(distance)… 1.4 MB 23.2%
#9 SELECT region, COUNT(*) AS n, AVG(distance)… 1.4 MB 23.2%
100% is not a bug. A query with no pushdownable filter must read everything, and that case is pinned by a regression test precisely so the number stays honest.

Most efficient scans

In history
Run Dataset SQL Duration Read Of dataset Pruned
#2 trips SELECT COUNT(*) AS total FROM trips WHERE region = 'EU' 9.06 ms 28.4 KB 0.5% 10f / 0rg
#7 trips SELECT COUNT(*) AS total FROM trips WHERE region = 'EU' 8.12 ms 28.4 KB 0.5% 10f / 0rg
#12 trips SELECT COUNT(*) AS total FROM trips WHERE region = 'EU' 7.99 ms 28.4 KB 0.5% 10f / 0rg
#5 trips SELECT region, COUNT(*) AS n, SUM(distance) AS total_di… 17.57 ms 1010.1 KB 16.5% 10f / 0rg
#10 trips SELECT region, COUNT(*) AS n, SUM(distance) AS total_di… 13.7 ms 1010.1 KB 16.5% 10f / 0rg
#15 trips SELECT region, COUNT(*) AS n, SUM(distance) AS total_di… 17.76 ms 1010.1 KB 16.5% 10f / 0rg
#4 trips SELECT region, COUNT(*) AS n, AVG(distance) AS avg_dist… 30.89 ms 1.4 MB 23.2% 0f / 0rg
#9 trips SELECT region, COUNT(*) AS n, AVG(distance) AS avg_dist… 28.05 ms 1.4 MB 23.2% 0f / 0rg
#14 trips SELECT region, COUNT(*) AS n, AVG(distance) AS avg_dist… 30.01 ms 1.4 MB 23.2% 0f / 0rg
#3 trips SELECT COUNT(*) AS total FROM trips WHERE trip_id > 2374 31.97 ms 1.6 MB 26.9% 0f / 0rg
These are the runs where pushdown did the most work. Pruned is files never opened and row-groups skipped on footer statistics.