SQL

success
SQL
SELECT region, COUNT(*) AS n, AVG(distance) AS avg_distance
FROM trips
GROUP BY region
ORDER BY n DESC

Dataset trips · ./data/trips

success 30.89 ms 3 rows 23.2% of dataset bytes read Run #4 details
3 rows 3 columns
Row regionnavg_distance
1 EU100,00020.0795
2 APAC100,00020.1004
3 US100,00020.1498
Bytes scanned / dataset bytes
23.2%
1.4 MB came off disk, out of 5.9 MB of compressed dataset bytes.
Execution
Duration 30.89 ms
Rows returned 3
Rows scanned 300000
Files pruned 0 of 15 files
Row-groups pruned 0 of 120 row-groups
Columns read 1 of 7 columns
Why it read that much
  • Only 1 of 7 columns came off disk — the other 6 were never referenced, and Parquet is columnar.
Where the bytes went
A naive full scan
5.9 MB
every byte of every column in every file
Columns not referenced
4.5 MB
Parquet is columnar, so these were never touched
Actually read
1.4 MB
what came off disk to answer the query

The middle bars are bytes a scan-everything engine would have read and this one did not. They add up to the full scan exactly — it is one number decomposed, not four separate measurements.

Row-group map (120 of 120 read)
region=EU/date=2024-01-01
region=EU/date=2024-01-02
region=EU/date=2024-01-03
region=EU/date=2024-01-04
region=EU/date=2024-01-05
region=US/date=2024-01-01
region=US/date=2024-01-02
region=US/date=2024-01-03
region=US/date=2024-01-04
region=US/date=2024-01-05
region=APAC/date=2024-01-01
region=APAC/date=2024-01-02
region=APAC/date=2024-01-03
region=APAC/date=2024-01-04
region=APAC/date=2024-01-05
read skipped — footer stats proved no match file never opened, so its row-group count is unknown by design