Datasets
A dataset is a registered prefix of Parquet files, local or in object storage. Registering one reads footers only — there is no load step and no copy of the data.
| Name | Location | Files | Compressed | Partitioned by | Columns | Runs | Median | Read | |
|---|---|---|---|---|---|---|---|---|---|
| trips | ./data/trips | 15 | 6.0 MB | region date | 9 | 15 | 30.0 ms | 33.3% | Query |
Median latency and read share are over all successful runs against each dataset.
Read is compressed bytes scanned as a share of the compressed bytes
those queries could have scanned.
Registering another prefix
The engine resolves any URI its filesystem layer understands — a local path or
an S3-compatible bucket. It opens one Parquet footer to infer the schema, parses
Hive-style key=value directory names into partition columns, and records
the file inventory. The data itself is never read or moved.
python manage.py register_dataset trips ./data/trips python manage.py register_dataset events s3://my-bucket/events