Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Large datasets

datui reads only the rows on screen to show a table, but a query, sort, aggregation or analysis may read the whole input.

datui --sample-rows 50000 https://d37ci6vzurychx.cloudfront.net/trip-data/yellow_tripdata_2025-01.parquet

ToDo
Page quickly through a large datasetPrefer Parquet: its footers hold the types and row counts, and it is read a row group at a time
Open local partitionsPass the directory (datui events/), so datui combines the footers and counts the rows
Analyze many rowsAnalyses read a sample, 100,000 rows by default; --sample-rows N changes it, 0 reads every row
Work on part of a large tableS draws a sample into memory: the query, Analysis, charts and export run on it, and export saves it
Chart many rowsCharts read [analysis] chart_rows rows, 10,000 by default, spread across the table; aggregate a long time series first to chart every step
Pivot a large tableFilter first: a pivot reads every row it covers to find its columns
Open compressed CSV, TSV or PSVPut --temp-dir on a disk with room for the uncompressed file
See what a format readsThe formats table: lazy scan, decompressed copy, converted to Arrow, or in memory. Past [read] memory_warning (1 GiB by default), datui asks before reading a file into memory
See what was readi → Resources for the buffer and the loading measurements; Notes for row groups and small files

[performance] streaming (on by default) runs what Polars can in batches; it is not a memory limit on every query. Performance has measured times to first rows and memory.

How large datasets open

A directory of more than 64 Parquet files, local or in the cloud, opens on the first and last files by name and reads the other footers in the background; the footer counts them on a line of its own. Until they are in:

  • The total row count is estimated from a random sample of 2,000 footers, read first: ~4.12B (est.) in the footer and on the Info panel.
  • Every empty cell shows as ∅.
  • New columns join the end of the table as they are found; Notes says how much was read.
  • A query, pivot or drill-down leaves the new columns out until you return to the data as opened.

Above 20,000 files the background pass reads a sample of the footers, and the row count reads the rest: the footer shows files 18,402 / 842,225 and Esc stops it, leaving the estimate. Counts read 256 footers at once.

FilesRow count
Up to 64Exact at once, from every footer
Up to 20,000Estimated, then exact when the background pass is done
Up to read.exact_count_files (50,000)Estimated, then counted in the background
MoreEstimated; c on the Info panel counts exactly, and so does End

A count keeps each file’s footer in the cache by its path, size, time and etag. Counting the dataset again, after a stop, or after files were added, reads only the footers it does not have.

Value counts of a dataset of files say in the footer how many of the files the read has reached.

While a directory or prefix is listed, the loading screen counts the files: Listing files: 412,000; Ctrl+O stops it. A large S3 or Google Cloud prefix is listed in parallel key ranges. Partition columns come from the listing: the newest file’s path names them, and the first and newest files’ values set their types.

The Notes tab flags layouts that make an open slow:

FindingWhy it matters
Median row group above 64 MiBA page may need a large row group read
More than 10,000 files, median below 1 MiBMany footer reads before the first rows
Different partition keys, such as date and dtThe partition columns differ across the dataset; key order alone is fine

-c read.parquet_schema=first skips datui’s union of the footers and lets Polars take one file’s schema; the partition-key check is skipped too.

Opening it again

The schema of a remote dataset, or of a local directory of more than 64 files, is cached by its URL or path. Opening it again lists the files; when no name, size, time or etag changed, no footer is read. The home screen uses the same cache to show a local directory’s full row count without reading a footer. The cache keeps 128 MiB of schemas and drops the dataset opened longest ago first. datui cache clear clears it with the rest of the cache.