Formats
datui reads 27 formats: Parquet, CSV, TSV, PSV, JSON, NDJSON, Arrow IPC, Avro, ORC, Excel, SafeTensors, GGUF, NMEA, GPX, WAV/AIFF audio, MIDI, SQLite, VCD, FIX, SDF, NumPy, ELF, ULog, DataFlash, candump, plain text, systemd journal, and binary formats you describe in a format spec.
The extension says the format; --format names it when the extension does
not, and text piped in is detected by content.
printf 'a,b\n1,2\n' > export.txt
datui --format csv export.txt
How each format is read
| Format | --format | Extensions | Read | Compressed | HTTP(S) | In a bucket | Bucket prefix |
|---|---|---|---|---|---|---|---|
| Parquet | parquet | .parquet | lazy scan | no | downloaded | in place | in place |
| CSV | csv | .csv | lazy scan | decompressed copy | downloaded | downloaded | in place |
| TSV | tsv | .tsv | lazy scan | decompressed copy | downloaded | downloaded | no |
| PSV | psv | .psv | lazy scan | decompressed copy | downloaded | downloaded | no |
| JSON | json | .json | in memory | no | downloaded | downloaded | no |
| NDJSON | jsonl | .jsonl, .ndjson | in memory | no | downloaded | downloaded | in place |
| Arrow IPC | arrow | .arrow, .arrows, .ipc, .feather | lazy scan | no | downloaded | in place | in place |
| Avro | avro | .avro | in memory | no | downloaded | downloaded | no |
| ORC | orc | .orc | in memory | no | downloaded | downloaded | no |
| Excel | excel | .xls, .xlsx, .xlsm, .xlsb | in memory | no | downloaded | downloaded | no |
| SafeTensors | safetensors | .safetensors, .safetensors.index.json | in memory | no | in place | in place | in place |
| GGUF | gguf | .gguf | in memory | no | in place | in place | in place |
| NMEA | nmea | .nmea | converted to Arrow | converted to Arrow | downloaded | downloaded | no |
| GPX | gpx | .gpx | converted to Arrow | converted to Arrow | downloaded | downloaded | no |
| WAV, BWF, RF64, AIFF | audio | .wav, .wave, .bwf, .rf64, .aif, .aiff, .aifc | lazy scan | no | downloaded | downloaded | no |
| MIDI | midi | .mid, .midi, .smf, .kar, .rmi | in memory | no | downloaded | downloaded | no |
| SQLite | sqlite | .db, .db3, .sqlite, .sqlite3 | lazy scan | no | downloaded | downloaded | no |
| VCD | vcd | .vcd | converted to Arrow | converted to Arrow | downloaded | downloaded | no |
| FIX | fix | none: by content | converted to Arrow | converted to Arrow | downloaded | downloaded | no |
| SDF | sdf | .sdf, .sd | converted to Arrow | converted to Arrow | downloaded | downloaded | no |
| NumPy | numpy | .npy, .npz | lazy scan | no | downloaded | downloaded | no |
| ELF | elf | .elf, .axf | in memory | no | downloaded | downloaded | no |
| ULog | ulog | .ulg | lazy scan | no | downloaded | downloaded | no |
| DataFlash | dataflash | none: by content | lazy scan | no | downloaded | downloaded | no |
| candump | candump | none: by content | lazy scan | no | downloaded | downloaded | no |
| Text | text | .log, .txt | lazy scan | decompressed copy | downloaded | downloaded | no |
| systemd journal | journal | none: by content | in memory | no | downloaded | downloaded | no |
| Arrow IPC stream | arrow | as Arrow IPC | converted to Arrow | no | downloaded | downloaded | downloaded |
| Format spec | its name | its match | lazy scan | decompressed copy | downloaded | downloaded | no |
| Read | What it means |
|---|---|
| lazy scan | Scanned where it is. Browsing reads a buffer of rows; queries, sorting and analysis may read the whole input |
| decompressed copy | Decompressed whole into a temporary file of the same format in the temp directory (--temp-dir), which is then scanned lazily. The file is removed on quit (temporary files) |
| converted to Arrow | Read through whole into a temporary Arrow IPC file in the temp directory, which is then scanned lazily. Removed on quit, like a decompressed copy; nothing is cached between sessions |
| in memory | Read whole into memory before the table appears. Past memory_warning in [read] ("1GiB" by default; 0 never asks), datui asks first: big.json: JSON reads 2.10 GB into memory. A model file’s table is one row per tensor, from the header, so it is small however large the model, and is never asked about; a MIDI file is at most 64 MiB |
- Compressed is a
.gz,.zst,.bz2or.xzfile;nomeans it does not open.-c read.decompress_in_memory=truereads compressed CSV, TSV, PSV and text in memory instead. - HTTP(S) is one file at an
http://orhttps://URL.downloadedcopies it to the temp directory first, then reads it as Read says. A model file’s header is fetched by range; a server that sends no ranges gets the download question. - In a bucket is one S3, GCS or Azure object.
in placereads only what is needed with ranged requests: a Parquet or Arrow IPC file’s footer and the rows shown, or a model file’s header.downloadedcopies the object to the temp directory first, after asking, then reads it as Read says; an Arrow stream is converted as it downloads, with no copy of the stream kept. - Bucket prefix is a prefix or glob read as one table. Only Parquet reads
hive partitions; the model files directly under a prefix are read by their
headers. An Arrow prefix scans its IPC files in place and downloads its
streams, one split of a Hugging Face cache as on disk; a glob of Arrow reads
IPC files only. A prefix marked
noopens one object at a time from the cloud source. - Standard input is written to a temporary file first, then read as Read says.
The home screen marks a file row that is not read lazily where it is:
decompresses, converts, in memory or downloads. The details pane and the Info panel’s
Resources tab say how it is read.
Detected by content
The format comes from the first bytes, unless --format or --compression
names it:
| First bytes | Read as |
|---|---|
| Parquet, Arrow IPC or Avro magic number, or an Arrow IPC stream’s schema message | that format |
SQLite format 3 | SQLite; a database of several tables needs --table |
| gzip, zstd, bzip2 or xz magic number | decompressed, then read by what is inside, as below: a text format, CSV or TSV, or lines |
{, the first line an object with __CURSOR and __REALTIME_TIMESTAMP | systemd journal |
[, then JSON | JSON |
{, the first line a whole object | NDJSON |
{, the object open past the first line, JSON so far | JSON |
an NMEA sentence ($GPGGA,) with a checksum that matches, or of a type receivers write | NMEA |
XML whose first element is <gpx | GPX |
a line with 8=FIX, a delimiter and 9= | FIX |
$date, $version, $timescale, $comment, $scope or $var, with an $end | VCD |
a V2000 or V3000 counts line, or M END with a data item or $$$$ | SDF |
| several lines with one field count, two or more, split at tabs | TSV |
| several lines with one field count, two or more, split at commas, quotes where CSV allows them | CSV |
| anything else | lines |
A comma or a tab alone is not a table: a log line with a comma in it stays a
line. An unnamed CSV the first lines do not show as one opens with --format csv; the Info panel’s notes say so. With --format csv, tsv or psv and no
--compression, compression still comes from the first bytes.