Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Formats

datui reads 27 formats: Parquet, CSV, TSV, PSV, JSON, NDJSON, Arrow IPC, Avro, ORC, Excel, SafeTensors, GGUF, NMEA, GPX, WAV/AIFF audio, MIDI, SQLite, VCD, FIX, SDF, NumPy, ELF, ULog, DataFlash, candump, plain text, systemd journal, and binary formats you describe in a format spec.

The extension says the format; --format names it when the extension does not, and text piped in is detected by content.

printf 'a,b\n1,2\n' > export.txt
datui --format csv export.txt

How each format is read

Format--formatExtensionsReadCompressedHTTP(S)In a bucketBucket prefix
Parquetparquet.parquetlazy scannodownloadedin placein place
CSVcsv.csvlazy scandecompressed copydownloadeddownloadedin place
TSVtsv.tsvlazy scandecompressed copydownloadeddownloadedno
PSVpsv.psvlazy scandecompressed copydownloadeddownloadedno
JSONjson.jsonin memorynodownloadeddownloadedno
NDJSONjsonl.jsonl, .ndjsonin memorynodownloadeddownloadedin place
Arrow IPCarrow.arrow, .arrows, .ipc, .featherlazy scannodownloadedin placein place
Avroavro.avroin memorynodownloadeddownloadedno
ORCorc.orcin memorynodownloadeddownloadedno
Excelexcel.xls, .xlsx, .xlsm, .xlsbin memorynodownloadeddownloadedno
SafeTensorssafetensors.safetensors, .safetensors.index.jsonin memorynoin placein placein place
GGUFgguf.ggufin memorynoin placein placein place
NMEAnmea.nmeaconverted to Arrowconverted to Arrowdownloadeddownloadedno
GPXgpx.gpxconverted to Arrowconverted to Arrowdownloadeddownloadedno
WAV, BWF, RF64, AIFFaudio.wav, .wave, .bwf, .rf64, .aif, .aiff, .aifclazy scannodownloadeddownloadedno
MIDImidi.mid, .midi, .smf, .kar, .rmiin memorynodownloadeddownloadedno
SQLitesqlite.db, .db3, .sqlite, .sqlite3lazy scannodownloadeddownloadedno
VCDvcd.vcdconverted to Arrowconverted to Arrowdownloadeddownloadedno
FIXfixnone: by contentconverted to Arrowconverted to Arrowdownloadeddownloadedno
SDFsdf.sdf, .sdconverted to Arrowconverted to Arrowdownloadeddownloadedno
NumPynumpy.npy, .npzlazy scannodownloadeddownloadedno
ELFelf.elf, .axfin memorynodownloadeddownloadedno
ULogulog.ulglazy scannodownloadeddownloadedno
DataFlashdataflashnone: by contentlazy scannodownloadeddownloadedno
candumpcandumpnone: by contentlazy scannodownloadeddownloadedno
Texttext.log, .txtlazy scandecompressed copydownloadeddownloadedno
systemd journaljournalnone: by contentin memorynodownloadeddownloadedno
Arrow IPC streamarrowas Arrow IPCconverted to Arrownodownloadeddownloadeddownloaded
Format specits nameits matchlazy scandecompressed copydownloadeddownloadedno
ReadWhat it means
lazy scanScanned where it is. Browsing reads a buffer of rows; queries, sorting and analysis may read the whole input
decompressed copyDecompressed whole into a temporary file of the same format in the temp directory (--temp-dir), which is then scanned lazily. The file is removed on quit (temporary files)
converted to ArrowRead through whole into a temporary Arrow IPC file in the temp directory, which is then scanned lazily. Removed on quit, like a decompressed copy; nothing is cached between sessions
in memoryRead whole into memory before the table appears. Past memory_warning in [read] ("1GiB" by default; 0 never asks), datui asks first: big.json: JSON reads 2.10 GB into memory. A model file’s table is one row per tensor, from the header, so it is small however large the model, and is never asked about; a MIDI file is at most 64 MiB
  • Compressed is a .gz, .zst, .bz2 or .xz file; no means it does not open. -c read.decompress_in_memory=true reads compressed CSV, TSV, PSV and text in memory instead.
  • HTTP(S) is one file at an http:// or https:// URL. downloaded copies it to the temp directory first, then reads it as Read says. A model file’s header is fetched by range; a server that sends no ranges gets the download question.
  • In a bucket is one S3, GCS or Azure object. in place reads only what is needed with ranged requests: a Parquet or Arrow IPC file’s footer and the rows shown, or a model file’s header. downloaded copies the object to the temp directory first, after asking, then reads it as Read says; an Arrow stream is converted as it downloads, with no copy of the stream kept.
  • Bucket prefix is a prefix or glob read as one table. Only Parquet reads hive partitions; the model files directly under a prefix are read by their headers. An Arrow prefix scans its IPC files in place and downloads its streams, one split of a Hugging Face cache as on disk; a glob of Arrow reads IPC files only. A prefix marked no opens one object at a time from the cloud source.
  • Standard input is written to a temporary file first, then read as Read says.

The home screen marks a file row that is not read lazily where it is: decompresses, converts, in memory or downloads. The details pane and the Info panel’s Resources tab say how it is read.

Detected by content

The format comes from the first bytes, unless --format or --compression names it:

First bytesRead as
Parquet, Arrow IPC or Avro magic number, or an Arrow IPC stream’s schema messagethat format
SQLite format 3SQLite; a database of several tables needs --table
gzip, zstd, bzip2 or xz magic numberdecompressed, then read by what is inside, as below: a text format, CSV or TSV, or lines
{, the first line an object with __CURSOR and __REALTIME_TIMESTAMPsystemd journal
[, then JSONJSON
{, the first line a whole objectNDJSON
{, the object open past the first line, JSON so farJSON
an NMEA sentence ($GPGGA,) with a checksum that matches, or of a type receivers writeNMEA
XML whose first element is <gpxGPX
a line with 8=FIX, a delimiter and 9=FIX
$date, $version, $timescale, $comment, $scope or $var, with an $endVCD
a V2000 or V3000 counts line, or M END with a data item or $$$$SDF
several lines with one field count, two or more, split at tabsTSV
several lines with one field count, two or more, split at commas, quotes where CSV allows themCSV
anything elselines

A comma or a tab alone is not a table: a log line with a comma in it stays a line. An unnamed CSV the first lines do not show as one opens with --format csv; the Info panel’s notes say so. With --format csv, tsv or psv and no --compression, compression still comes from the first bytes.