DATUI-FORMATS(7) Miscellaneous DATUI-FORMATS(7)

datui-formats - the formats datui reads, and format specs

datui reads 27 formats: Parquet, CSV, TSV, PSV, JSON, NDJSON, Arrow IPC, Avro, ORC, Excel, SafeTensors, GGUF, NMEA, GPX, WAV/AIFF audio, MIDI, SQLite, VCD, FIX, SDF, NumPy, ELF, ULog, DataFlash, candump, plain text, systemd journal, and binary formats you describe in a format spec.

The extension says the format; --format names it when the extension does not, and text piped in is detected by content.

printf 'a,b\n1,2\n' > export.txt
datui --format csv export.txt

--format: parquet
Extensions: .parquet
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: in place
Bucket prefix: in place
--format: csv
Extensions: .csv
Read: lazy scan
Compressed: decompressed copy
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: in place
--format: tsv
Extensions: .tsv
Read: lazy scan
Compressed: decompressed copy
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: psv
Extensions: .psv
Read: lazy scan
Compressed: decompressed copy
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: json
Extensions: .json
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: jsonl
Extensions: .jsonl, .ndjson
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: in place
--format: arrow
Extensions: .arrow, .arrows, .ipc, .feather
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: in place
Bucket prefix: in place
--format: avro
Extensions: .avro
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: orc
Extensions: .orc
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: excel
Extensions: .xls, .xlsx, .xlsm, .xlsb
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: safetensors
Extensions: .safetensors, .safetensors.index.json
Read: in memory
Compressed: no
HTTP(S): in place
In a bucket: in place
Bucket prefix: in place
--format: gguf
Extensions: .gguf
Read: in memory
Compressed: no
HTTP(S): in place
In a bucket: in place
Bucket prefix: in place
--format: nmea
Extensions: .nmea
Read: converted to Arrow
Compressed: converted to Arrow
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: gpx
Extensions: .gpx
Read: converted to Arrow
Compressed: converted to Arrow
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: audio
Extensions: .wav, .wave, .bwf, .rf64, .aif, .aiff, .aifc
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: midi
Extensions: .mid, .midi, .smf, .kar, .rmi
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: sqlite
Extensions: .db, .db3, .sqlite, .sqlite3
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: vcd
Extensions: .vcd
Read: converted to Arrow
Compressed: converted to Arrow
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: fix
Extensions: none: by content
Read: converted to Arrow
Compressed: converted to Arrow
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: sdf
Extensions: .sdf, .sd
Read: converted to Arrow
Compressed: converted to Arrow
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: numpy
Extensions: .npy, .npz
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: elf
Extensions: .elf, .axf
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: ulog
Extensions: .ulg
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: dataflash
Extensions: none: by content
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: candump
Extensions: none: by content
Read: lazy scan
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: text
Extensions: .log, .txt
Read: lazy scan
Compressed: decompressed copy
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: journal
Extensions: none: by content
Read: in memory
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
--format: arrow
Extensions: as Arrow IPC
Read: converted to Arrow
Compressed: no
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: downloaded
--format: its name
Extensions: its match
Read: lazy scan
Compressed: decompressed copy
HTTP(S): downloaded
In a bucket: downloaded
Bucket prefix: no
Scanned where it is. Browsing reads a buffer of rows; queries, sorting and analysis may read the whole input
Decompressed whole into a temporary file of the same format in the temp directory (--temp-dir), which is then scanned lazily. The file is removed on quit (temporary files)
Read through whole into a temporary Arrow IPC file in the temp directory, which is then scanned lazily. Removed on quit, like a decompressed copy; nothing is cached between sessions
Read whole into memory before the table appears. Past memory_warning in [read] ("1GiB" by default; 0 never asks), datui asks first: big.json: JSON reads 2.10 GB into memory. A model file's table is one row per tensor, from the header, so it is small however large the model, and is never asked about; a MIDI file is at most 64 MiB

The home screen marks a file row that is not read lazily where it is: decompresses, converts, in memory or downloads. The details pane and the Info panel's Resources tab say how it is read.

The format comes from the first bytes, unless --format or --compression names it:

that format
SQLite format 3
SQLite; a database of several tables needs --table
decompressed, then read by what is inside, as below: a text format, CSV or TSV, or lines
{, the first line an object with __CURSOR and __REALTIME_TIMESTAMP
systemd journal
[, then JSON
JSON
{, the first line a whole object
NDJSON
{, the object open past the first line, JSON so far
JSON
NMEA
GPX
FIX
$date, $version, $timescale, $comment, $scope or $var, with an $end
VCD
SDF
TSV
CSV
lines

A comma or a tab alone is not a table: a log line with a comma in it stays a line. An unnamed CSV the first lines do not show as one opens with --format csv; the Info panel's notes say so. With --format csv, tsv or psv and no --compression, compression still comes from the first bytes.

A format spec is a TOML file that describes a binary format, or a family of delimited text files, so datui opens it as a table.

l2feed.toml

name = "acme.l2feed"
description = "Level 2 capture"
match = { glob = ["*.l2"], magic = "L2FD" }
endian = "le"
[header]
fields = [{ name = "magic", type = "str", size = 4 }, { name = "count", type = "u1" }]
[records]
count = "header.count"
fields = [

{ name = "ts", type = "u8", time = "ns" },
{ name = "symbol", type = "str", size = 8 },
{ name = "side", type = "u1", enum = { 1 = "BUY", 2 = "SELL" } },
{ name = "price", type = "u4", scale = 4, null = "max" }, ]

make_day_l2.py writes day.l2, a file in that format:

import ctypes
class Header(ctypes.LittleEndianStructure):

_layout_ = "ms"
_pack_ = 1 # no padding between fields
_fields_ = [
("magic", ctypes.c_char * 4),
("count", ctypes.c_uint8),
] class Record(ctypes.LittleEndianStructure):
_layout_ = "ms"
_pack_ = 1
_fields_ = [
("ts", ctypes.c_uint64), # nanoseconds since 1970
("symbol", ctypes.c_char * 8),
("side", ctypes.c_uint8), # 1 BUY, 2 SELL
("price", ctypes.c_uint32), # in ten-thousandths; the largest value is null
] records = [
Record(1709294400000000000, b"MSFT", 1, 4105000),
Record(1709294400000500000, b"AAPL", 2, 0xFFFFFFFF), ] with open("day.l2", "wb") as f:
f.write(bytes(Header(b"L2FD", len(records))))
for record in records:
f.write(bytes(record))

python3 make_day_l2.py
datui formats check ./l2feed.toml day.l2
datui --format ./l2feed.toml day.l2
mkdir -p formats
cp l2feed.toml formats/
DATUI_FORMATS_PATH=formats datui day.l2
DATUI_FORMATS_PATH=formats datui formats
Checks the spec and prints the file's header and first rows
Reads the file with that spec file
Reads it with the spec of that name, from the search path
Reads it with the spec whose match it fits, from the search path
Lists the specs and dictionaries on the search path

A spec reads fixed-size records, records that carry their length, several message types in one stream, or compressed blocks; a kind = "delimited" spec reads CSV-like text with lines above its header. Specs are data: no scripts or expressions, and every size read from a file is bounded. --format also takes a spec's http(s)://, s3://, gs:// or az:// URL, fetched once as the open starts; a spec file is at most 1 MiB.

l2feed.toml above is a whole spec. Its table has the columns ts (a datetime), symbol, side and price (a decimal with four places, null where the field holds its largest value), one row per record. Format spec reference lists every field type and key.

The format's name, namespaced: acme.l2feed. --format takes it
Shown by datui formats and the Documentation view
An https:// link to the format's own documentation, for the Documentation view
Which files are this format: glob (a pattern or a list), magic (a string or a list of bytes) at magic_offset (default 0), where (header values)
le (default), be, or auto: big-endian when the magic (at least two bytes) reads reversed. For fields without their own suffix
rows (default): one file of records. columns: a directory with one file per field
[header] fields
Fields read once from the start of the file. Later parts refer to them
[header] size
The header's size when it is more than its fields; a number or a header field
[records] fields
The fields of one record, in order
[records] size
The record's size, at least the fields' sum (the default); the rest is skipped
[records] count
How many records there are, such as "header.count" or "footer.n"
[records] framing
fixed (default), length_prefixed, variant or sync: see records of different sizes
[records] ring
A ring buffer of fixed records: the oldest record's index, such as "header.write_idx"; rows start there and wrap
[records] checksum
A checksum in each record: see checks
[[variants]]
Record layouts a type field picks: see variants
[footer]
Fields at the end of the file: see footer
[blocks]
Records in blocks, each compressed on its own: see blocks
[sections.NAME]
A part of the file, offset and size from the header, that string_at fields point into
[capture]
Records in the UDP payloads of a pcap or pcapng capture: see captures
[files]
A directory tree of the format's files as one table: see a tree of files

Datui reads every *.toml in these places, in order. The first spec of each name wins, as with PATH.

~/.config/datui/formats/
Your own specs
$DATUI_FORMATS_PATH
Directories separated by : (; on Windows), such as a checked-out repository of a team's specs
[formats] path in config
More directories. Lists add up across imported config files

datui formats lists each spec, what it matches (magic L2FD · version 3 · glob *.l2, or no match), the file it came from, any copy of the same name it overrides, and the files that could not be read, with the line and column of each problem. The same places hold FIX log dictionaries: QuickFIX XML files and TOML files of kind = "fix", listed after the specs.

datui formats check SPEC [FILE] checks one spec, by name or file. With a file, it prints the warnings and the first ten rows. Given a QuickFIX dictionary, it checks that, and with a FIX log says how many messages it matches and which of its tags they hold. It exits non-zero on an error, so a repository of specs can run it in CI.

That spec, whatever the file is called
The spec of that name
Read as it is, as before, unless a delimited spec matches a .csv, .tsv or .psv
That spec
That spec

A spec with match.where matches only a file whose header holds those values, so one spec per version can share a glob and a magic. In a spec, in place of its match line:

match = { glob = "*.l2", magic = "L2FD", where = { "header.version" = 3 } }

When two specs match the same way, the first on the search path reads the file. The bar shows 2 formats match, and the Notes tab names the others. A file no spec matches opens as it does without specs; a local file no reader takes either opens in the hex view, where B reads it with a spec and r lines the bytes up in records while you write one.

T on a file of several record types lists the whole file, then each type with its column count, and opens the one picked (day.itch/add), clearing the query, filters and sort.

b on the table picks another spec and reads the file again with it, clearing the query, filters and sort. The list starts with the spec the file was read with, then the others that matched it the same way, then every other spec on the search path for a file (or for a directory of column files). The Notes tab of i says which spec read the file and why (matched by magic L2FD · version 3), the header's values, and any bytes left out.

The home screen and its search name a file the same way. A file whose name says nothing (no extension, or .bin) is matched by magic and where against the first 4 KiB the listing reads from it anyway, up to 256 files a directory; nothing more is read. Its row reads the spec's name, and its details:

acme.l2feed file
The spec's file, cut in the middle to fit: ~/…/formats/l2feed.toml
What named the file, a chip per condition: [magic L2FD] [version 3] when the magic and the header did, [glob *.l2] when the glob did
3 columns (spec) and each column's type, when the spec alone says them (fixed records, no size from the header); otherwise on open
For a spec with variants: 2 types (spec) and each record type's column count (add 5 · cancel 3). → lists the record types

A chip with several values ([glob *.l2 *.lvl2]) takes any of them. The listing names a file by its glob without reading it, so a glob-named row shows no where values; the open still checks them. Chips are drawn without brackets where the header tint shows. Text values are quoted where it does not ([kind "A"]), and a magic that is not text is hex (7f 45 4c 46).

Loggers and instruments write a metadata line and a units line above a padded header. A spec of kind = "delimited" holds the CSV options for such a family of files, so they open with no flags: from the command line, from the home screen, compressed, or as a directory.

instrument.toml

name = "acme.instrument-log"
kind = "delimited"
match = { magic = "#device_info" }
comment = "#"
skip_initial_space = true
header_rows = { name = 3, unit = 2 }
metadata_line = 1
[columns]
time = { from = ["Lcl Date", "Lcl Time", "UTCOfst"], as = "datetime" }
bus1volts = { description = "Main bus voltage" }

flight.csv

#device_info, log_version="1.03", model="Unit 7, rev B", serial="123"
#yyyy-mm-dd, hh:mm:ss, hh:mm, degrees, volts, deg F

Lcl Date, Lcl Time, UTCOfst, Latitude, bus1volts, T1 Temp
, , , , 25.1, 187.2 2024-03-01, 10:00:00, -05:00, 40.100000, 25.0, 180.0

datui formats check ./instrument.toml flight.csv
datui --format ./instrument.toml flight.csv

A delimited spec takes the keys of the config's [[csv]](../reference/settings.md#csv), plus match, kind, the layout keys, [columns], description and documentation.

delimited. Default binary
glob and magic, as for other format specs. magic compares the start of the first line
One character, "tab", "\t" or a code such as "0x1f", as --delimiter takes. Default ,, or the one the file's name implies
Lines that start with it are skipped wherever they are
true: ignore the spaces after a delimiter
{ name = N, unit = M }: the line that names the columns and the line that gives their units. name may be a list of lines, joined with header_join (default a space). A number or a list is name alone. A header line is never data
What joins the pieces of a name from several lines
A line of key="value" or key=value pairs, separated by commas, for the Info panel. It must not be data: above the last header line, within skip_lines, or a comment line
A value, or a list, read as null: "NA", or "COL=-999" for one column
Lines to pass over before the header
[columns]
Column types and derived columns, below, and what columns mean: description and unit

Lines count from 1 at the top of the file. Each option the spec sets replaces the config's; a flag typed on the command line (--delimiter, --comment, --skip-initial-space, --header-rows, --skip-lines) wins over the spec. The options the spec does not set keep theirs. datui --delimiter ';' formats check SPEC FILE reads the file as an open with those flags would, and names the flags that override the spec. The header lines and the metadata line are the only lines read apart from the CSV reader.

Units

A unit sits beside its column's type on the table's type row (f64 · deg F), in a Unit column on the Info panel's Schema tab, and in chart axis titles (T1 Temp (deg F)). A filter, sort or drill keeps them, and so does a query, pivot or melt for each column it carries unchanged, renamed or not. A column a query computes has no unit, even under the name of one that had.

Metadata

The Info panel's Metadata tab lists the metadata line's pairs, under its leading item when it has one (device_info). A line that is not pairs is shown as it is. For a directory, the first file's line is shown.

Column types

A column of the file takes a type, beside its unit and description:

typed.toml

name = "acme.typed-log"
kind = "delimited"
match = { magic = "#device_info" }
comment = "#"
skip_initial_space = true
header_rows = { name = 3, unit = 2 }
metadata_line = 1
[columns]
"Lcl Date" = { type = "date", format = "%Y-%m-%d" }
Latitude = { type = "f64", description = "GPS latitude" }
bus1volts = { type = "f64", unit = "V" }

typed.csv

#device_info, log_version="1.03"
#yyyy-mm-dd, degrees, volts

Lcl Date, Latitude, bus1volts 2024-03-01, 40.100000, 25.0 2024-03-01, , n/a

datui formats check ./typed.toml typed.csv
Text as it is, never typed by read.infer_types: 02134 keeps its zero
true/false or 1/0, in any case
Signed integers
Unsigned integers
Decimals. f32 keeps about 7 significant digits
With format, a strftime format; without, the format is inferred
1d, 2h30m, -1w2d

Use i64 and f64 unless a narrower type is wanted for an export or to hold values to a range. A value is trimmed first, and one that does not fit the type, or is out of an integer type's range, is null. The first time the Info panel opens, one pass counts them, and the Notes tab says how many per column: RPM: 2 values out of range for u8, read as null. A typed column the file does not have is a note, not an error, since the files of a family differ. A typed column is the same type in every file read together, and read.infer_types leaves it alone. type beside from or as is refused: a derived column takes its type from as. The same types, and the derived columns, are on hand in the table: Column types.

Derived columns

from: a date and a time, and optionally a UTC offset such as -05:00, +0530 or -5; or one column of text
Column: A datetime. With an offset it is in UTC
from: one column
Column: A date
from: one column
Column: A time of day

A column of the file takes description and unit alone (bus1volts = { description = "Main bus voltage" }), and a derived one takes them beside from and as. They show in the Documentation view. A unit there is documentation only: the type row shows the units line's.

format = "%Y-%m-%d %H:%M:%S" gives the strftime format of the text, a date and a time joined with a space; without it the format is inferred. A value that does not parse is null. The column goes before the first column it is made from, which stays; one named after a column it is made from replaces that column, and its unit. There is no expression language: anything more is a query.

Matching

A delimited spec matches a file whose name says no format datui reads, or says .csv, .tsv or .psv, compressed or not. A directory, or a glob such as 'logs/log_*.csv', is read through the spec its first file with text matches. H on the Info panel's Schema tab reads the file without a header, and without its derived columns.

Several files

Files read together through a spec are matched by column name, so logs from different writer versions stack:

The column is null in its rows
It takes the type the other files give it. A value further on that is not of that type stops the read, naming the file and the column
The column is f64
The column is text
The first file's unit; a note lists the units seen

The Notes tab lists the columns not every file has. Columns keep the order the files first have them in.

Garmin and TXi are trademarks of Garmin Ltd. or its subsidiaries; datui is not affiliated with or endorsed by Garmin.

The repository's contrib/formats/garmin-txi.toml reads the data logs a Garmin TXi writes: the airframe line as metadata, the units line, time in UTC, and each column typed. Copy it into ~/.config/datui/formats/ to open the logs, or a directory of them, with no flags. A twin fills the E2 columns and a single leaves them blank. A log written before a GPS fix has blank date and GPS cells.

garmin-log.csv

#airframe_info, log_version="1.03", airframe_name="Example 182", tail_number="N12345", system_id="0000EXAMPLE", unit="GDU1",
#yyy-mm-dd, hh:mm:ss,   hh:mm,  ident,      degrees,      degrees,  ft msl,     kt,    rpm,   deg F,   deg F,  bool,      #

Lcl Date, Lcl Time, UTCOfst, AtvWpt, Latitude, Longitude, AltMSL, IAS, E1 RPM, E1 CHT1, E1 EGT1, OnGrnd, LogIdx
, , , , , , , 0.0, 980.0, 210.0, 1105.0, 1, 1 2024-05-04, 09:12:01, -04:00, KXYZ, 41.0000000, -74.0000000, 350.0, 0.0, 1000.0, 215.0, 1120.0, 1, 2 2024-05-04, 09:12:02, -04:00, KXYZ, 41.0000100, -74.0000100, 350.0, 12.5, 1800.0, 230.0, 1250.0, 0, 3

datui formats check contrib/formats/garmin-txi.toml garmin-log.csv

The spec is refused, with its line and column
The open fails, showing the bytes found
The whole records open; a note shows the bytes left over
The whole records open, with a note
The records before it open; a note says where the rest was left out
A checksum_ok column, true or false for each record, rather than an error
The file is downloaded first, then read

In a spec's [records]:

checksum = { algo = "crc16-ccitt", field = "crc", from = "len", to = "crc" }

A record checksum covers the bytes from the from field (default: the record's start) up to the to field (default: the checksum's own field). algo is crc16-ccitt, crc16-xmodem, crc16-modbus, crc16-arc, crc32, crc32c, sum8 or xor8; a footer checksum takes the same names.

day.l2.zst, .gz, .bz2 and .xz are decompressed to a temporary file before they are read: the loading screen says Decompressing, then Reading records, and Esc stops either. The glob matches the name without the compression suffix, and magic is read from the decompressed bytes.

A spec's file is read as a lazy scan, or from a decompressed copy when compressed. It is memory-mapped, and only the columns and rows on screen are decoded. Records that are not all one size, and blocks, are indexed by one pass when the file opens. That pass keeps where each record starts (5 bytes a record, up to 64M records), so a query reads every column from there rather than walking the records again, and a file opened again with the same spec is not walked again. Scrolling to the last row of a gigabyte file reads only the rows shown. A sort, filter, query, chart or analysis reads every row of the columns it uses, a batch at a time on the streaming engine ([performance] streaming, on by default). A table holds at most 4,294,967,295 rows; records past that are not shown, and the dataset's notes say so.

A file that grows while it is open keeps the rows it had; open it again to read the rest. A file cut short by another program is refused at the next read rather than read past its end.

Every field type and key a format spec takes. Each example below is a whole spec; datui formats check ./spec.toml checks one.

Type names follow Kaitai Struct <https://kaitai.io>. Widths are in bytes.

Unsigned and signed integers of 1 to 8 bytes, u3 and s6 included
Floats; f2 is a half float
A bfloat16
LEB128 varints, vs zigzag-encoded. Records only
One byte, nonzero is true
Text of size bytes, its NUL and space padding trimmed
Text up to a NUL, at most size bytes when given. Records only
Raw bytes of size
size bytes skipped, no column

A le or be suffix (u4be, s2le, bf2be) overrides the spec's endian.

The column's name. Every field but pad has one
Example: "price"
Bytes of a str, strz, bytes or pad: a number, a header or footer field, an earlier field of the record, or rest, what is left of the record
Example: 8, "header.len", "len", "rest"
Added to a size read from a field
Example: -4
Of a str or strz: utf8 (default), latin1, utf16le or utf16be
Example: "latin1"
That many values side by side: one Array column
Example: 10
With count (up to 1024), columns name_0 to name_9 instead of an Array
Example: true
A stored value that means no value: the type's smallest or largest value, a NaN, or this number, which the type must be able to hold
Example: "min", "max", "nan", -1
Implied decimal places: the integer becomes a Decimal
Example: 4
value * factor + offset, as a float
Example: 0.1, -40.0
Codes and labels; an unlisted code reads as its number
Example: { 1 = "BUY", 2 = "SELL" }
A count of days, s, ms, us or ns: a datetime, or a date for days. A float counts fractions too
Example: "ns"
What the count is since. Default 1970-01-01
Example: 2000-01-01
An integer such as 20240102, as a date
Example: "yyyymmdd"
With time, a count since midnight: a time of day
Example: true
The day those times are on, from a header field that reads as a date (date = "yyyymmdd", time = "days"), a datetime, or text such as 2024-01-02: a datetime
Example: "header.trade_date"
In the columns layout, the file in the directory holding the field. Default: its name
Example: "px.dat"
In the columns layout, where the column starts in one file: see columns layout
Example: "header.px_off"
An integer indexes a list of symbols in a file beside the data: a categorical. format is lines (default), nul or str:N
Example: { file = "../sym", format = "lines" }
What the column means, for the Documentation view. Not read
Example: "Limit price"
The column's unit, for the Documentation view. Not read
Example: "USD"

These keys are for record fields only:

Each value is the change from the record before; the running sum is shown. "block" starts the sum again in each block
Example: true, "block"
Bit fields of an integer, each its own column. A width of 1 is a bool
Example: [{ name = "valid", bit = 0 }, { name = "mode", bit = 4, width = 3, enum = { 0 = "IDLE" } }]
A counted run of items: a List of Structs. Takes no type
Example: { count = "n_levels", fields = [...] }
An unsigned offset into a [sections.strings] part of the file, where NUL-terminated text is
Example: "strings"

A field takes at most one of time (or date), scale, factor and enum.

A field refers to an earlier one by name, never by an expression. In the header, size = "len" reads an earlier header field; anywhere, header.NAME and footer.NAME do. In a record, "len" reads an earlier field of the same record. A record whose own field gives its size is length_prefixed.

name = "acme.messages"
match = { glob = "*.msg" }
[records]
framing = "length_prefixed"
size = "len"          # the field that holds each record's length
size_adjust = 2       # the length leaves out its own two bytes
fields = [{ name = "len", type = "u2" }, { name = "msg", type = "str", size = "rest" }]
Every record takes size, or what its fields take
A field of the record gives its size: size = "len", with size_adjust
The variant the type field picks gives the size: its size, or what its fields take
Each record starts with the sync marker ("1ACFFC1D", "0xEB90" or a list of bytes); bytes between records are skipped and counted in a note
true: the length is written again after the record, as Fortran unformatted files do. Needs size = "len"
Each record starts at a multiple of this (1 to 65536), counted from the first: align = 2 for IFF and RIFF chunks

Records that are not all one size are walked once when the file opens, and the start of every 1024th is kept, so a scroll anywhere reads from the nearest one.

Variants

name = "acme.orders"
match = { glob = "*.ord" }
[records]
framing = "length_prefixed"
size = "len"
size_adjust = 2
fields = [{ name = "len", type = "u2" }, { name = "kind", type = "str", size = 1 }]
type = "kind"
[[variants]]
name = "add"
when = "A"
fields = [{ name = "ref", type = "u8" }, { name = "shares", type = "u4" }, { name = "price", type = "u4", scale = 4 }]
[[variants]]
name = "exec"
when = ["E", "C"]
fields = [{ name = "ref", type = "u8" }, { name = "shares", type = "u4" }]

fields (or [records.common] fields) are the fields every record starts with. type names the common field that picks the variant: an integer, or text, compared with its padding trimmed. type = { field = "kind", type = "u1" } declares it in place.

Shown in the type column
The type value, or a list of them, that picks it
The fields after the common ones
The whole record's size, when more than its fields take
What a record of this type is, for the Documentation view

All records make one table, with a type column naming each one's variant. A column of a field one variant lacks is null in that variant's rows. A record of a type no variant names shows as ?X when its size is known (length_prefixed); otherwise the read stops there, with a note.

A field name two variants share is one column, so it must be the same field in both: the same type, count, encoding, bits and group, and the same time, scale, factor or enum. Otherwise the spec is refused: variants: `px` is a different field in two variants; one column has one type, so name them apart. A variant's field cannot take a common field's name either (a second field named `kind`), nor be named type.

datui --table add day.ord opens one variant as its own table: only its records and its columns.

On the home screen, each variant of a file a spec reads is a record type: its details read records 2 types (spec) and each type's column count (add 5 · exec 4). Enter opens every record; → lists the record types, one row each at day.ord/add, and Enter on one opens it alone. That path opens the record type on the command line too, and is what recents record.

A spec says what its files mean with the words a catalog uses. Ctrl+E on a file the spec reads, and the Info panel's Documentation tab once it is open, show it in the Documentation view. None of it changes how a file is read.

name = "acme.quotes"
description = "Quotes and trades from the Acme feed"
documentation = "https://example.com/acme-feed.pdf"
match = { glob = "*.acq" }
[records]
framing = "variant"
type = "kind"
fields = [{ name = "kind", type = "u1" }, { name = "ts", type = "u8", time = "ns", description = "When the exchange sent it" }]
[[variants]]
name = "quote"
when = 1
description = "The best bid and offer"
fields = [{ name = "bid", type = "u4", scale = 4, unit = "USD" }, { name = "ask", type = "u4", scale = 4, unit = "USD" }]
[[variants]]
name = "trade"
when = 2
description = "A trade on the book"
fields = [

{ name = "px", type = "u4", scale = 4, description = "Trade price", unit = "USD" },
{ name = "side", type = "u1", enum = { 1 = "BUY", 2 = "SELL" }, description = "The aggressor's side" }, ]
Where: The spec, a variant, a field
Says: What the format, the record type, the column or the header or footer field is
Where: The spec
Says: An https:// link to the format's own documentation
Where: A field
Says: The column's, or the header or footer field's, unit
Where: A record field
Says: Its codes and labels are the column's value legend

A flattened field's note goes to each of its columns, bid_0, bid_1 and on. Named [header] and [footer] fields with a description or unit are listed in sections of their own.

A delimited spec takes description and unit in [columns], for a column of the file or a derived one: temp = { description = "Air temperature", unit = "deg F" }. A column's declared type shows beside its unit: Latitude f64 · deg GPS latitude.

Each text is trimmed, and an empty one is refused. Where a catalog lists the same file, its description and its documentation link stand over the spec's. Where both note a column, each of the catalog's description, unit and values stands over the spec's when the catalog gives it, and the spec's fills the rest. The spec's other notes and its record types stay.

name = "acme.counted"
match = { glob = "*.cnt" }
[records]
count = "footer.n"
fields = [{ name = "v", type = "u2" }]
[footer]
fields = [{ name = "n", type = "u4" }, { name = "crc", type = "u4" }]
checksum = { algo = "crc32", field = "crc" }

The footer is read from the end of the file, so its fields' sizes are written in the spec, or it gives size. Later parts refer to its fields as footer.NAME: a record count, or a block index's offset. checksum checks the bytes before the footer against a footer field; a mismatch is a note.

name = "acme.blocks"
match = { glob = "*.blk" }
[blocks]
header = [{ name = "clen", type = "u4" }, { name = "rawlen", type = "u4" }]
size = "clen"
compression = "zstd"
uncompressed = "rawlen"
[records]
fields = [{ name = "v", type = "u4", delta = "block" }]

The data after the file's header is a run of blocks: a block header, then size bytes of records. The records' framing applies inside each block.

The fields at the start of each block
The bytes after the block header: a number or a block header field
none (default), gzip, deflate, zlib, zstd, lz4, lz4_block, snappy, snappy_framed, brotli, bzip2 or xz. Or a code in the block header: { field = "codec", values = { 0 = "none", 1 = "zstd" } }
The block header field with the decompressed size. Required for lz4_block
The block header field counting its records; missing records are null
{ at = "footer.index_off", count = "footer.n_blocks", fields = [...] }: a block index read instead of walking the blocks. Its entries need an integer offset, and may give rows

Only the block headers are read when the file opens. A block is decompressed when its rows are first read, up to 256 MiB each, and the last few are kept. A block that will not decompress is left out, with a note.

name = "acme.multicast"
match = { glob = "*.pcap" }
[capture]
header = [{ name = "session", type = "str", size = 10 }, { name = "seq", type = "u8" }, { name = "count", type = "u2" }]
count = "count"
time = "captured"
[records]
fields = [{ name = "price", type = "u4" }]

The file is a pcap or pcapng capture, told apart by its magic. Each UDP payload holds the records, after the payload header; count names the header field counting them, and time adds a column with each packet's capture time. Packets that are not UDP are left out and counted in a note. A capture spec has no [header], [footer] or [blocks].

name = "acme.trades"
match = { glob = "*.bin" }
[files]
path = "{date:%Y%m%d}/{venue}/trades.bin"
[records]
fields = [{ name = "price", type = "f8" }]

datui --format acme.trades store/ reads every file under store/ that the pattern matches as one table, with a column for each part: a date for a part with a format, text otherwise. A file whose part does not parse as its date is left out, with a note.

name = "kdb.trades"
layout = "columns"
endian = "be"
[records]
fields = [{ name = "price", type = "f8" }, { name = "size", type = "s8" }]

With it, datui --format kdb.trades db/trades/ reads db/trades/price and db/trades/size as two columns of one table, as kdb+ splays a table. A [header] describes the start of each file. A glob in a columns spec matches the directory.

When the header lists where each column starts in one file, every field gives offset and the spec reads that one file:

name = "acme.packed"
layout = "columns"
[header]
fields = [{ name = "n", type = "u4" }, { name = "px_off", type = "u4" }]
[records]
count = "header.n"
fields = [{ name = "px", type = "f8", offset = "header.px_off" }]

datui(1), datui-formats(1)

The datui documentation: <https://derekwisong.github.io/datui/>

Report bugs at <https://github.com/derekwisong/datui/issues>.

Derek Wisong and the datui contributors.

Copyright © 2026 Derek Wisong

datui is free software under the MIT License.

2026-10-06 datui 0.4.0