| DATUI-FORMATS(7) | Miscellaneous | DATUI-FORMATS(7) |
datui-formats - the formats datui reads, and format specs
datui reads 27 formats: Parquet, CSV, TSV, PSV, JSON, NDJSON, Arrow IPC, Avro, ORC, Excel, SafeTensors, GGUF, NMEA, GPX, WAV/AIFF audio, MIDI, SQLite, VCD, FIX, SDF, NumPy, ELF, ULog, DataFlash, candump, plain text, systemd journal, and binary formats you describe in a format spec.
The extension says the format; --format names it when the extension does not, and text piped in is detected by content.
printf 'a,b\n1,2\n' > export.txt datui --format csv export.txt
The home screen marks a file row that is not read lazily where it is: decompresses, converts, in memory or downloads. The details pane and the Info panel's Resources tab say how it is read.
The format comes from the first bytes, unless --format or --compression names it:
A comma or a tab alone is not a table: a log line with a comma in it stays a line. An unnamed CSV the first lines do not show as one opens with --format csv; the Info panel's notes say so. With --format csv, tsv or psv and no --compression, compression still comes from the first bytes.
A format spec is a TOML file that describes a binary format, or a family of delimited text files, so datui opens it as a table.
l2feed.toml
name = "acme.l2feed"
description = "Level 2 capture"
match = { glob = ["*.l2"], magic = "L2FD" }
endian = "le"
[header]
fields = [{ name = "magic", type = "str", size = 4 }, { name = "count", type = "u1" }]
[records]
count = "header.count"
fields = [
{ name = "ts", type = "u8", time = "ns" },
{ name = "symbol", type = "str", size = 8 },
{ name = "side", type = "u1", enum = { 1 = "BUY", 2 = "SELL" } },
{ name = "price", type = "u4", scale = 4, null = "max" },
]
make_day_l2.py writes day.l2, a file in that format:
import ctypes class Header(ctypes.LittleEndianStructure):
_layout_ = "ms"
_pack_ = 1 # no padding between fields
_fields_ = [
("magic", ctypes.c_char * 4),
("count", ctypes.c_uint8),
] class Record(ctypes.LittleEndianStructure):
_layout_ = "ms"
_pack_ = 1
_fields_ = [
("ts", ctypes.c_uint64), # nanoseconds since 1970
("symbol", ctypes.c_char * 8),
("side", ctypes.c_uint8), # 1 BUY, 2 SELL
("price", ctypes.c_uint32), # in ten-thousandths; the largest value is null
] records = [
Record(1709294400000000000, b"MSFT", 1, 4105000),
Record(1709294400000500000, b"AAPL", 2, 0xFFFFFFFF), ] with open("day.l2", "wb") as f:
f.write(bytes(Header(b"L2FD", len(records))))
for record in records:
f.write(bytes(record))
python3 make_day_l2.py datui formats check ./l2feed.toml day.l2 datui --format ./l2feed.toml day.l2 mkdir -p formats cp l2feed.toml formats/ DATUI_FORMATS_PATH=formats datui day.l2 DATUI_FORMATS_PATH=formats datui formats
A spec reads fixed-size records, records that carry their length, several message types in one stream, or compressed blocks; a kind = "delimited" spec reads CSV-like text with lines above its header. Specs are data: no scripts or expressions, and every size read from a file is bounded. --format also takes a spec's http(s)://, s3://, gs:// or az:// URL, fetched once as the open starts; a spec file is at most 1 MiB.
l2feed.toml above is a whole spec. Its table has the columns ts (a datetime), symbol, side and price (a decimal with four places, null where the field holds its largest value), one row per record. Format spec reference lists every field type and key.
Datui reads every *.toml in these places, in order. The first spec of each name wins, as with PATH.
datui formats lists each spec, what it matches (magic L2FD · version 3 · glob *.l2, or no match), the file it came from, any copy of the same name it overrides, and the files that could not be read, with the line and column of each problem. The same places hold FIX log dictionaries: QuickFIX XML files and TOML files of kind = "fix", listed after the specs.
datui formats check SPEC [FILE] checks one spec, by name or file. With a file, it prints the warnings and the first ten rows. Given a QuickFIX dictionary, it checks that, and with a FIX log says how many messages it matches and which of its tags they hold. It exits non-zero on an error, so a repository of specs can run it in CI.
A spec with match.where matches only a file whose header holds those values, so one spec per version can share a glob and a magic. In a spec, in place of its match line:
match = { glob = "*.l2", magic = "L2FD", where = { "header.version" = 3 } }
When two specs match the same way, the first on the search path reads the file. The bar shows 2 formats match, and the Notes tab names the others. A file no spec matches opens as it does without specs; a local file no reader takes either opens in the hex view, where B reads it with a spec and r lines the bytes up in records while you write one.
T on a file of several record types lists the whole file, then each type with its column count, and opens the one picked (day.itch/add), clearing the query, filters and sort.
b on the table picks another spec and reads the file again with it, clearing the query, filters and sort. The list starts with the spec the file was read with, then the others that matched it the same way, then every other spec on the search path for a file (or for a directory of column files). The Notes tab of i says which spec read the file and why (matched by magic L2FD · version 3), the header's values, and any bytes left out.
The home screen and its search name a file the same way. A file whose name says nothing (no extension, or .bin) is matched by magic and where against the first 4 KiB the listing reads from it anyway, up to 256 files a directory; nothing more is read. Its row reads the spec's name, and its details:
A chip with several values ([glob *.l2 *.lvl2]) takes any of them. The listing names a file by its glob without reading it, so a glob-named row shows no where values; the open still checks them. Chips are drawn without brackets where the header tint shows. Text values are quoted where it does not ([kind "A"]), and a magic that is not text is hex (7f 45 4c 46).
Loggers and instruments write a metadata line and a units line above a padded header. A spec of kind = "delimited" holds the CSV options for such a family of files, so they open with no flags: from the command line, from the home screen, compressed, or as a directory.
instrument.toml
name = "acme.instrument-log"
kind = "delimited"
match = { magic = "#device_info" }
comment = "#"
skip_initial_space = true
header_rows = { name = 3, unit = 2 }
metadata_line = 1
[columns]
time = { from = ["Lcl Date", "Lcl Time", "UTCOfst"], as = "datetime" }
bus1volts = { description = "Main bus voltage" }
flight.csv
#device_info, log_version="1.03", model="Unit 7, rev B", serial="123" #yyyy-mm-dd, hh:mm:ss, hh:mm, degrees, volts, deg F
Lcl Date, Lcl Time, UTCOfst, Latitude, bus1volts, T1 Temp
, , , , 25.1, 187.2 2024-03-01, 10:00:00, -05:00, 40.100000, 25.0, 180.0
datui formats check ./instrument.toml flight.csv datui --format ./instrument.toml flight.csv
A delimited spec takes the keys of the config's [[csv]](../reference/settings.md#csv), plus match, kind, the layout keys, [columns], description and documentation.
Lines count from 1 at the top of the file. Each option the spec sets replaces the config's; a flag typed on the command line (--delimiter, --comment, --skip-initial-space, --header-rows, --skip-lines) wins over the spec. The options the spec does not set keep theirs. datui --delimiter ';' formats check SPEC FILE reads the file as an open with those flags would, and names the flags that override the spec. The header lines and the metadata line are the only lines read apart from the CSV reader.
Units
A unit sits beside its column's type on the table's type row (f64 · deg F), in a Unit column on the Info panel's Schema tab, and in chart axis titles (T1 Temp (deg F)). A filter, sort or drill keeps them, and so does a query, pivot or melt for each column it carries unchanged, renamed or not. A column a query computes has no unit, even under the name of one that had.
Metadata
The Info panel's Metadata tab lists the metadata line's pairs, under its leading item when it has one (device_info). A line that is not pairs is shown as it is. For a directory, the first file's line is shown.
Column types
A column of the file takes a type, beside its unit and description:
typed.toml
name = "acme.typed-log"
kind = "delimited"
match = { magic = "#device_info" }
comment = "#"
skip_initial_space = true
header_rows = { name = 3, unit = 2 }
metadata_line = 1
[columns]
"Lcl Date" = { type = "date", format = "%Y-%m-%d" }
Latitude = { type = "f64", description = "GPS latitude" }
bus1volts = { type = "f64", unit = "V" }
typed.csv
#device_info, log_version="1.03" #yyyy-mm-dd, degrees, volts
Lcl Date, Latitude, bus1volts 2024-03-01, 40.100000, 25.0 2024-03-01, , n/a
datui formats check ./typed.toml typed.csv
Use i64 and f64 unless a narrower type is wanted for an export or to hold values to a range. A value is trimmed first, and one that does not fit the type, or is out of an integer type's range, is null. The first time the Info panel opens, one pass counts them, and the Notes tab says how many per column: RPM: 2 values out of range for u8, read as null. A typed column the file does not have is a note, not an error, since the files of a family differ. A typed column is the same type in every file read together, and read.infer_types leaves it alone. type beside from or as is refused: a derived column takes its type from as. The same types, and the derived columns, are on hand in the table: Column types.
Derived columns
A column of the file takes description and unit alone (bus1volts = { description = "Main bus voltage" }), and a derived one takes them beside from and as. They show in the Documentation view. A unit there is documentation only: the type row shows the units line's.
format = "%Y-%m-%d %H:%M:%S" gives the strftime format of the text, a date and a time joined with a space; without it the format is inferred. A value that does not parse is null. The column goes before the first column it is made from, which stays; one named after a column it is made from replaces that column, and its unit. There is no expression language: anything more is a query.
Matching
A delimited spec matches a file whose name says no format datui reads, or says .csv, .tsv or .psv, compressed or not. A directory, or a glob such as 'logs/log_*.csv', is read through the spec its first file with text matches. H on the Info panel's Schema tab reads the file without a header, and without its derived columns.
Several files
Files read together through a spec are matched by column name, so logs from different writer versions stack:
The Notes tab lists the columns not every file has. Columns keep the order the files first have them in.
Garmin and TXi are trademarks of Garmin Ltd. or its subsidiaries; datui is not affiliated with or endorsed by Garmin.
The repository's contrib/formats/garmin-txi.toml reads the data logs a Garmin TXi writes: the airframe line as metadata, the units line, time in UTC, and each column typed. Copy it into ~/.config/datui/formats/ to open the logs, or a directory of them, with no flags. A twin fills the E2 columns and a single leaves them blank. A log written before a GPS fix has blank date and GPS cells.
garmin-log.csv
#airframe_info, log_version="1.03", airframe_name="Example 182", tail_number="N12345", system_id="0000EXAMPLE", unit="GDU1", #yyy-mm-dd, hh:mm:ss, hh:mm, ident, degrees, degrees, ft msl, kt, rpm, deg F, deg F, bool, #
Lcl Date, Lcl Time, UTCOfst, AtvWpt, Latitude, Longitude, AltMSL, IAS, E1 RPM, E1 CHT1, E1 EGT1, OnGrnd, LogIdx
, , , , , , , 0.0, 980.0, 210.0, 1105.0, 1, 1 2024-05-04, 09:12:01, -04:00, KXYZ, 41.0000000, -74.0000000, 350.0, 0.0, 1000.0, 215.0, 1120.0, 1, 2 2024-05-04, 09:12:02, -04:00, KXYZ, 41.0000100, -74.0000100, 350.0, 12.5, 1800.0, 230.0, 1250.0, 0, 3
datui formats check contrib/formats/garmin-txi.toml garmin-log.csv
In a spec's [records]:
checksum = { algo = "crc16-ccitt", field = "crc", from = "len", to = "crc" }
A record checksum covers the bytes from the from field (default: the record's start) up to the to field (default: the checksum's own field). algo is crc16-ccitt, crc16-xmodem, crc16-modbus, crc16-arc, crc32, crc32c, sum8 or xor8; a footer checksum takes the same names.
day.l2.zst, .gz, .bz2 and .xz are decompressed to a temporary file before they are read: the loading screen says Decompressing, then Reading records, and Esc stops either. The glob matches the name without the compression suffix, and magic is read from the decompressed bytes.
A spec's file is read as a lazy scan, or from a decompressed copy when compressed. It is memory-mapped, and only the columns and rows on screen are decoded. Records that are not all one size, and blocks, are indexed by one pass when the file opens. That pass keeps where each record starts (5 bytes a record, up to 64M records), so a query reads every column from there rather than walking the records again, and a file opened again with the same spec is not walked again. Scrolling to the last row of a gigabyte file reads only the rows shown. A sort, filter, query, chart or analysis reads every row of the columns it uses, a batch at a time on the streaming engine ([performance] streaming, on by default). A table holds at most 4,294,967,295 rows; records past that are not shown, and the dataset's notes say so.
A file that grows while it is open keeps the rows it had; open it again to read the rest. A file cut short by another program is refused at the next read rather than read past its end.
Every field type and key a format spec takes. Each example below is a whole spec; datui formats check ./spec.toml checks one.
Type names follow Kaitai Struct <https://kaitai.io>. Widths are in bytes.
A le or be suffix (u4be, s2le, bf2be) overrides the spec's endian.
These keys are for record fields only:
A field takes at most one of time (or date), scale, factor and enum.
A field refers to an earlier one by name, never by an expression. In the header, size = "len" reads an earlier header field; anywhere, header.NAME and footer.NAME do. In a record, "len" reads an earlier field of the same record. A record whose own field gives its size is length_prefixed.
name = "acme.messages"
match = { glob = "*.msg" }
[records]
framing = "length_prefixed"
size = "len" # the field that holds each record's length
size_adjust = 2 # the length leaves out its own two bytes
fields = [{ name = "len", type = "u2" }, { name = "msg", type = "str", size = "rest" }]
Records that are not all one size are walked once when the file opens, and the start of every 1024th is kept, so a scroll anywhere reads from the nearest one.
Variants
name = "acme.orders"
match = { glob = "*.ord" }
[records]
framing = "length_prefixed"
size = "len"
size_adjust = 2
fields = [{ name = "len", type = "u2" }, { name = "kind", type = "str", size = 1 }]
type = "kind"
[[variants]]
name = "add"
when = "A"
fields = [{ name = "ref", type = "u8" }, { name = "shares", type = "u4" }, { name = "price", type = "u4", scale = 4 }]
[[variants]]
name = "exec"
when = ["E", "C"]
fields = [{ name = "ref", type = "u8" }, { name = "shares", type = "u4" }]
fields (or [records.common] fields) are the fields every record starts with. type names the common field that picks the variant: an integer, or text, compared with its padding trimmed. type = { field = "kind", type = "u1" } declares it in place.
All records make one table, with a type column naming each one's variant. A column of a field one variant lacks is null in that variant's rows. A record of a type no variant names shows as ?X when its size is known (length_prefixed); otherwise the read stops there, with a note.
A field name two variants share is one column, so it must be the same field in both: the same type, count, encoding, bits and group, and the same time, scale, factor or enum. Otherwise the spec is refused: variants: `px` is a different field in two variants; one column has one type, so name them apart. A variant's field cannot take a common field's name either (a second field named `kind`), nor be named type.
datui --table add day.ord opens one variant as its own table: only its records and its columns.
On the home screen, each variant of a file a spec reads is a record type: its details read records 2 types (spec) and each type's column count (add 5 · exec 4). Enter opens every record; → lists the record types, one row each at day.ord/add, and Enter on one opens it alone. That path opens the record type on the command line too, and is what recents record.
A spec says what its files mean with the words a catalog uses. Ctrl+E on a file the spec reads, and the Info panel's Documentation tab once it is open, show it in the Documentation view. None of it changes how a file is read.
name = "acme.quotes"
description = "Quotes and trades from the Acme feed"
documentation = "https://example.com/acme-feed.pdf"
match = { glob = "*.acq" }
[records]
framing = "variant"
type = "kind"
fields = [{ name = "kind", type = "u1" }, { name = "ts", type = "u8", time = "ns", description = "When the exchange sent it" }]
[[variants]]
name = "quote"
when = 1
description = "The best bid and offer"
fields = [{ name = "bid", type = "u4", scale = 4, unit = "USD" }, { name = "ask", type = "u4", scale = 4, unit = "USD" }]
[[variants]]
name = "trade"
when = 2
description = "A trade on the book"
fields = [
{ name = "px", type = "u4", scale = 4, description = "Trade price", unit = "USD" },
{ name = "side", type = "u1", enum = { 1 = "BUY", 2 = "SELL" }, description = "The aggressor's side" },
]
A flattened field's note goes to each of its columns, bid_0, bid_1 and on. Named [header] and [footer] fields with a description or unit are listed in sections of their own.
A delimited spec takes description and unit in [columns], for a column of the file or a derived one: temp = { description = "Air temperature", unit = "deg F" }. A column's declared type shows beside its unit: Latitude f64 · deg GPS latitude.
Each text is trimmed, and an empty one is refused. Where a catalog lists the same file, its description and its documentation link stand over the spec's. Where both note a column, each of the catalog's description, unit and values stands over the spec's when the catalog gives it, and the spec's fills the rest. The spec's other notes and its record types stay.
name = "acme.counted"
match = { glob = "*.cnt" }
[records]
count = "footer.n"
fields = [{ name = "v", type = "u2" }]
[footer]
fields = [{ name = "n", type = "u4" }, { name = "crc", type = "u4" }]
checksum = { algo = "crc32", field = "crc" }
The footer is read from the end of the file, so its fields' sizes are written in the spec, or it gives size. Later parts refer to its fields as footer.NAME: a record count, or a block index's offset. checksum checks the bytes before the footer against a footer field; a mismatch is a note.
name = "acme.blocks"
match = { glob = "*.blk" }
[blocks]
header = [{ name = "clen", type = "u4" }, { name = "rawlen", type = "u4" }]
size = "clen"
compression = "zstd"
uncompressed = "rawlen"
[records]
fields = [{ name = "v", type = "u4", delta = "block" }]
The data after the file's header is a run of blocks: a block header, then size bytes of records. The records' framing applies inside each block.
Only the block headers are read when the file opens. A block is decompressed when its rows are first read, up to 256 MiB each, and the last few are kept. A block that will not decompress is left out, with a note.
name = "acme.multicast"
match = { glob = "*.pcap" }
[capture]
header = [{ name = "session", type = "str", size = 10 }, { name = "seq", type = "u8" }, { name = "count", type = "u2" }]
count = "count"
time = "captured"
[records]
fields = [{ name = "price", type = "u4" }]
The file is a pcap or pcapng capture, told apart by its magic. Each UDP payload holds the records, after the payload header; count names the header field counting them, and time adds a column with each packet's capture time. Packets that are not UDP are left out and counted in a note. A capture spec has no [header], [footer] or [blocks].
name = "acme.trades"
match = { glob = "*.bin" }
[files]
path = "{date:%Y%m%d}/{venue}/trades.bin"
[records]
fields = [{ name = "price", type = "f8" }]
datui --format acme.trades store/ reads every file under store/ that the pattern matches as one table, with a column for each part: a date for a part with a format, text otherwise. A file whose part does not parse as its date is left out, with a note.
name = "kdb.trades"
layout = "columns"
endian = "be"
[records]
fields = [{ name = "price", type = "f8" }, { name = "size", type = "s8" }]
With it, datui --format kdb.trades db/trades/ reads db/trades/price and db/trades/size as two columns of one table, as kdb+ splays a table. A [header] describes the start of each file. A glob in a columns spec matches the directory.
When the header lists where each column starts in one file, every field gives offset and the spec reads that one file:
name = "acme.packed"
layout = "columns"
[header]
fields = [{ name = "n", type = "u4" }, { name = "px_off", type = "u4" }]
[records]
count = "header.n"
fields = [{ name = "px", type = "f8", offset = "header.px_off" }]
datui(1), datui-formats(1)
The datui documentation: <https://derekwisong.github.io/datui/>
Report bugs at <https://github.com/derekwisong/datui/issues>.
Derek Wisong and the datui contributors.
Copyright © 2026 Derek Wisong
datui is free software under the MIT License.
| 2026-10-06 | datui 0.4.1 |