Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Signals and logs

Recordings, captures and logs open as tables: audio, MIDI, waveforms, GPS tracks, flight and CAN logs, FIX sessions, molecule files, ELF symbol tables and the systemd journal.

FormatExtensionsRead--tableInfo tab
Audio.wav, .wave, .bwf, .rf64, .aif, .aiff, .aifclazyAudio
MIDI.mid, .midi, .smf, .kar, .rmiin memoryMIDI
VCD.vcdconverted to ArrowVCD
NMEA, GPX.nmea, .gpxconverted to ArrowNMEA: fixes, GGA, RMC, VTG, GSA, GSV, GLL, ZDA, sentencesGPS
ULog, DataFlash.ulg; DataFlash by contentlazya topic, a message typeULog, DataFlash
candumpby contentlazyframes, signals, a messageCAN
FIXby contentconverted to ArrowFIX
SDF.sdf, .sdconverted to ArrowSDF
ELF.elf, .axfin memorysymbols, sectionsELF
systemd journalby contentin memoryJournal

How each format is read says what lazy scan, converted to Arrow and in memory mean. A file of several tables opens the home screen inside it, a row per table; downloaded or piped in, it is refused with the tables’ names, and --table picks one.

Audio

make_take_wav.py writes one second of a tone on the left channel:

import ctypes
import math
import wave


class Frame(ctypes.LittleEndianStructure):
    _fields_ = [("left", ctypes.c_int16), ("right", ctypes.c_int16)]


frames = b"".join(bytes(Frame(int(8000 * math.sin(i / 20)), 0)) for i in range(48000))
with wave.open("take.wav", "wb") as w:
    w.setnchannels(2)
    w.setsampwidth(2)  # bytes a sample
    w.setframerate(48000)
    w.writeframes(frames)
python3 make_take_wav.py
datui take.wav
datui -c read.audio_float=true take.wav

An uncompressed audio file opens as a table with one row per sample frame: frame, seconds from the start, and one column per channel. The file is mapped and only the frames on screen are decoded, so a recording of many gigabytes opens at once and scrolls to any point as fast as to the first. The row count comes from the file’s size.

ContainersSamples
WAV, Broadcast WAV, RF64/BW64, WAVE_FORMAT_EXTENSIBLE8-, 16-, 24- and 32-bit integer; 32- and 64-bit float
AIFF, AIFF-C (NONE, twos, sowt, fl32, fl64, in24, in32)The same
  • Channels are ch1, ch2, … An extensible file’s channel mask names them instead: L, R, C, LFE, BL, BR, SL, SR, and so on.
  • Integer samples stay integer: 24-bit is i32, and 8-bit WAV, stored unsigned, is shown signed. [read] audio_float shows them as f32 in [-1, 1]; float samples are never rescaled.
  • A data size of 0 or a placeholder, as a recorder writes until it stops, is read as everything to the end of the file. A size past the end of the file is cut to what the file holds, and the Audio tab says so. A plain WAV past 4 GiB, whose 32-bit data size wrapped, is read to its whole length.
  • Files are recognized by their first bytes too, so a WAV or AIFF with any name opens.
  • Compressed audio (A-law, mu-law, ADPCM, MP3, FLAC) is refused with its name.

Press i for the Audio tab: the format, sample rate, length, the Broadcast WAV (bext), iXML and LIST INFO fields, and the cue /MARK markers with their labels.

A line chart of a long recording draws each step’s lowest and highest sample (Charting), and a full Data Quality run reports clipping, runs of zeros and DC offset.

MIDI

printf 'MThd\0\0\0\6\0\0\0\1\1\340MTrk\0\0\0\26\0\220\74\100\203\140\200\74\0\0\220\100\100\203\140\200\100\0\0\377\57\0' > song.mid
datui song.mid

A Standard MIDI File opens as a table with one row per event, track by track in file order.

ColumnHolds
trackThe track, from 1
tickTicks from the start of the track
secondsSeconds from the start, through the tempo map
kindnote_on, note_off, cc, program, pitch_bend, poly_aftertouch, channel_aftertouch, sysex, sysex_escape, or a meta event: tempo, time_signature, key_signature, track_name, instrument, lyric, marker, cue, text, copyright, end_of_track, …; or a system message a file should not hold but some do: clock, start, stop, song_position, …
channel1-16, as a sequencer numbers them
note, note_nameThe note number and its name, middle C (60) as C4
velocityFor note_on and note_off
controllerThe controller number of a cc
valueThe cc value, program, pressure, pitch bend (-8192 to 8191), tempo in microseconds per quarter, key signature in sharps (negative for flats), a sysex’s length, or a system message’s data
lengthOn a note_on, the seconds until its note_off; null for a note that never ends
textMeta text, tempo as 120 bpm, 6/8, D major, or sysex bytes in hex
  • A note_on at velocity 0 is a note_off, as the specification says.
  • Formats 0 and 1 share one tempo map, from tempo events in any track; each format 2 track keeps its own. With SMPTE timing, seconds follows the frame rate and tempo events do not change it.
  • Meta text is read as UTF-8, or as Latin-1 when it is not.
  • Files are recognized by their first bytes too, so a MIDI file with any name opens, and so does one in a RIFF MIDI (.rmi) wrapper.
  • A track that runs past the end of the file, an event cut short, or fewer tracks than the header says is refused with an error. Of a directory, a file that cannot be read is left out; the Notes tab says so and the MIDI tab lists each one with why.
  • channel, note, velocity and controller are u8, track is u16, and value is i32.
  • A file over 64 MiB is refused, or left out of a directory. An open of more than 10 million events in all is refused.
  • A real-time byte in a track keeps running status, as on the wire; a sysex, meta or system common message cancels it, as the specification says.

Press i for the MIDI tab: format, timing, length, tempo, meter, key and each track’s name, events, notes and channels. Notes that never end are counted on the Notes tab.

VCD

counter.vcd

$timescale 1 ns $end
$scope module tb $end
$var wire 1 ! clk $end
$var wire 4 " count [3:0] $end
$upscope $end
$enddefinitions $end
#0
0!
b0000 "
#5
1!
b0001 "
#10
0!
#15
1!
b0010 "
datui counter.vcd

A VCD file from an HDL simulator or logic analyzer opens as a long table, one row per value change of each signal, read once into a temporary Arrow IPC file.

ColumnHolds
timeThe change’s time: a Duration in nanoseconds for a timescale of 1 ns or coarser; for ps and fs, an integer count of them (a Notes line says which)
signalThe dotted scope path and name with its bit range: tb.dut.count[3:0]
valueThe value as written, a short vector padded to the signal’s width (b1 of a 4-bit signal is 0001; bx is xxxx); a real’s text
intThe value as an integer, when it is binary with no x or z and fits 64 bits
widthThe signal’s width from its $var
  • A file with another name opens when it starts with a VCD section, such as $date or $timescale.
  • An identifier declared at two paths (an alias) gives a row for each.
  • Press i for the VCD tab: timescale, date, version, comments, the number of value changes and their time span, and each signal’s type, width and identifier. The Notes tab counts tokens that are not VCD and changes to undeclared identifiers.
  • A token is at most 1 MiB, and a header holds at most 1,048,576 signals, 256 scopes deep.

The wide table, one row per time and one column per signal, each carried forward from its last change, is this SQL query on counter.vcd above; for another dump, replace the signal names tb.clk and tb.count[3:0] and the columns named for them. It runs on a copy shipped with the docs, counter.vcd:

SELECT time,
       MAX(clk) OVER (PARTITION BY clk_n) AS clk,
       MAX(count) OVER (PARTITION BY count_n) AS count
FROM (
  SELECT *, COUNT(clk) OVER (ORDER BY time) AS clk_n,
            COUNT(count) OVER (ORDER BY time) AS count_n
  FROM (
    SELECT time,
           MAX(CASE WHEN signal = 'tb.clk' THEN value END) AS clk,
           MAX(CASE WHEN signal = 'tb.count[3:0]' THEN value END) AS count
    FROM df GROUP BY time
  )
)
ORDER BY time

The inner GROUP BY is the pivot: one row per time, null where a signal did not change. Each COUNT(...) OVER numbers the runs between changes, and MAX over a run fills it with the change that starts it. Without the fill, Pivot (p) with Index time, Columns signal, Values value and Aggregate last gives the same table with nulls between changes.

GPS logs

make_drive_nmea.py writes three seconds of fixes, each an RMC and a GGA sentence:

def sentence(body):
    """$BODY*CS, where CS is the XOR of the body's bytes, in hex."""
    checksum = 0
    for c in body:
        checksum ^= ord(c)
    return f"${body}*{checksum:02X}\n"


with open("drive.nmea", "w") as f:
    for second in range(3):
        t = f"1200{second:02d}.00"
        f.write(sentence(f"GPRMC,{t},A,4042.6142,N,07400.4168,W,10.5,90.0,010324,,,A"))
        f.write(sentence(f"GPGGA,{t},4042.6142,N,07400.4168,W,1,08,0.9,10.0,M,-34.0,M,,"))

ride.gpx

<?xml version="1.0"?>
<gpx version="1.1" creator="docs"><trk><name>ride</name><trkseg>
<trkpt lat="40.71" lon="-74.00"><ele>10</ele><time>2024-03-01T12:00:00Z</time></trkpt>
<trkpt lat="40.72" lon="-74.01"><ele>12</ele><time>2024-03-01T12:00:05Z</time></trkpt>
</trkseg></trk></gpx>
python3 make_drive_nmea.py
datui drive.nmea
datui --table GGA drive.nmea
datui ride.gpx

An NMEA 0183 log or a GPX file is read once into a temporary Arrow IPC file, then scanned; memory stays at one batch of rows however long the log. Several logs, named together or as a directory of them, open as one table with a file column first; a column one file lacks is null in its rows, and the Notes tab counts across the files. A file with another name, such as capture.log, opens when its first complete line is an NMEA sentence or its first element is <gpx. .nmea.gz and the other compressions are read as they are decompressed.

NMEA opens as one row per fix, merged from each second’s GGA, RMC, VTG and GLL sentences:

ColumnHolds
timeUTC. NMEA dates only RMC and ZDA; every other time of day takes the last date, a day on when it passes midnight
lat, lonDecimal degrees, negative south and west
altMeters above mean sea level (GGA)
speed, courseMeters per second; degrees true
sats, hdopSatellites used and horizontal dilution (GGA)
fixnone, gps, dgps, pps, rtk, rtk float, estimated, manual or simulated
gapSeconds since the fix before; see the gap column
checksum_okEvery sentence of the fix matched its checksum; null when none had one

--table opens one sentence type instead, with all its fields: GGA, RMC, VTG, GSA, GSV (a row per satellite), GLL, ZDA, or sentences (every sentence as written, with its line number, vendor sentences included). On the home screen, Enter on a log opens its fixes and → lists these tables (drive.nmea/GSV). The Info panel’s GPS tab gives the time span, the bounds and the count of each sentence type. Lines that are not NMEA are skipped; the Info panel’s Notes tab counts them, and the sentences that fail their checksum. When the log has sentence types the table on screen does not show, such as GSV beside the fixes, the Schema tab names the other tables with how many sentences each has.

A time of day is dated by the last RMC or ZDA before it. Rows read before the first one are dated back from it when it comes within the first 65,536 rows; when it comes later, those rows keep a null time. A log with neither sentence has no dates, and time is null throughout.

GPX opens as one row per trkpt, rtept and wpt:

ColumnHolds
time, lat, lon, eleThe point’s time (UTC), position and elevation
kindtrack, route or waypoint
track, track_nameThe track or route, numbered from 0 in each kind, and its name
segmentThe track segment, numbered from 0 in its track
gapSeconds since the point before in the same track segment; see the gap column
the restThe point’s other fields (name, sym, sat, hdop…) and each leaf of its <extensions> by its name without the namespace (hr, cad, atemp), as numbers when every value is one

A file cut off mid-element opens with the points before the cut, and says so in Notes.

The gap column

gap is not in the file: datui adds it so a dropout can be sorted and checked. It is the seconds from the row before (the fix before, or the point before in the same GPX track segment) to this one.

gap isWhen
nullThe first fix or point; one without a time
nullTime steps back more than 5 seconds: a receiver reset, or logs joined together
nullNMEA not yet dated and more than an hour passed: whole days could be hidden in it
negative, down to -5Time steps back a little, as a receiver’s clock settles
across midnightAn undated NMEA time earlier than the one before, within the hour, is taken as past midnight

To look at a track:

To seeDo
A rough mapChart, XY, Scatter, X axis lon, Y series lat
DropoutsSort by gap, largest first; or in Data Quality declare a range for gap, such as at most 2, and each dropout is Out of range
Speed spikesThe same for speed; Analysis also counts its outliers

Flight logs

Replace <LOG> with a PX4 ULog or ArduPilot DataFlash log:

datui <LOG>.ulg
datui <LOG>.ulg --table vehicle_status
datui <LOG>.BIN/GPS

Both formats describe their own messages; no format spec is needed. One pass indexes the log, then each table is decoded from a map of the file where it is shown. q at a table comes back to the log’s list of tables without reading the log again.

PX4 ULog (.ulg)
A table per topicNamed for the topic; sensor_accel.0, sensor_accel.1 when it has several instances. timestamp is a duration since boot; nested types are outer.inner, outer[0].inner for an array of them; a number array is an Array column, a char array text. _padding fields are left out
logged_messagestimestamp, level (error, warning, info, …), tag, message
parametersname, type, value, and the timestamp of a change made in flight (null for the value the log started with)
Info tabThe version, dropouts, info messages (sys_name, ver_hw, …) and each parameter’s starting value
ArduPilot DataFlash (.bin)
A table per message typeNamed for the type (GPS, ATT, PARM, …), a column per label, typed by its format character
TimeUS, TimeMSA duration since boot
c, C, e, EHundredths, as a float
LDegrees (latitude, longitude), as a float
aAn Array of 32 i16
UnitsFrom FMTU and UNIT, on the Info panel’s Schema tab; an integer field FMTU gives a multiplier (MULT) is scaled by it
Info tabMessage types and their record counts, formats and lengths
  • A ULog file is known by its first bytes. A DataFlash log is known by its first record, an FMT that defines FMT, whatever it is called.
  • A damaged stretch is passed over to the next ULog sync marker or DataFlash record header; a log cut off mid-message keeps what it holds. The Notes tab says how many bytes were passed over.
  • ULog appended data (written after a crash) is read with the rest.
  • At most 67,108,864 messages are indexed in one log.

CAN logs

candump.log

(1706689000.100000) can0 123#A00F000000000000
(1706689000.200000) can0 123#B80B000000000000
(1706689000.300000) can0 456#01

vehicle.dbc

VERSION ""

BO_ 291 Engine: 8 ECU
 SG_ rpm : 0|16@1+ (0.25,0) [0|16383.75] "rpm" Vector__XXX
datui candump.log
datui candump.log --dict vehicle.dbc --table Engine

One pass indexes the log, then each frame is read from its line where it is shown. A candump log opens by its content, whatever it is called: (1706689000.123456) can0 123#DEADBEEF as candump -l and -L write it (## for CAN FD, #R for a remote request), or can0 123 [4] DE AD BE EF as candump prints it, with or without a timestamp in front.

frames
tsThe timestamp: a datetime for wall-clock time (-l, -ta), a duration for time since the start; null when the line has none
ifaceThe interface: can0, vcan0
idThe id in hex: three digits standard, eight extended
extWhether the id is extended
dlcThe data length code
dataThe data bytes
fd, flagsWhether it is a CAN FD frame, and its flags (BRS, ESI)
kinddata, remote or error

With a dictionary that names the log’s messages, the log opens the home screen inside it, like a directory: a table per message with frames, frames, and signals.

TableColumns
A message, by its DBC namets and a column per signal: factor and offset applied, an integer while they keep it one; value names (VAL_) as text; a multiplexed signal null in the frames its multiplexer does not select. Units are on the Info panel’s Schema tab
signalsOne row per decoded value: ts, message, signal, value (a float) and unit, in time order

Signals in Intel and Motorola byte order, signed and unsigned, and floats (SIG_VALTYPE_) are read; a signal past the end of a short frame is null. Extended multiplexing (SG_MUL_VAL_) is not; the Notes tab says so.

CAN log dictionaries

DBC dictionaries are found where format specs are: the formats directory of the config directory, $DATUI_FORMATS_PATH, and [formats] path. A DBC dictionary there applies to every interface. A TOML file names one for an interface: replace <DBC_FILE> with the dictionary’s name, beside the TOML file or a full path, and <INTERFACE> with the interface:

kind = "dbc"
file = "<DBC_FILE>"
[match]
interface = "<INTERFACE>"

They are read in that order, then --dict FILE; where two name a message of the same id, the later one is read. Press i for the CAN tab: frames, interfaces, the dictionaries read and the frames none of them names, and each message’s id, frames, signals and comment.

FIX logs

session.log

2024-03-01 12:00:00.001 OUT 8=FIX.4.4|9=65|35=D|49=BUYSIDE|56=BROKER|11=ord1|55=MSFT|54=1|38=100|40=2|44=410.5|10=000|
2024-03-01 12:00:00.020 IN 8=FIX.4.4|9=70|35=8|49=BROKER|56=BUYSIDE|11=ord1|55=MSFT|54=1|150=0|39=0|14=0|10=000|
datui session.log

A log of FIX tag=value messages, delimited by SOH, | or ^A, opens as one row per message, read once into a temporary Arrow IPC file. A file of any name opens when a line in its first 4 KiB holds 8=FIX, a delimiter and 9=; messages may be one per line or back to back.

ColumnHolds
prefixThe text before 8=FIX on the line, such as a log timestamp; only when a line has one
directionin or out, from a word in the prefix: IN, OUT, <, >, RECV, SENT and the like
sessionA session in the prefix: FIX.4.4:SENDER->TARGET
a column per tagNamed from the dictionary (35 is MsgType, 55 is Symbol), in the order the tags first appear; a tag no dictionary names keeps its number
<name>_codeBeside an enumerated tag: the code, where the tag’s column shows its name (54=1 is Buy)
<name>_restBeside a tag repeated within a message, as a repeating group’s tags are: a list of its later values; the tag’s column keeps the first
body_length_okTag 9 matches the message’s length; null for a message cut short of tag 10
checksum_okTag 10 matches the message’s checksum; null for a message cut short
  • Prices, quantities and amounts are numbers, integers and sequence numbers i64, UTC timestamps (52, 60) datetimes, dates dates and Y/N booleans, when every value of the tag reads as one; otherwise text.
  • A length-tagged value (95/96 RawData, 90/91, 93/89, 212/213 XmlData and the encoded text fields) is read by its length, so it may hold the delimiter or a newline.
  • A bad message stays: its checks are false, and the Notes tab counts them, the lines with no message, and messages cut short.
  • A message is at most 1 MiB and holds at most 4,096 fields; at most 4,096 tags become columns.
  • Press i for the FIX tab: messages per BeginString, the dictionaries read with the log, and each column’s tag number and the names the dictionaries give it.
  • Binary FIX encodings (SBE, FAST) are not read.

FIX log dictionaries

The built-in dictionary is FIX 4.2, 4.4 and 5.0 SP2 together, the newest version’s names winning. Venues and brokers add their own tags (5000-9999 and 10000 up), so dictionaries can be added: on the format search path, or with --dict FILE.

Form
QuickFIX XML (.xml)A QuickFIX or QuickFIX/J data dictionary, read as it is: its fields, types and enums. It applies to the messages of its version’s BeginString
TOML (.toml, kind = "fix")As below

Continuing from the FIX example above, a TOML dictionary for the broker’s messages, checked against the log:

broker.toml

name = "acme.fix.broker-x"
kind = "fix"
match = { sender = "BROKER", begin_string = "FIX.4.4" }
tags = { 9001 = "AlgoName", 9002 = { name = "Urgency", type = "int", enum = { 1 = "Low", 2 = "High" } } }
datui formats check ./broker.toml session.log
datui --dict broker.toml session.log
Key
nameA namespaced name, such as acme.fix.broker-x
matchOptional. sender (49), target (56), begin_string (8): the dictionary applies only to messages with these values
tagsTag number to a name, or to name, type (int, float, price, qty, string, char, timestamp, date, bool, length, data), enum (code to name) and, for a length tag, data (the tag it sizes)

The built-in dictionary comes first, then each matching dictionary on the search path in order, then --dict; a later one renames a tag or adds to its enums. One log can hold two counterparties that name tag 9001 differently: each message is read with its own, the column falls back to the tag number, and the FIX tab shows both names. datui formats lists dictionaries beside the format specs, and datui formats check NAME [LOG] checks one, and with a log says how many messages it matches and which of its tags they hold.

The built-in dictionary is generated from QuickFIX’s data dictionaries. This product includes software developed by quickfixengine.org (http://www.quickfixengine.org/).

SDF

datui https://raw.githubusercontent.com/rdkit/rdkit/bfc98b529561d11e4a20a64f272c5f6900393cb2/Docs/Book/data/solubility.train.sdf

An SDF (structure-data) file of molecules, as PubChem, ChEMBL and screening libraries publish them, opens as one row per record ($$$$), read once into a temporary Arrow IPC file. The atom and bond blocks are passed over, never held.

ColumnHolds
nameThe molecule’s name, the record’s first line; null when blank
atoms, bondsFrom the counts line, or a V3000 COUNTS line
a column per data itemEach > <FIELD> (also > <FIELD>, > <FIELD> (ID), > 25 <FIELD>, > DT12), in the order first seen; null in a record without it. Integers or floats when every value is one, text otherwise
  • A value of several lines keeps them, joined by newlines.
  • A record that names a field twice keeps the first; the Notes tab counts the rest.
  • A line or value is at most 1 MiB, and a file has at most 4,096 fields.
  • Press i for the SDF tab: the record count, and each field’s type and how many records hold it.
  • Aqueous solubility (SDF) in the home screen’s example datasets is one to try: 1,025 molecules with SOL as a float and SOL_classification as text. Sort by SOL, or filter SOL_classification to (C) high.

ELF

On Linux, /bin/sh is an ELF file:

datui /bin/sh
datui /bin/sh --table sections

On the home screen, Enter on an ELF file opens its symbols and → lists both tables.

ColumnHolds
nameThe symbol’s name; a Rust name demangled, without its hash. C++ names stay mangled
addrIts address (u64)
sizeIts size in bytes
kindfunc, object, section, file, common, tls, ifunc or notype
bindlocal, global, weak or unique
sectionThe section it is in; UND for undefined, ABS for absolute, COMMON
regionflash when its section is loaded and not written (code, constants), ram when it is written (.data, .bss); null for what is not loaded

The sections table has name, addr, size, flags (as readelf writes them: W write, A alloc, X execute, …), kind and region.

  • Sort by size and group by section or region to see what fills flash and RAM.
  • The symbol table is .symtab, or .dynsym for a stripped library.
  • .elf and .axf files open by name; any file that starts with \x7fELF opens too when named on the command line.
  • At most 10 million symbols are read; the Notes tab says how many more there are.

Press i for the ELF tab: class, machine, type, entry point, the bytes in flash and in RAM, and each section’s address, size and flags.

systemd journal

journalctl -o json output is read as the journal, from a pipe or a file:

journalctl -o json -n 1000 | datui
jd() { journalctl -o json "$@" | datui; }
jd -n 100 -p info

Live, as entries are written:

journalctl -o json -f | datui -f -
ColumnWhat
time__REALTIME_TIMESTAMP as a UTC datetime
levelPRIORITY as emerg, alert, crit, err, warning, notice, info, debug, ordered by severity: select where level <= "err" keeps errors and worse, and a sort puts emerg first
_SYSTEMD_UNITThe unit, or SYSLOG_IDENTIFIER when no entry has one
_PID, MESSAGEThen the rest of the fields as they came, and bookkeeping (__CURSOR, __SEQNUM, _BOOT_ID, …) last
  • Every field is a column, one first seen late in the journal included. Values stay text as journalctl writes them; PRIORITY is kept beside level.
  • A MESSAGE journalctl wrote as bytes (not UTF-8, or with control characters) is shown as text, lossily; the Info panel says how many.
  • The Info panel’s Journal tab gives the time span, the entries, and the units, boots and hosts, with the entries per unit.
  • From a pipe, the entries show as they arrive, as for any pipe; a field first seen after the table opened joins as a column when the stream ends, and the Journal tab is read again over every entry.
  • A journal file is read whole into memory. Narrow a large journal with --since, -u or -b.
  • Copy as Python reads a journal file with pl.scan_ndjson and derives the same columns.