Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Info panel

i at the table opens the Info panel: the dataset’s schema, what its format records, how it is stored and read, and what datui noticed about it. i or Esc closes it. While the dataset has unread notes, i opens on the Notes tab; after that, on Model, Audio, MIDI or VCD for those files, and on Schema for anything else.

The panel’s footer names the keys that work now. ← → (or Tab Shift+Tab) switch tabs from anywhere, and the body always has the keys: the rail marks the schema’s current row. When the schema is taller than the panel, the footer counts the columns out of view.

TabShows
SchemaRow and column counts, column types, schema source, file coverage, and a Parquet file’s per-column compression. A column’s unit, when a delimited format spec read a unit row or the file names one (DataFlash, a DBC dictionary). A dataset a catalog documents adds what each column means (About), the selected column’s note and codes below, and the documentation’s link
DocumentationA dataset a catalog lists, or one inside it, or a file whose format spec documents it: the page Ctrl+E shows on the home screen (Documentation view). ↑ ↓ move, Enter opens a column’s legend, o opens a link in the browser after asking (how), y copies a line’s link or value
Format’s ownWhat the file says besides its rows, named for its format; see the table below
MetadataThe metadata line a delimited format spec names, as key and value; appears for files read through one
ResourcesFile size, how the file is read, buffered memory, and loading measurements
PartitionsPartition columns for a hive-partitioned dataset
NotesSchema differences, skipped files and other findings; appears when there are notes

The file size, and a Parquet, Arrow IPC, Avro or ORC file’s tab, are read in the background the first time the panel opens for a dataset: the few KB at an end of the file the open read too, and none of its rows. Until they arrive the size and the tab read reading...; a file that cannot be read shows why in their place. Remote sources, directories, globs, datasets of several files, and compressed or streamed copies have no file size and no such tab.

Column types

Enter on the Schema tab, or Change type… in the cell menu (a right click on a cell), changes the column’s type for the view: the same names, formats and rules a format spec’s type takes. Combine into datetime… in the cell menu, on a text, date or time column, makes a column as a spec’s derived column does.

ChoiceDoes
A type name (i64, f64, str, bool, …)The column reads as it at once. A value that does not fit is null
date, time, datetimeA format next: each line shows what it makes of the column’s first value (03/04/2024 → 2024-04-03), and a format typed (%d.%m.%Y) is the format
as readThe type the read gave the column
  • The type row shows a changed type in the accent, and the footer says typed zip (or typed 3) before the filters and sort.
  • The first time the change is made, one pass counts the values it made null, and the Notes tab says how many: code: 1 value not i64, read as null.
  • Filters, charts, analysis, find, value counts and exports see the new type; a Parquet export writes it.
  • A saved view keeps each change as a spec’s [columns] entry says it, { "name": "zip", "type": "str" }. On data without the column, the step is left out with a note.
  • A query, a pivot or a melt starts from the data as read, without the changes.

Tabs by format

Each format has its own tab beside Schema, or none, and a file that holds several tables lists them on the home screen as places inside it (shop.db/orders, book.xlsx/Sales) that recents record and Enter opens.

FormatTabLists inside the file
ParquetParquetno
Arrow IPCArrowno
AvroAvrono
ORCORCno
ExcelExceltables
SQLiteSQLitetables
CSV, TSV, PSV, JSON, NDJSONnoneno
SafeTensors, GGUFModelno
NMEAGPStables
GPXGPSno
audioAudiono
MIDIMIDIno
VCDVCDno
FIXFIXno
SDFSDFno
NumPyNumPytables
ELFELFtables
ULogULogtables
DataFlashDataFlashtables
candumpCANtables
systemd journalJournalno
textnoneno

CSV, TSV, PSV, JSON, NDJSON and plain text have no tab of their own: text holds its rows and nothing else. An .xlsx or .xlsm workbook lists its worksheets from its directory, without reading them; an .xls or .xlsb file keeps its worksheet names where only reading the workbook finds them, so it opens its first worksheet and --table or T names another. On the Excel and SQLite tabs, ↑ ↓ move a cursor over the worksheets or tables and Enter opens the one under it in place of this one.

A Hugging Face cache directory lists its splits inside it the same way (cache/test), above the files they are made of.

Every key of the panel is in the keyboard reference. The row count is the dataset’s, not the page’s; D at the table shows the column types in a second header row.

On a CSV, TSV or PSV file, H on the Schema tab reads the first row as data, under column_1, column_2, …, and again as column names. It reads the file again, so the query, filters and sort are cleared, and the panel closes. The footer offers it only for those files.

Model

For a SafeTensors or GGUF file, i opens on the Model tab:

LineShows
FormatSafeTensors or GGUF v3, the tensor count, and the file count for a sharded checkpoint
ParametersThe sum of every tensor’s parameters, in full and abbreviated (8.0B), and the size of the tensor data
TypesEach dtype or quantization type’s share of the parameters, largest first
MetadataKey and value: SafeTensors __metadata__ and the index’s metadata, or GGUF’s key/value pairs

Values are shown whole up to 64 KiB, so a chat template wraps over as many lines as it takes; a longer value, such as a whole tokenizer.json, ends with how much more there is. Arrays of up to 16 items are listed; longer ones, such as a tokenizer’s vocabulary, show their length ([128,256 strings]). Across several files, the first file to name a key gives its value.

Audio

For a WAV, BWF, RF64 or AIFF file, i opens on the Audio tab:

LineShows
FormatWAV, WAV (Broadcast WAV), RF64, AIFF or AIFF-C, the channel count and the sample rate
Samples24-bit integer, with the valid bits when fewer; the encoding; and whether [read] audio_float is on
FramesThe frame count, the length (1:02:03.250) and the size of the sample data
WarningsA data size the file does not hold, frames past the 4,294,967,295 a table holds, or bytes after the last whole frame
Metadatabext.* (description, originator, origination, time reference, coding history), ixml.* (project, scene, take, tape, note) and the iXML document itself, info.* from LIST INFO, AIFF’s name and annotation, then each marker: its time, frame, region length and label

A data size of 0 or a placeholder, as a recorder leaves it, says so: the frames are counted from the file’s size.

MIDI

For a MIDI file, i opens on the MIDI tab, or on Notes first when a note never ends or a file could not be read:

LineShows
FormatMIDI format 1, the timing (480 ticks per quarter, or SMPTE frames), and the track count
LengthThe time of the last event (2:05.250), the event count, and the notes, with how many never end
TempoThe first tempo, and the range and number of changes when it changes; the first time and key signatures
CopyrightThe first copyright notice, when there is one
TracksEach track’s number and name, its events, notes, channels and instrument name

For a directory of songs, the lines are totals and the tempo range, and the list is the files that could not be read, with why.

File format tabs

For a VCD dump, i opens on the VCD tab; for every other format the tab sits beside Schema.

TabLinesList
ParquetRows and row groups, the rows in each; compressed and uncompressed size and the codecs; format version and writer; the footer’s metadata keysColumns: each one’s least and greatest value and its nulls, from the row groups’ statistics, where every group has them
ArrowRecord batches and dictionaries; columns and byte orderMetadata: the schema’s and the footer’s key and value
AvroThe record’s name, fields and codec; its documentationField docs: each field’s documentation, where it has one
ORCRows and stripes, the rows per stripe; format version and compressionMetadata: what the writer kept, key and value
ExcelWorksheets, how many are hidden or hold no cells; the worksheet openedWorksheets: each one’s range and size (A1:D100, 100 × 4), as the opened worksheet’s cells or the other worksheets’ declarations give it
SQLitePage size and pages; schema version, user version and text encoding; tables and views; whether row counts are storedTables: each one’s kind, columns and the rows ANALYZE stored for it. No table is counted to fill it
GPSNMEA: rows of the table opened, sentences and lines. GPX: points, tracks, routes and waypoints. Both: the time span and the latitude and longitude bounds of the rowsSentences (NMEA): each type and how many
VCDTimescale, signal and scope counts; value changes and their time span; $date, $version, $commentSignals: each path with its type, width and identifier
FIXMessages per BeginString; the dictionaries read with the log, each with what it matches and how many messagesTags: each column with its tag number and the names the dictionaries give it, each dictionary’s when they differ
SDFRecords, fields, and how many records are V3000Fields: each one’s type and how many records hold it
NumPyShape, type, order (C or Fortran) and format version; for an archive’s array, the archive and how many arrays it holdsFields: each one’s type, subarray shape and byte offset
ELFClass, machine, type, entry point; bytes in loaded, unwritten sections (flash) and in written ones (RAM); the symbol countSections: each one’s address, size and flags
ULogVersion, topic tables, dropoutsInfo and parameters: each info message, and each parameter’s starting value
DataFlashMessage types with records and defined; records; whether the log has unitsMessages: each type’s records, format characters and length
CANFrames, interfaces, whether timestamps are wall-clock; each dictionary read and what it matches; frames no dictionary namesMessages: each one’s id, frames, signals and comment

A list of more than 10,000 shows the first 10,000 and how many more there are.

Notes

Notes explain what datui found while listing files and reading metadata:

FindingEvidence
Missing columns, conflicting types, or types widened for readingParquet footers
Empty files, large row groups, or many small filesListing and footers
Inconsistent partition keysFile paths
Unreadable footers or files skipped because of their formatListing and metadata reads
Plain files from a Delta, Iceberg or Hudi tableDirectory markers
Rows excluded because a filter/sort column has incompatible typesSchema metadata and the active view

Each note states its scope, such as in all 6,541 footers or in 20,000 of 200,000 footers (sample). Background metadata reads can update these findings. Notes reuse information gathered during loading; they do not scan the data values. For null rates, duplicates and other content checks, use Data Quality.

When all footers are available, a missing-column note may identify the first partition containing a column, or a single partition where it appears. Datui omits these patterns when metadata is sampled or partition names cannot be reliably ordered, such as part=2 and part=10.

Lake tables: datui reads their plain files without applying table metadata. The displayed row count may include deleted rows and superseded versions. See lake table directories.

Each note is one sentence and a line beneath it saying what it is based on, so in 1 of 3 files never stands for files datui has not looked at. The list shows whole notes only, never a claim without its basis; when it is taller than the panel, the corner counts the notes out of view.

Enter on a type-conflict note offers read as text: the column is read from the files that disagree too, at the type each wrote, so the values the conflict hid show, and the marks and the note go. Nothing is listed or read from the footers again. A filter or sort on the column then compares text, and a note says so. See files that disagree.

Most notes describe the dataset as opened. Filter/sort exclusion notes follow the active view and disappear when those controls are cleared. Queries, pivots and drill-downs hide dataset notes until you reset or return to the original level.

Unread notes accent the i key and open on the Notes tab. Set notes_accent = false under [display] in the config to disable the accent; the notes are still collected and the tab still appears.

Measurements

The Resources tab reports work datui can measure:

MetricMeaning
ListingTime spent finding files, plus the number found
FootersTime spent reading Parquet metadata, plus footer reads
Last pageTime spent fetching the visible rows, plus files read
TotalListing time plus footer time

Listing and Footers appear for paths handled by datui’s metadata reader, including Parquet directories and remote sources. Paths delegated directly to Polars may omit those measurements. Last page is available on either route, unless the dataset is already known to be empty.

Footer reads can exceed the number of files: schema and row-count passes may read the same footer more than once. Use the Listing count for dataset size. A dataset opened from cached metadata reads no footers and shows no Footers row.