Info panel
i at the table opens the Info panel: the dataset’s schema, what its format records, how it is stored and read, and what datui noticed about it. i or Esc closes it. While the dataset has unread notes, i opens on the Notes tab; after that, on Model, Audio, MIDI or VCD for those files, and on Schema for anything else.
The panel’s footer names the keys that work now. ← → (or Tab Shift+Tab) switch tabs from anywhere, and the body always has the keys: the rail marks the schema’s current row. When the schema is taller than the panel, the footer counts the columns out of view.
| Tab | Shows |
|---|---|
| Schema | Row and column counts, column types, schema source, file coverage, and a Parquet file’s per-column compression. A column’s unit, when a delimited format spec read a unit row or the file names one (DataFlash, a DBC dictionary). A dataset a catalog documents adds what each column means (About), the selected column’s note and codes below, and the documentation’s link |
| Documentation | A dataset a catalog lists, or one inside it, or a file whose format spec documents it: the page Ctrl+E shows on the home screen (Documentation view). ↑ ↓ move, Enter opens a column’s legend, o opens a link in the browser after asking (how), y copies a line’s link or value |
| Format’s own | What the file says besides its rows, named for its format; see the table below |
| Metadata | The metadata line a delimited format spec names, as key and value; appears for files read through one |
| Resources | File size, how the file is read, buffered memory, and loading measurements |
| Partitions | Partition columns for a hive-partitioned dataset |
| Notes | Schema differences, skipped files and other findings; appears when there are notes |
The file size, and a Parquet, Arrow IPC, Avro or ORC file’s tab, are read in the
background the first time the panel opens for a dataset: the few KB at an end of the
file the open read too, and none of its rows. Until they arrive the size and the tab
read reading...; a file that cannot be read shows why in their place. Remote
sources, directories, globs, datasets of several files, and compressed or streamed
copies have no file size and no such tab.
Column types
Enter on the Schema tab, or Change type… in the cell menu (a
right click on a cell), changes the column’s type for the view: the same names,
formats and rules a format spec’s type
takes. Combine into datetime… in the cell menu, on a text, date or time
column, makes a column as a spec’s derived column
does.
| Choice | Does |
|---|---|
A type name (i64, f64, str, bool, …) | The column reads as it at once. A value that does not fit is null |
date, time, datetime | A format next: each line shows what it makes of the column’s first value (03/04/2024 → 2024-04-03), and a format typed (%d.%m.%Y) is the format |
as read | The type the read gave the column |
- The type row shows a changed type in the accent, and the footer says
typed zip(ortyped 3) before the filters and sort. - The first time the change is made, one pass counts the values it made null,
and the Notes tab says how many:
code: 1 value not i64, read as null. - Filters, charts, analysis, find, value counts and exports see the new type; a Parquet export writes it.
- A saved view keeps each change as a spec’s
[columns]entry says it,{ "name": "zip", "type": "str" }. On data without the column, the step is left out with a note. - A query, a pivot or a melt starts from the data as read, without the changes.
Tabs by format
Each format has its own tab beside Schema, or none, and a file that holds several
tables lists them on the home screen as places inside it
(shop.db/orders, book.xlsx/Sales) that recents record and Enter opens.
| Format | Tab | Lists inside the file |
|---|---|---|
| Parquet | Parquet | no |
| Arrow IPC | Arrow | no |
| Avro | Avro | no |
| ORC | ORC | no |
| Excel | Excel | tables |
| SQLite | SQLite | tables |
| CSV, TSV, PSV, JSON, NDJSON | none | no |
| SafeTensors, GGUF | Model | no |
| NMEA | GPS | tables |
| GPX | GPS | no |
| audio | Audio | no |
| MIDI | MIDI | no |
| VCD | VCD | no |
| FIX | FIX | no |
| SDF | SDF | no |
| NumPy | NumPy | tables |
| ELF | ELF | tables |
| ULog | ULog | tables |
| DataFlash | DataFlash | tables |
| candump | CAN | tables |
| systemd journal | Journal | no |
| text | none | no |
CSV, TSV, PSV, JSON, NDJSON and plain text have no tab of their own: text holds its rows and nothing else.
An .xlsx or .xlsm workbook lists its worksheets
from its directory, without reading them; an .xls or .xlsb file keeps its worksheet
names where only reading the workbook finds them, so it opens its first worksheet and
--table or T names another. On the Excel and SQLite tabs,
↑ ↓ move a cursor over the worksheets or tables and
Enter opens the one under it in place of this one.
A Hugging Face cache directory lists its
splits inside it the same way (cache/test), above the files they are made of.
Every key of the panel is in the keyboard reference. The row count is the dataset’s, not the page’s; D at the table shows the column types in a second header row.
On a CSV, TSV or PSV file, H on the Schema tab reads the first row
as data, under column_1, column_2, …, and again as column names. It reads
the file again, so the query, filters and sort are cleared, and the panel
closes. The footer offers it only for those files.
Model
For a SafeTensors or GGUF file, i opens on the Model tab:
| Line | Shows |
|---|---|
| Format | SafeTensors or GGUF v3, the tensor count, and the file count for a sharded checkpoint |
| Parameters | The sum of every tensor’s parameters, in full and abbreviated (8.0B), and the size of the tensor data |
| Types | Each dtype or quantization type’s share of the parameters, largest first |
| Metadata | Key and value: SafeTensors __metadata__ and the index’s metadata, or GGUF’s key/value pairs |
Values are shown whole up to 64 KiB, so a chat template wraps over as many
lines as it takes; a longer value, such as a whole tokenizer.json, ends with
how much more there is.
Arrays of up to 16 items are listed; longer ones, such as a tokenizer’s
vocabulary, show their length ([128,256 strings]). Across several files, the
first file to name a key gives its value.
Audio
For a WAV, BWF, RF64 or AIFF file, i opens on the Audio tab:
| Line | Shows |
|---|---|
| Format | WAV, WAV (Broadcast WAV), RF64, AIFF or AIFF-C, the channel count and the sample rate |
| Samples | 24-bit integer, with the valid bits when fewer; the encoding; and whether [read] audio_float is on |
| Frames | The frame count, the length (1:02:03.250) and the size of the sample data |
| Warnings | A data size the file does not hold, frames past the 4,294,967,295 a table holds, or bytes after the last whole frame |
| Metadata | bext.* (description, originator, origination, time reference, coding history), ixml.* (project, scene, take, tape, note) and the iXML document itself, info.* from LIST INFO, AIFF’s name and annotation, then each marker: its time, frame, region length and label |
A data size of 0 or a placeholder, as a recorder leaves it, says so: the frames are counted from the file’s size.
MIDI
For a MIDI file, i opens on the MIDI tab, or on Notes first when a note never ends or a file could not be read:
| Line | Shows |
|---|---|
| Format | MIDI format 1, the timing (480 ticks per quarter, or SMPTE frames), and the track count |
| Length | The time of the last event (2:05.250), the event count, and the notes, with how many never end |
| Tempo | The first tempo, and the range and number of changes when it changes; the first time and key signatures |
| Copyright | The first copyright notice, when there is one |
| Tracks | Each track’s number and name, its events, notes, channels and instrument name |
For a directory of songs, the lines are totals and the tempo range, and the list is the files that could not be read, with why.
File format tabs
For a VCD dump, i opens on the VCD tab; for every other format the tab sits beside Schema.
| Tab | Lines | List |
|---|---|---|
| Parquet | Rows and row groups, the rows in each; compressed and uncompressed size and the codecs; format version and writer; the footer’s metadata keys | Columns: each one’s least and greatest value and its nulls, from the row groups’ statistics, where every group has them |
| Arrow | Record batches and dictionaries; columns and byte order | Metadata: the schema’s and the footer’s key and value |
| Avro | The record’s name, fields and codec; its documentation | Field docs: each field’s documentation, where it has one |
| ORC | Rows and stripes, the rows per stripe; format version and compression | Metadata: what the writer kept, key and value |
| Excel | Worksheets, how many are hidden or hold no cells; the worksheet opened | Worksheets: each one’s range and size (A1:D100, 100 × 4), as the opened worksheet’s cells or the other worksheets’ declarations give it |
| SQLite | Page size and pages; schema version, user version and text encoding; tables and views; whether row counts are stored | Tables: each one’s kind, columns and the rows ANALYZE stored for it. No table is counted to fill it |
| GPS | NMEA: rows of the table opened, sentences and lines. GPX: points, tracks, routes and waypoints. Both: the time span and the latitude and longitude bounds of the rows | Sentences (NMEA): each type and how many |
| VCD | Timescale, signal and scope counts; value changes and their time span; $date, $version, $comment | Signals: each path with its type, width and identifier |
| FIX | Messages per BeginString; the dictionaries read with the log, each with what it matches and how many messages | Tags: each column with its tag number and the names the dictionaries give it, each dictionary’s when they differ |
| SDF | Records, fields, and how many records are V3000 | Fields: each one’s type and how many records hold it |
| NumPy | Shape, type, order (C or Fortran) and format version; for an archive’s array, the archive and how many arrays it holds | Fields: each one’s type, subarray shape and byte offset |
| ELF | Class, machine, type, entry point; bytes in loaded, unwritten sections (flash) and in written ones (RAM); the symbol count | Sections: each one’s address, size and flags |
| ULog | Version, topic tables, dropouts | Info and parameters: each info message, and each parameter’s starting value |
| DataFlash | Message types with records and defined; records; whether the log has units | Messages: each type’s records, format characters and length |
| CAN | Frames, interfaces, whether timestamps are wall-clock; each dictionary read and what it matches; frames no dictionary names | Messages: each one’s id, frames, signals and comment |
A list of more than 10,000 shows the first 10,000 and how many more there are.
Notes
Notes explain what datui found while listing files and reading metadata:
| Finding | Evidence |
|---|---|
| Missing columns, conflicting types, or types widened for reading | Parquet footers |
| Empty files, large row groups, or many small files | Listing and footers |
| Inconsistent partition keys | File paths |
| Unreadable footers or files skipped because of their format | Listing and metadata reads |
| Plain files from a Delta, Iceberg or Hudi table | Directory markers |
| Rows excluded because a filter/sort column has incompatible types | Schema metadata and the active view |
Each note states its scope, such as in all 6,541 footers or
in 20,000 of 200,000 footers (sample). Background metadata reads can update
these findings. Notes reuse information gathered during loading; they do not
scan the data values. For null rates, duplicates and other content checks,
use Data Quality.
When all footers are available, a missing-column note may identify the first
partition containing a column, or a single partition where it appears.
Datui omits these patterns when metadata is sampled or partition names cannot
be reliably ordered, such as part=2 and part=10.
Lake tables: datui reads their plain files without applying table metadata. The displayed row count may include deleted rows and superseded versions. See lake table directories.
Each note is one sentence and a line beneath it saying what it is based on,
so in 1 of 3 files never stands for files datui has not looked at. The list
shows whole notes only, never a claim without its basis; when it is taller
than the panel, the corner counts the notes out of view.
Enter on a type-conflict note offers read as text: the column is read from the files that disagree too, at the type each wrote, so the values the conflict hid show, and the marks and the note go. Nothing is listed or read from the footers again. A filter or sort on the column then compares text, and a note says so. See files that disagree.
Most notes describe the dataset as opened. Filter/sort exclusion notes follow the active view and disappear when those controls are cleared. Queries, pivots and drill-downs hide dataset notes until you reset or return to the original level.
Unread notes accent the i key and open on the Notes tab. Set
notes_accent = false under [display] in the config
to disable the accent; the notes are still collected and the tab still
appears.
Measurements
The Resources tab reports work datui can measure:
| Metric | Meaning |
|---|---|
| Listing | Time spent finding files, plus the number found |
| Footers | Time spent reading Parquet metadata, plus footer reads |
| Last page | Time spent fetching the visible rows, plus files read |
| Total | Listing time plus footer time |
Listing and Footers appear for paths handled by datui’s metadata reader, including Parquet directories and remote sources. Paths delegated directly to Polars may omit those measurements. Last page is available on either route, unless the dataset is already known to be empty.
Footer reads can exceed the number of files: schema and row-count passes may read the same footer more than once. Use the Listing count for dataset size. A dataset opened from cached metadata reads no footers and shows no Footers row.