Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Export data

e writes the view to a file: every row and column as queried, filtered and sorted.

Export a CSV

  1. Open Premier League (2020-21) from Example datasets and run the goals query.
  2. Press e and type goals.csv in Path. The field starts with a suggested name, the dataset’s with -export; typing replaces it, and Enter takes it as it stands.
  3. Press Enter. If the file exists, confirm whether to overwrite it.

The status line says Exported to goals.csv; a long path is cut from its start, so the file name stays. The file holds a header and 380 matches, in the query’s order:

Round,match_date,home,away,goals
4,2020-10-04,Aston Villa,Liverpool,9
22,2021-02-02,Manchester Utd,Southampton,9
14,2020-12-20,Manchester Utd,Leeds United,8

match_date is written as a date because the query casts it; a STRPTIME result alone is a datetime and exports as 2020-10-04T00:00:00.000000.

The file contains all matching rows and displayed columns, not just the page on screen. Numbers use their raw values, without display formatting.

A view with a sample writes the sample’s rows: S on a remote table, then e to sample.parquet, keeps a local sample of it. The export waits until the sample is drawn.

FormatExtensionOptions
CSV.csvDelimiter, header, compression
TSV.tsvHeader, compression; the delimiter is a tab
PSV.psvHeader, compression; the delimiter is a pipe
Parquet.parquet
JSON.jsonCompression; one array
NDJSON.jsonl, .ndjsonCompression; one object per line
Arrow IPC.arrow, .ipc, .feather
Avro.avro

Each export format also offers Source file for a dataset whose files disagree; see below.

The dialog starts on the format datui read the data as, where it writes that format: a TSV file exports as TSV, a Parquet part file with no extension as Parquet. Otherwise it keeps the last format picked.

Excel, ORC, NMEA, GPX, VCD, FIX, SDF and SQLite can be read but not written.

Lists and structs

A by query, SQL ARRAY_AGG or a nested source file gives list, array or struct columns. CSV, TSV and PSV have no such types, so they write each of those cells as JSON text; the dialog says so when the view has one.

Table showsCSV cell
[1, 2][1,2]
["a", "b"]["a","b"]
{1,"a"}{"x":1,"label":"a"}

A null list or struct is an empty field; NaN and infinity inside one are null, which JSON has no other spelling for. Polars reads the text back with str.json_decode.

Parquet, Arrow, JSON and NDJSON keep lists, arrays and structs as they are.

Binary

CSV, TSV, PSV, JSON and NDJSON have no bytes type, so they write a binary value as standard base64 text, inside lists and structs too: hi is written aGk=. Polars reads it back with str.decode("base64"). Parquet, Arrow and Avro keep the bytes.

Durations

CSV has no duration type, so a CSV export writes a duration as ISO 8601 text in seconds: the text JSON and NDJSON exports write, and the same alone or inside a list or struct. It is exact in every unit, down to the nanosecond.

Table showsCSV cell
1h 2m 3s 4msPT3723.004S
-1s -500ms-PT1.5S
1500nsPT0.0000015S
0msP0D

Parquet and Arrow keep the duration type; Avro writes microseconds, below.

Dates past the calendar

A date or datetime past the calendar’s range, such as a sentinel of i64::MIN + 1 microseconds, has no calendar text. CSV, JSON and NDJSON write its stored number, as the table shows it: -9223372036854775807 us since 1970-01-01 UTC. Parquet and Arrow keep the value.

Avro types

Avro keeps booleans, 32- and 64-bit integers and floats, strings, binary, dates, millisecond and microsecond datetimes, lists and structs. Other columns are converted, inside lists and structs too:

ColumnAvro
ArrayList
Categorical, enumString
8- and 16-bit integer32-bit integer
16-bit float32-bit float
Unsigned 32- and 64-bit integer64-bit integer; a value past its range fails the export
Decimal, 128-bit integerString of the exact value, such as 327.68
Nanosecond datetimeMicrosecond datetime; digits past the microsecond are dropped
Datetime with a time zoneThe same instant in UTC, without the zone
Time, duration64-bit integer of microseconds (a time counts from midnight)
NullString

An Avro name is letters, digits and _, and does not start with a digit, so Avro export renames columns and struct fields that are not: any other character becomes _, a leading digit gets a _ in front, and a name that is then taken gets _2, _3 and so on. A valid name is never changed. The dialog says so when the view has such a name.

ColumnAvro field
my colmy_col
2024_2024
délaid_lai
a-b, beside a_ba_b_2

A renamed field keeps its original name as its doc in the file’s schema.

Keys

The dialog opens on the path and takes the keys every dialog takes:

KeyAction
↓ ↑ or Tab Shift+TabMove between format, path and options
← →Change the format, or the compression; in the path, move the cursor
SpaceToggle a checkbox; the next format or compression
Ctrl+P Ctrl+NIn the path: the paths exported to before
EnterExport, from anywhere in the form. On a blank path the form says “Enter a file path.” instead
?Help
EscClose without exporting

Typing a path with a known extension selects the matching format, and a trailing .gz/.zst/.bz2/.xz sets the compression (out.csv.gz selects CSV, gzipped); picking a format afterward rewrites the typed extension to match, so the file’s name and its bytes agree.

Overwrite confirmation defaults to No: ← → or Tab pick Overwrite or No, Enter confirms the one picked, and Esc declines. Declining returns to the form with your path intact. The chart’s export dialog asks the same way.

The CSV Delimiter starts as --delimiter when given, else a comma. Only its first ASCII character counts, and Tab moves focus, so a tab cannot be typed there: pick TSV.

Format lists the formats on its row, the chosen one highlighted, and the rows under it are the chosen format’s options: they come and go as ← → step the format. Where the row is too narrow for them all, it shows the chosen one alone, ‹ TSV ›. A click on a format chooses it.

Overwriting

An export is written to a hidden file beside the destination, named .datui-XXXXXX-<name>, and moved into place only once every byte, including the compressed file’s end, is written and synced to disk. The status line says Exported to only after that move.

CaseResult
The export failsThe old file keeps its bytes and permissions; a new export leaves no file. The hidden file is removed. The dialog comes back as you left it, the reason under the fields: fix the path and press Enter
You confirmed the overwriteThe file is replaced whole. On Linux and macOS it keeps its permission bits
A file appears at the path after you pressed Enter without being askedIt is left alone and the export fails
The file is read-only or you may not write it, or the path is a directory, a pipe or a deviceThe export fails before anything is written
The directory is not writableThe export fails, even where the file itself is writable: the hidden file is made in the directory
The path is a symbolic linkThe file it points to is replaced; the link stays

The replacement is a new file: the old file’s owner, ACLs, extended attributes and hard links are not carried over, and on Windows its attributes are the defaults. If datui is killed during an export, the hidden file can be left behind. A power cut just after the move can leave the old file in place. On filesystems with no atomic no-replace rename and no hard links, such as FAT and some network mounts, a file created at the path in the instant before the move can still be replaced. Chart and Data Quality report exports work the same way.

Large exports

ExportWritten
CSV, TSV and PSV without compression, ParquetStreamed: written in batches as the rows are read; the export never holds the whole view
Compressed CSV, TSV and PSV, JSON, NDJSON, Arrow IPC, AvroThe whole view is read into memory, then written

Streaming needs streaming on in [performance], the default, and a build with the streaming feature; without either, every export reads the whole view first. Streaming bounds what the export holds, not what the view needs: a sort, a by or GROUP BY query, or a join still holds its input in memory before the first row is written. The status line counts the bytes written so far.

Source file

For a dataset with missing columns or conflicting types, Source file adds the original file path to each exported row. This helps trace empty values back to the files that produced them.

The ∅, · and ≠ cell markers all export as null. Use the source path with the original file’s schema to distinguish their causes.

The option is offered only for those datasets, and is off by default. If the data already has a column called source_file, that column is left alone and datui’s is added at the end as source_file_1, or the next free number.

Charts export separately, to PNG or EPS, from the chart view.