Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Value counts

F at the table counts the rows holding each value of the column cursor’s column, with a summary of the column above them. Esc goes back.

Move the cursor with h l or g. ← → on the counts move to the previous or next column, and the table’s cursor goes with them.

Which carrier flies most

Open NYC flights (2013) from Example datasets on the home screen, press g, type carrier, Enter, then F:

Value Counts · carrier · all 336,776 rows
 Rows 336776   Distinct 16   Nulls 0

 carrier   Count▼        %  Cum %
▎UA         58665    17.4%  17.4%   ████████████████████████████████████████████
 B6         54635    16.2%  33.6%   █████████████████████████████████████████
 EV         54173    16.1%  49.7%   ████████████████████████████████████████▋
 DL         48110    14.3%  64.0%   ████████████████████████████████████▏

Four carriers fly almost two thirds of the flights. Enter on B6 shows its 54,635 flights as a drill-down, Group: carrier=B6; Esc comes back to the counts.

→ moves to flight, a number: it opens as a histogram of its counts, under a summary that answers whether the column can be added up:

 Rows 336776   Distinct 3844   Nulls 0   Sum 664096549   Mean 1971.9236   Min 1
 Max 8500

c turns a number’s histogram into the listing of its values, and back.

Histogram

A number column’s counts show as a histogram: 40 bins from the least value to the greatest, or a bin per value when an integer column spans fewer than 40. When the tails reach ten times past the 1st to 99th percentile, the bins span that range instead and the values outside it are counted under the plot: 1,207 values outside p1-p99. The bins are made from the counts, so they are exact wherever the counts are. c shows the listing; another column opens as its own type says.

What the screen shows

PartWhat it says
HeaderThe column, and what was counted: all 336,776 rows, or sample of 100,000 of 657,752 rows
SummaryRows, Distinct (null not among them) and Nulls; Sum, Mean, Min and Max for numbers; Min and Max for dates and times. A sample has no Sum
RowsEach value’s rows, its percent of all of them, the running percent, and a bar beside the most common value’s
∅The nulls, on a row of their own: ranked by their rows when sorted by count, last when sorted by value
other (N values)Past the 1,000 most common values, the rest in one row, without a bar

Values are written as the table writes them: , groups digits here too.

Keys

KeyAction
↑ ↓ or j kMove
PgUp PgDnA page
Home End or GFirst and last row
← → or h lPrevious or next column; a column counted before shows at once
EnterThe rows holding the value, as a drill-down. Esc there comes back
sSort by count or by value; the header’s mark says which
cA number’s histogram, or the listing of its values
aCount every row, when the counts are of a sample
yCopy the counts as TSV: every value with its count, percent and cumulative percent
eExport the same table to a file
? F1Help
EscBack to the table. While every row is being counted, stop and keep the sample

What is read

The counts are of the view: the query, the filters and a drill-down apply. One pass over the column counts every row, in the background; Esc or another column stops it.

A remote view, or one over 100 times the sample size (10,000,000 rows at the default), held in one Parquet or IPC file, is sampled first instead: a few of its row groups are read, and the header says sample of. a then counts every row. A view the sampler would have to read whole anyway, such as a directory of files or a CSV, is counted exactly from the start. The sample size is [analysis] sample_rows; 0 always counts every row.