Value counts
F at the table counts the rows holding each value of the column cursor’s column, with a summary of the column above them. Esc goes back.
Move the cursor with h l or g. ← → on the counts move to the previous or next column, and the table’s cursor goes with them.
Which carrier flies most
Open NYC flights (2013) from Example datasets on the home screen, press
g, type carrier, Enter, then F:
Value Counts · carrier · all 336,776 rows
Rows 336776 Distinct 16 Nulls 0
carrier Count▼ % Cum %
▎UA 58665 17.4% 17.4% ████████████████████████████████████████████
B6 54635 16.2% 33.6% █████████████████████████████████████████
EV 54173 16.1% 49.7% ████████████████████████████████████████▋
DL 48110 14.3% 64.0% ████████████████████████████████████▏
Four carriers fly almost two thirds of the flights. Enter on B6
shows its 54,635 flights as a drill-down, Group: carrier=B6; Esc
comes back to the counts.
→ moves to flight, a number: it opens as a histogram of its
counts, under a summary that answers whether the column can be added up:
Rows 336776 Distinct 3844 Nulls 0 Sum 664096549 Mean 1971.9236 Min 1
Max 8500
c turns a number’s histogram into the listing of its values, and back.
Histogram
A number column’s counts show as a histogram: 40 bins from the least value to
the greatest, or a bin per value when an integer column spans fewer than 40.
When the tails reach ten times past the 1st to 99th percentile, the bins span
that range instead and the values outside it are counted under the plot:
1,207 values outside p1-p99. The bins are made from the counts, so they are
exact wherever the counts are. c shows the listing; another column
opens as its own type says.
What the screen shows
| Part | What it says |
|---|---|
| Header | The column, and what was counted: all 336,776 rows, or sample of 100,000 of 657,752 rows |
| Summary | Rows, Distinct (null not among them) and Nulls; Sum, Mean, Min and Max for numbers; Min and Max for dates and times. A sample has no Sum |
| Rows | Each value’s rows, its percent of all of them, the running percent, and a bar beside the most common value’s |
∅ | The nulls, on a row of their own: ranked by their rows when sorted by count, last when sorted by value |
other (N values) | Past the 1,000 most common values, the rest in one row, without a bar |
Values are written as the table writes them: , groups digits here too.
Keys
| Key | Action |
|---|---|
| ↑ ↓ or j k | Move |
| PgUp PgDn | A page |
| Home End or G | First and last row |
| ← → or h l | Previous or next column; a column counted before shows at once |
| Enter | The rows holding the value, as a drill-down. Esc there comes back |
| s | Sort by count or by value; the header’s mark says which |
| c | A number’s histogram, or the listing of its values |
| a | Count every row, when the counts are of a sample |
| y | Copy the counts as TSV: every value with its count, percent and cumulative percent |
| e | Export the same table to a file |
| ? F1 | Help |
| Esc | Back to the table. While every row is being counted, stop and keep the sample |
What is read
The counts are of the view: the query, the filters and a drill-down apply. One pass over the column counts every row, in the background; Esc or another column stops it.
A remote view, or one over 100 times the sample size (10,000,000 rows at the
default), held in one Parquet or IPC file, is
sampled first instead: a few of its row groups are
read, and the header says sample of. a then counts every row. A
view the sampler would have to read whole anyway, such as a directory of
files or a CSV, is counted exactly from the start. The sample size is
[analysis] sample_rows; 0 always counts every row.