Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Analysis

a opens Analysis: Describe, Distribution, Correlation Matrix and Data Quality, over a sample of the table.

A list of tools sits on the right: ↑ ↓ pick a tool and Enter opens it, moving into its pane. Tab or Shift+Tab moves between the list and the result. Esc in the result goes back to the list, and Esc there returns to the table. Results last while the table shows the same rows: a again shows them as you left them.

Analysis runs on the data as you see it, after any query and filters, unless the sample is set to read the source.

Describe

Summary statistics per column, like Polars’ describe: count, nulls, mean, standard deviation, min, 25th percentile, median, 75th percentile and max. Date, datetime, time and duration columns get all of them but the standard deviation, written as the table writes them; text columns get min and max. When they do not all fit, the header counts those out of view (+4 →) and ← → scroll to the last. Distribution scrolls its columns the same way.

On NYC yellow taxis (January 2025), 3,475,226 trips: press a, Enter on Describe, set Random seed to 1 in the Sample form and press Enter. The header reads Describe · sample of 100,000 of 3,475,226 rows. fare_amount has a mean of 16.92 and a median of 12.47, from −595.20 to 950.00: refunds and typos are part of the data. passenger_count and four other columns have 16,000 nulls.

Distribution

Compares each numeric column against fourteen distributions — Normal, Log-Normal, Uniform, Power Law, Exponential, Beta, Gamma, Chi-Squared, Student’s t, Poisson, Bernoulli, Binomial, Geometric and Weibull — and names the one the values are consistent with, along with the Shapiro-Francia normality statistic and p-value, coefficient of variation, outlier count (past 1.5 IQR from the quartiles or 3 standard deviations from the mean, over every value in the sample), skewness and kurtosis.

WhatHow
ParametersMaximum likelihood for normal, log-normal, uniform, exponential, gamma (Minka’s approximation), beta, Weibull, power law (from the smallest value), Poisson, Bernoulli and geometric; moments for chi-squared, Student’s t and binomial
P-valueKolmogorov-Smirnov on up to 500 values, calibrated by 199 samples drawn from the fit and refitted, so estimating the parameters from the data is accounted for. <0.005 when none of the samples came close
VerdictAmong the families not rejected (p of 0.01 or more), the lowest AIC; a simpler family that holds is named instead unless the richer one is decisively better (AIC 10 or more lower). No clear fit when every family is rejected
n/aThe family cannot describe the values: a log-normal of negative values, a Poisson of fractions
ColumnSays
DistributionThe family the values are consistent with; Constant when the column has one value, No clear fit when every family is rejected
P-valueThe fit’s p-value; with no clear fit, the best any family managed
Shapiro-FranciaThe normality statistic W’, from 0 to 1, higher is more normal, on up to 5,000 values
SF p-valueRoyston’s approximation: how likely a W’ this low is if the values were normal
CVCoefficient of variation: standard deviation over mean, spread independent of scale
OutliersValues past the outlier bounds, and their share of the sample
SkewnessAsymmetry: positive is right-tailed, negative left-tailed
KurtosisTail heaviness; 3.0 is normal
ColorWhen
Green or cyanA p-value of 0.05 or more
YellowOutliers 5–20%, |skewness| of 1 or more, kurtosis 1 or more away from 3, CV above 1, or a p-value between 0.01 and 0.05
RedNo clear fit, outliers above 20%, extreme skewness or kurtosis, or a p-value of 0.01 or less

A p-value is how surprising the values would be if they came from the fitted distribution, not the probability that they did. The tests assume independent draws: a time series such as a price over years is dependent, and its histogram need not match any family.

Press Enter on a column for the detail view: the verdict, then a Q-Q plot and a histogram comparing the values with the family chosen in the list, drawn with that family’s fitted parameters. ↑ ↓ choose another family to compare with, which does not change the verdict; s toggles the histogram between linear and log scale; Esc returns to the table.

DetailShows
FitThe verdict and its p-value; it stays as you choose other families
SF, Skew, Kurt, CVAs in the list
Median, Mean, StdThe middle value, the average, and the standard deviation
Q-Q plotThe values against the chosen family’s quantiles; points on the diagonal mean a good match, and where they leave it the data differs
HistogramThe values in bins, with the chosen family’s density drawn over them in the theme’s secondary series color
DistributionsEach family with its p-value, highest first; choosing an n/a family says why

The list’s last row names the histogram’s scale. Log needs positive values: asked for on others, the histogram stays linear and the scale reads Linear in the warning color.

On the same taxi sample, every column reads No clear fit, which is honest for fares, distances and tips. fare_amount has 10,459 outliers (10.5%), skewness 6.92 and kurtosis 279.05; its detail view shows the median, 12.47, as Describe does.

Correlation matrix

Pairwise correlations between every numeric column, colored by strength. Move around with the arrow keys and press Enter on a cell for the pair: the coefficient with a plain reading of it, R², the p-value, and how many row pairs it was computed from, out of the total rows. Enter on a diagonal cell does nothing. Correlation is undefined for a constant column. r draws a new sample from the matrix, not from inside the pair’s detail.

KeyAction
mMethod: Pearson r or Spearman ρ, named in the title
EnterThe selected pair

Spearman is Pearson’s r of the ranks: it reads any relation that only rises or only falls, where Pearson reads a straight line. Both use the rows where the two columns hold a value, and both come from the one run, so m reads nothing. Spearman ranks at most 64 Mi values (rows times numeric columns); past that, as when every row of a large table is read, the matrix has Pearson only and s chooses a smaller sample.

Cells show three decimal places and the pair four. A value that would round to 1 but is not exactly 1 shows as 0.999 (0.9999 in the pair), so 1.000 always means a perfect relation.

On Palmer penguins, first hide rownames, the host’s row number, so it is not treated as a measurement: s, Tab, ↓, v, Enter. Then a, Correlation Matrix, Enter to read all 344 rows. flipper_length_mm against body_mass_g is r = 0.871, strong positive, from 342 of 344 rows.

Data quality

Choose Data Quality to check the rows in scope and get a report: what is likely wrong, what depends on intent, and which columns are clean. It opens on its Setup, which starts from the same sample as every other tool and reads nothing until you run it.

See Check data quality for the workflow and the reference for Setup, keys and metric definitions.

Sampling

Every analysis tool reads the same sample: which rows, how they are picked, how many, and the seed. The first tool you run on a dataset shows the Sample form in its pane with the cursor in it: change a setting or press Enter to run with the form as it stands; Esc goes back to the tool list. After that, every tool you pick runs at once on the same sample. A value the data does not hold, or rows that match nothing, is refused with what the data does hold. s opens the form again from any tool; Enter applies it and runs the tool on screen again, and Esc discards the edit. The other tools’ results go with the old sample, so switching tools compares like with like. The header says what was read: Describe · sample of 100,000 of 36,839,175 rows · source year=2020..2022.

SettingChoices
Rows fromAll rows (the table as shown, with its count), The source, unfiltered (only when a filter or query changes the rows), Partitions, Files, Row range, Time range; a choice appears only when the table has it
MethodRandom (default), Equal per value, First rows, Every row
Per value ofFor Equal per value: the column to split by; partition columns come first
Sample sizeRows, or rows per value for Equal per value, typed over the one shown: 50000, 50,000, 50k, 250k, 2m. Read on Enter; a size that is not one says why. The default is [analysis] sample_rows
Random seedFor Random and Equal per value: any whole number, typed over the one shown; the same seed reads the same rows, so 0 or 1 is a sample anyone can repeat. r draws a new one

Each kind of rows brings its own settings, with what it needs to know:

Rows fromSettings
PartitionsPartition: the column. Values: one value, a list (2019,2021) or an inclusive range (2020..2022) compared in the column’s own type; the values the source holds are listed under it
FilesFiles: their numbers, like 1,3; the numbered files are listed under it and ticked as they are typed, PgUp PgDn scrolls
Row rangeFrom row and To row, inclusive and 1-based, in the order the table shows; it starts as the whole table
Time rangeColumn, From and Before: dates or RFC 3339 timestamps, the Before date not included

Partitions, files and time ranges read the source, ignoring the query and filters.

MethodWhat it reads
RandomA seeded random sample across all the rows chosen
Equal per valueUp to the sample size from each value of a column, so a small partition is represented beside a large one. At most 2,000,000 rows are kept: past that, every value keeps the same smaller number, and the header says what it was lowered to. Refused past 10,000 values
First rowsThe first rows in order: the fastest read, and only the head
Every rowNo sampling
KeyAction
sOpen the Sample form
vView the sample’s rows in the table: sort, filter, query, copy or export them; Esc returns to the tool
rDraw another sample (a new seed)
aRead every row, after confirming the count; sets the method to Every row
EscCancel a run in progress

How a spread sample is read depends on the source:

SourceSampleReads
One Parquet or IPC file, unfiltered50 runs of rows at seeded places across itThe row groups those runs fall in
Anything else: a directory or hive table, a filter, a query, CSVA seeded uniform sample, kept while the rows stream pastEvery row once, holding only the sample

The sort is left out of an analysis read: no statistic depends on it. While a cancelled run is still finishing, no tool starts another read: a, r, v and a new run wait. Data Quality reads the same sample, at the same size, as every other tool.

A view with its own sample (S at the table) is read whole by every tool: the header says Reads the view's sample 100,000 of 3.48M, and s edits the view’s sample, drawing it again before the tool runs; Every row there takes it away.

The sample’s starting size is analysis.sample_rows; 0 starts at every row. For one run, --sample-rows N:

[analysis]
sample_rows = 100000