Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Fuzzing

Datui fuzzes the hand-written parsers and matchers that run on untrusted input, using cargo-fuzz and libFuzzer. The targets live in fuzz/.

What is fuzzed

TargetSurfaceWhat it checks
parse_queryquery::parse_queryA tokenizer and recursive-descent parser that slices token vectors by index. Malformed input must return Err, never panic.
sql_group_plansql_group::planReads every SQL statement to decide whether a GROUP BY result drills: resolves keys by ordinal, alias and expression and writes a statement of its own. Any text must give a plan or None, never panic.
number_formatnumfmt::NumberFormatwidth_* computes a display width arithmetically, write_* renders into a fixed 64-byte stack buffer and returns the width it produced. The two must agree, and both must equal the characters actually appended. Table columns are sized from these numbers, so a disagreement corrupts the layout instead of failing visibly.
fuzzy_matchfuzzy::best_matchReturned positions must be valid, strictly ascending character indices into the haystack, one per needle character. The home screen highlights matches by indexing with them.
glob_matchnumfmt::GlobA backtracking wildcard matcher, checked for hangs and for its wildcard-free fast path agreeing with equality.
ipc_stream_headipc_stream::is_stream_head, stdin::sniffReads a length from a file’s first bytes and checks the flatbuffer it names is an Arrow schema message, on any file being opened and any pipe. Any bytes must give an answer, never a panic (Polars’ own schema reader panics on some column types), and a pipe that is a stream must be read as one.
config_parseconfig::AppConfig, config::ColorParserValidation and merging of user TOML, and color strings that get sliced by byte offset after a byte-length check.
audio_headeraudio::read_header, audio::AudioSourceThe WAV, RF64 and AIFF chunk walker, which slices by sizes, counts and offsets the file states, and the sample decoder, which reads at offsets worked out from the header. A corrupt header must be an error, never a panic or an allocation sized by the file; a header that parses must give frames inside the file, and decoding them must work.
midi_filemidi::parse, midi::buildA hand-written Standard MIDI File parser that slices by chunk lengths, variable-length deltas and event lengths read from the file, and keeps running status between events. A corrupt file must be an error, never a panic or an allocation sized by the file; a file that parses must build its table, one row per event.
model_headermodel_files::read_safetensors, model_files::read_gguf, model_files::read_header_ranged_fromHand-written readers for model file headers that allocate and skip by lengths read from the file. Every input goes to both; a corrupt header must be an error, never a panic or an allocation sized by the file, and a header that parses must build its table. Read again by range, in ranges of a few bytes, each finds the same header or fails as the file reader does.
format_specformats::Spec, fixed_recordsA binary format spec and a file it reads, split at the first NUL byte. A spec parses or fails with a line and column; a file reads or fails; every row the reader counts decodes, and a window of the rows matches the same rows read from the start.
gps_parsegps::nmea::NmeaReader, gps::gpx::GpxReaderHand-written readers for GPS logs that take the file a piece at a time. The first byte picks the NMEA table and the size of the pieces, so every line, tag and entity is cut somewhere. Never a panic, a frame of another schema, or a coordinate off the globe; every length is bounded by the reader.
vcd_parsevcd::VcdReaderA hand-written reader of VCD tokens that takes the file a piece at a time, with token, header text, scope depth and signal bounds. The first byte picks the piece size. Never a panic or a batch of another schema; the rows add up and the header stays within its bounds.
fix_parsefix::FixReaderA hand-written reader of FIX messages a piece at a time: tag and value framing, length-tagged values read by the length they state, checksums and body lengths. The first byte picks the piece size. Never a panic; the rows add up to the messages, and the last batch renames and types into a frame that collects.
fix_dictfix::dict::DictionaryQuickFIX XML data dictionaries, read by a hand-written scanner of tags and attributes, and the TOML form. Any text must parse or fail, never panic, and a dictionary that parses keeps its names within bounds.
sdf_parsesdf::SdfReaderA hand-written reader of SDF records a piece at a time, with line, value and field bounds. The first byte picks the piece size. Never a panic; the rows add up to the records and no value passes its bound.
can_parsedbc::parse, candump::index, candump::DecodedA DBC dictionary and a candump log, split at the first NUL. DBC statements are read across lines with bounded counts and lengths; each line of the log is read as a frame; each message the dictionary names is decoded from its frames, Intel and Motorola bits, signed and multiplexed. Never a panic, and every decoded table has a row per frame.
text_lineslines::guess, lines::LineIndex, lines::LinesWhat text no signature claims is (JSON, CSV or TSV on evidence, lines otherwise), from a head whole or cut anywhere, and the line index over it. Any bytes must give an answer, never a panic; the index has a row per line, the same whether built at once or read on as the bytes grow, and every row decodes.
elf_symbolself::read, elf::demangleELF headers, section headers and symbol tables read by the object crate at offsets and sizes the file gives, and the rows and Info tab built from them. A corrupt file must be an error, never a panic or an allocation sized by the file; a file that reads gives its two tables.
flight_logulog::index, dataflash::index, indexed::IndexedRecordsULog and DataFlash logs walked by the sizes and type ids they give, their message types defined by their own format records; the first byte picks the reader. A corrupt log must be passed over or end the pass, never a panic or an allocation sized by the file, and every message the index records decodes.
numpy_headernumpy::parse_literal, numpy::parse_header, numpy::open_inThe .npy header’s Python dict literal, parsed by hand, and the structured types in it, whose offsets, itemsizes and subarray shapes come from the file. A corrupt header must be an error, never a panic or an allocation sized by the file; a header that parses must give columns inside the bytes on hand, and its rows must decode.
hex_inputhex_view::parse_offset, hex_view::parse_pattern, hex_view::findThe hex view’s offset and byte-pattern parsers and its search, on what was typed (up to the first NUL) and the file after it. An offset that parses is inside the file, a pattern that parses fits its bound, every match reported is one and the first is the one a plain scan finds, and the byte inspector reads any bytes without a panic.

Layout

PathWhat
fuzz/src/<target>.rsThe target’s body: run, its input type and every check
fuzz/fuzz_targets/<target>.rslibfuzzer-sys decodes the input and calls run
tests/fuzz_corpus_test.rsReplays fuzz/corpus/<target>/ through the same run, as an ordinary test

The test decodes each input as libfuzzer-sys does (Arbitrary::arbitrary_take_rest over Unstructured), runs the empty input as libFuzzer does, and fails on any panic, even one the code catches, as the fuzzer’s panic hook does. Its decoding matches the fuzzer’s only while Cargo.lock and fuzz/Cargo.lock resolve the same arbitrary; the test checks that too. Put new checks in run. A new target needs a line in the test, which fails until its corpus is replayed.

Running

Replay every committed corpus input, without building the fuzzers:

./scripts/dev/test.sh integration fuzz_corpus_test

To fuzz, install the tool once:

cargo install cargo-fuzz --locked

Then, from the repository root:

./scripts/code/fuzz.sh list                       # the target names
./scripts/code/fuzz.sh build                      # build them all
./scripts/code/fuzz.sh replay                     # replay the committed corpus and exit
./scripts/code/fuzz.sh run parse_query            # fuzz until interrupted
./scripts/code/fuzz.sh run parse_query -- -max_total_time=60

replay loads every committed corpus input into the instrumented binaries, runs each once, and generates no new test inputs. It checks the same inputs as the test above, after a much longer build.

What CI does

Every pull request replays the corpus through tests/fuzz_corpus_test.rs, as part of the test suite. It is a regression gate: it re-runs the inputs already known to be interesting and fails if one of them starts crashing again. It does not look for new bugs. The Fuzz targets job runs cargo check --manifest-path fuzz/Cargo.toml --locked so the fuzz crate keeps compiling.

The Nightly workflow runs each target for ten minutes against fresh input, with AddressSanitizer on, as a matrix so one slow target does not consume another’s budget. Crashing inputs are uploaded as build artifacts.

Each target restores the previous coverage corpus before running and minimizes it with cargo fuzz cmin afterward. Separate cache restore/save steps preserve new inputs even when a target crashes.

Every release runs DATUI_FUZZ_SANITIZER=address ./scripts/code/fuzz.sh replay over all committed corpora from the tag. Publishing requires this job to pass, independently of Nightly. Failed replays upload crashing inputs as build artifacts.

The corpus

fuzz/corpus/ is committed, but it is a seed corpus, not the full coverage corpus.

The four targets that take text are seeded with inputs a person can read: parse_query from the parser’s own unit tests and the query examples throughout docs/, sql_group_plan from the planner’s unit tests, config_parse from the TOML blocks in docs/, and format_spec from the specs in the format spec pages, each with a file after it. Anything named regression-* is an input that once crashed a target, kept so the replay test notices if it ever crashes again.

The other three take structured input that arbitrary decodes from raw bytes, so a hand-written seed would mean nothing. Those directories hold a bounded sample of minimized inputs from a real run, capped at 64 files each.

midi_file takes the bytes as they are. Its seeds are small files: format 0, 1 and 2 with running status, sysex and meta events, SMPTE timing, a RIFF MIDI wrapper, and a track cut short.

model_header takes the bytes as they are. Its seeds are small model file headers: SafeTensors with and without __metadata__, GGUF v3 in both byte orders with strings, arrays and tensors of several types, and GGUF v2. gps_parse takes the bytes as they are. Its seeds are a short NMEA log (every sentence type it reads, a prefixed line, a vendor sentence, out-of-range coordinates) and a GPX file (a DOCTYPE, CDATA, entities, namespaced extensions), each behind several first bytes, and a nesting past the depth bound.

vcd_parse, fix_parse and sdf_parse take the bytes as they are, behind a first byte that picks the piece size. Their seeds are a small dump (scopes, an alias, vectors, a real, x and z), a picosecond timescale and a time past i64; FIX messages with SOH, |, ^A and ; delimiters, a prefix, a repeating group, a length-tagged value holding the delimiter, a bad checksum and a message cut short; and SDF records with a blank name, V3000 counts, multi-line values and a missing $$$$. fix_dict takes text: a QuickFIX XML dictionary, a TOML one, and a broken one of each.

audio_header takes the bytes as they are too. Its seeds are tiny audio files: 16-bit PCM, a Broadcast WAV with bext, iXML, cue and LIST chunks, extensible float with a channel mask, RF64 with ds64, 8-bit with a placeholder data size, and AIFF and AIFF-C (sowt, fl32) with a marker.

Commit a regression-* input for each fixed crash. Keep routine coverage inputs in the fuzzing cache. Minimize any additional seeds before committing them:

./scripts/code/fuzz.sh cmin parse_query

When a target fails

libFuzzer writes the offending input to fuzz/artifacts/<target>/. Reproduce it by passing that file instead of a corpus directory, replacing <HASH> with the file’s:

./scripts/code/fuzz.sh run parse_query fuzz/artifacts/parse_query/crash-<HASH>

Fix the bug, then copy the input into fuzz/corpus/<target>/ so the replay test keeps it fixed. A minimal reproducer usually deserves a unit test next to the code as well.

Sanitizer

Local runs omit the sanitizer by default. Enable AddressSanitizer when checking dependency memory errors or running a longer fuzzing session:

DATUI_FUZZ_SANITIZER=address ./scripts/code/fuzz.sh run parse_query

The Nightly and Release workflows run this configuration; pull requests do not.

Sanitizer builds use substantial memory. Nightly and Release limit parallel compilation to two jobs; use the same limit if your build is killed:

CARGO_BUILD_JOBS=2 DATUI_FUZZ_SANITIZER=address ./scripts/code/fuzz.sh run parse_query

Why these run on stable

cargo-fuzz reaches for -Z sanitizer, which is normally nightly-only, and scripts/code/fuzz.sh sets RUSTC_BOOTSTRAP=1 to allow it on stable instead.

Polars currently enables an incompatible internal code path on nightly. The wrapper uses stable with RUSTC_BOOTSTRAP=1 to compile the fuzz targets.

If a future Polars release fixes the nightly path, the flag can be dropped and the scripts switched to cargo +nightly fuzz.