Skip to content

Add a strict parsing mode that reports invalid quoting - #436

Open
TutTrue wants to merge 1 commit into
BurntSushi:masterfrom
TutTrue:strict-quoting
Open

TutTrue wants to merge 1 commit into
BurntSushi:masterfrom
TutTrue:strict-quoting

Conversation

@TutTrue

@TutTrue TutTrue commented Oct 1, 2026

Copy link
Copy Markdown

This adds the strict parsing mode from #77.

Strict mode is opt-in through strict(bool) on both csv_core::ReaderBuilder and csv::ReaderBuilder. It does not change the parse. It only reports fields and records with invalid quoting:

  • a quote inside a field that does not start with a quote (a"b, or "a" with a leading space),
  • anything other than a delimiter, a record terminator or the end of the data after a closing quote ("a"b, "a" ,); doubled quotes and escaped quotes are not closing quotes,
  • the data ending inside a quoted field.

Comment lines are not checked, and strict mode has no effect when quoting is disabled.

csv-core. A bad field is returned as ReadFieldResult::Invalid { record_end } instead of Field { record_end }, and a record with a bad field as ReadRecordResult::Invalid instead of Record. Everything else a call returns (bytes read and written, output, field ends) is the same as without strict mode. As suggested in the issue, there are no new DFA states. build_dfa computes, per transition, whether strict mode rejects it (from the NFA transition), and stores that in a table next to has_output. The end of the data inside a quoted field is checked in the final transition. Both the DFA and the NFA are covered, and reset clears a pending invalid field.

The two result enums are exhaustive, so the new variants are a breaking change for csv-core callers that match on them. I updated the matches in the docs, README, example and benchmarks.

csv. Reading a bad record returns a new ErrorKind::InvalidQuoting { pos }, where pos is the record's position. The record is consumed, so the next read continues with the following record. A bad record still counts in later record numbers, but it never becomes the header row (the next good row does), and it is never compared by the flexible(false) length check. The deserialize iterators report a bad header row as an error item before falling back to the next row as headers, instead of silently deserializing without headers.

While testing records split across calls with the NFA, I found that read_record_nfa used output_pos as an index into the caller's new output slice and returned a cumulative count. It now counts output per call and offsets field ends, like the DFA path does.

Tests are in csv-core/tests/strict.rs and tests/strict.rs. Besides the hand-written cases, they generate documents with known bad fields across delimiters, quotes, escape and double-quote settings, terminators, comments, chunk sizes and small output buffers. They check the strict results against the non-strict reader call by call, for both the DFA and the NFA.

Closes #77

This adds an opt-in `strict` option to both `csv_core::ReaderBuilder` and
`csv::ReaderBuilder`. The parse itself does not change; strict mode only
reports fields and records whose quoting is invalid:

* a quote in a field that does not start with a quote,
* anything other than a delimiter, a record terminator or the end of the
  data after the closing quote of a quoted field,
* the data ending inside a quoted field.

In csv-core, such a field is reported as `ReadFieldResult::Invalid` and such
a record as `ReadRecordResult::Invalid`, in place of `Field` and `Record`.
The DFA gets a per-transition "invalid" table computed from the NFA, so the
check costs one extra table lookup and no new states.

In csv, reading such a record returns an `ErrorKind::InvalidQuoting` error
with the record's position. The record is consumed, so reading continues
with the next one. It is never used as the header row and is not compared
by the `flexible` length check.

This also fixes the NFA `read_record` path for records split across calls:
it now counts output per call and offsets field ends by the output already
written, like the DFA does.

Closes BurntSushi#77
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

add a "strict" parsing mode

1 participant