Repository navigation
Conversation
This adds an opt-in `strict` option to both `csv_core::ReaderBuilder` and `csv::ReaderBuilder`. The parse itself does not change; strict mode only reports fields and records whose quoting is invalid: * a quote in a field that does not start with a quote, * anything other than a delimiter, a record terminator or the end of the data after the closing quote of a quoted field, * the data ending inside a quoted field. In csv-core, such a field is reported as `ReadFieldResult::Invalid` and such a record as `ReadRecordResult::Invalid`, in place of `Field` and `Record`. The DFA gets a per-transition "invalid" table computed from the NFA, so the check costs one extra table lookup and no new states. In csv, reading such a record returns an `ErrorKind::InvalidQuoting` error with the record's position. The record is consumed, so reading continues with the next one. It is never used as the header row and is not compared by the `flexible` length check. This also fixes the NFA `read_record` path for records split across calls: it now counts output per call and offsets field ends by the output already written, like the DFA does. Closes BurntSushi#77
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds the strict parsing mode from #77.
Strict mode is opt-in through
strict(bool)on bothcsv_core::ReaderBuilderandcsv::ReaderBuilder. It does not change the parse. It only reports fields and records with invalid quoting:a"b, or"a"with a leading space),"a"b,"a" ,); doubled quotes and escaped quotes are not closing quotes,Comment lines are not checked, and strict mode has no effect when quoting is disabled.
csv-core. A bad field is returned as
ReadFieldResult::Invalid { record_end }instead ofField { record_end }, and a record with a bad field asReadRecordResult::Invalidinstead ofRecord. Everything else a call returns (bytes read and written, output, field ends) is the same as without strict mode. As suggested in the issue, there are no new DFA states.build_dfacomputes, per transition, whether strict mode rejects it (from the NFA transition), and stores that in a table next tohas_output. The end of the data inside a quoted field is checked in the final transition. Both the DFA and the NFA are covered, andresetclears a pending invalid field.The two result enums are exhaustive, so the new variants are a breaking change for csv-core callers that match on them. I updated the matches in the docs, README, example and benchmarks.
csv. Reading a bad record returns a new
ErrorKind::InvalidQuoting { pos }, whereposis the record's position. The record is consumed, so the next read continues with the following record. A bad record still counts in later record numbers, but it never becomes the header row (the next good row does), and it is never compared by theflexible(false)length check. The deserialize iterators report a bad header row as an error item before falling back to the next row as headers, instead of silently deserializing without headers.While testing records split across calls with the NFA, I found that
read_record_nfausedoutput_posas an index into the caller's new output slice and returned a cumulative count. It now counts output per call and offsets field ends, like the DFA path does.Tests are in
csv-core/tests/strict.rsandtests/strict.rs. Besides the hand-written cases, they generate documents with known bad fields across delimiters, quotes, escape and double-quote settings, terminators, comments, chunk sizes and small output buffers. They check the strict results against the non-strict reader call by call, for both the DFA and the NFA.Closes #77