Skip to content

feat: extract TextContent from text, JSON, and XML files - #14

Closed
davebartlett442 wants to merge 1 commit into
docbox-nz:mainfrom
davebartlett442:feature/text-json-xml-content
Closed

davebartlett442 wants to merge 1 commit into
docbox-nz:mainfrom
davebartlett442:feature/text-json-xml-content

Conversation

@davebartlett442

Copy link
Copy Markdown
Contributor

Summary

  • Add a native text processor for any text/* MIME, plus JSON (application/json, +json) and XML (application/xml, +xml)
  • Decode source bytes as UTF-8 (BOM stripped, lossy fallback) and emit TextContent plus a single searchable page
  • Leave LibreOffice conversion first in the dispatch order so text/html, text/spreadsheet, and application/xhtml+xml stay on the existing PDF path

Test plan

  • Unit tests in packages/docbox-processing/tests/test_process_text.rs cover plain text, markdown, csv, JSON, XML, BOM stripping, empty files, and unrelated MIME types
  • CI docbox-processing tests on this PR
  • Upload a .txt, .md, .json, and .xml file and confirm generated TextContent plus search indexing
  • Confirm HTML/office files still convert through LibreOffice as before

Made with Cursor

Decode text/*, application/json, and application/xml as source text so they are searchable without a LibreOffice conversion.

Co-authored-by: Cursor <cursoragent@cursor.com>
@jacobtread

Copy link
Copy Markdown
Member

Have opened another PR expanding on this with some changes. #15 have moved these changes over to there

@jacobtread jacobtread closed this Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants