Skip to content

Add rich file type support: Office, HTML, EPUB, RTF, EML, ZIP - #1

Merged
drompincen merged 1 commit into
mainfrom
feature/rich-file-types
Mar 22, 2026
Merged

drompincen merged 1 commit into
mainfrom
feature/rich-file-types

Conversation

@drompincen

Copy link
Copy Markdown
Owner

Summary

  • 13 new file formats extracted using 100% Java open-source libraries
  • DOCX/DOC, PPTX/PPT, XLSX/XLS via Apache POI 5.3.0
  • ODT/ODP/ODS via ODF XML ZIP extraction
  • HTML/HTM and EPUB via Jsoup
  • RTF via JDK built-in RTFEditorKit (zero new deps)
  • EML via Jakarta Mail (angus-mail 2.0.3)
  • ZIP recursive extraction with 50MB/500-entry bomb protection
  • Extension defaults updated in MCP server, CLI client, and interactive CLI
  • README: added Supported File Types table

Test plan

  • 65 tests passing, 0 failures (mvn test)
  • 7 new TextExtractorTest methods covering PPTX, DOCX, XLSX, HTML, RTF, ZIP, and isSupportedExtension
  • POI 5.x API fixes: HSLFSlideShow(InputStream), slide iteration replacing removed PowerPointExtractor
  • .html routing fixed: removed from TEXT_EXTENSIONS so Jsoup path (script-stripping) takes precedence

🤖 Generated with Claude Code

New extractors in TextExtractor (13 new formats, all pure Java):
- DOCX/DOC — Apache POI (XWPFWordExtractor, HSLFSlideShow)
- PPTX/PPT — Apache POI (XMLSlideShow, HSLFSlideShow slide iteration)
- XLSX/XLS — Apache POI (XSSFWorkbook/HSSFWorkbook + DataFormatter)
- ODT/ODP/ODS — ODF XML ZIP extraction (content.xml tag stripping)
- HTML/HTM — Jsoup (script/style stripped, body text only)
- EPUB — Jsoup (ZIP+XHTML spine extraction)
- RTF — JDK RTFEditorKit (zero new deps)
- EML — Jakarta Mail (headers + multipart body, HTML parts via Jsoup)
- ZIP — JDK ZipInputStream (recursive text entry extraction, 50MB/500 entry limit)

Dependencies added: poi-ooxml 5.3.0, poi-scratchpad 5.3.0, odfdom-java 0.12.0,
jsoup 1.18.1, angus-mail 2.0.3

Extension defaults updated in MCP server, CLI client, and interactive CLI.
README: added Supported File Types table.

Tests: 7 new TextExtractorTest methods (PPTX/DOCX/XLSX/HTML/RTF/ZIP/isSupportedExtension)
All 65 tests passing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@drompincen
drompincen merged commit 0e39fa7 into main Mar 22, 2026
1 check passed
@drompincen
drompincen deleted the feature/rich-file-types branch March 22, 2026 17:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant