Repository navigation
feat: extract TextContent from text, JSON, and XML files (expanded) - #15
Conversation
Decode text/*, application/json, and application/xml as source text so they are searchable without a LibreOffice conversion.
and add option to enable application file processing
54946ce to
5bd4935
Compare
|
The main thing I'd be interested in is text/plain file handling, as a first class supported docbox mime type, with indexing and rendering as plain text. Also interested in markdown files. |
|
And thumbnails etc for plain text files, using a monospace font. This would cover all text file types. |
|
E.g. user can upload notes.md or notes.txt and docbox knows exactly how to handle it. |
Not sure that its very useful to be generating thumbnails for plain text files, seems like it would be a waste of compute & space for something that could be done on the client side (with more customization) as theres lots to consider like text wrapping, font size, ..etc text is small so it can be streamed in pretty easily and just displayed how the client wants it rather than baking in an unchanging image The smaller thumbnails will just be a white background with a bunch of garbled black lines due to the scale, and using the large thumbnail would make less sense than just loading the text content and displaying it properly for nice formatting, colors, wrapping etc Same goes for markdown where you're going to want to handle your own styling and formatting for the markdown and the a server rendered thumbnail itself isn't going to provide much value (especially since it would involve building a markdown and text rendering engine for the server) |
|
We can develop rendering in react for text. But would like to be able to at the very least store and retrieve text files in docbox, with the same text extraction method as for pdf etc. |
Expanded changes for #14
@davebartlett442 due to the nature of .json and .xml files I've put these behind a optional configuration option (and or environment variable) as not everyone will want to index these files especially since they are usually going to be machine readable content and will just pollute the search index so I have separated them from the text file processing you've added.
If you're just looking to render .json and .xml in the UI it would be better to just modify the display logic to handle xml and json by getting the file raw content (rather than generated text content, since its already text) and displaying that rather than indexing the machine readable content