23 points beatrizalmeidaf 1 hour ago 4 comments
beatrizalmeidaf 1 hour ago | parent
I built this because extracting PDFs into Markdown/JSON often loses
reading order, tables, formulas, figures, and their original locations.
The goal is a lightweight document extraction pipeline that preserves document structure and bounding boxes while exporting to Markdown, JSON, Excel and Word.
I'm also working on structure-aware semantic chunking for RAG, so retrieved chunks can retain their section, page and exact visual location in the PDF.
The project is open source and I'd love feedback on the architecture, extraction quality, and useful use cases.
thatcherc 26 minutes ago | parent
This looks fantastic! The table and formula extraction features are especially interesting. My immediate question is: can this be integrated into Zotero? Most of the PDFs I read are research papers and extracting tables and formulas directly from my zotero collection would be super handy.
archeantus 11 minutes ago | parent
Great work, thanks for sharing.