Fast and efficient PDF text extraction without storage concerns.
Project details
Papero is a powerful PDF text extraction API designed to capture the structure of documents accurately. Extract text into formats like Markdown, JSON, Word, and Excel, all while ensuring files are never stored. Ideal for RAG, search, and LLM applications, it runs seamlessly in the browser or as a Python API, leveraging CPU-only processing.
papero-pdf-text-extractor is a fast and open-source API designed for efficient PDF text extraction. It provides detailed document structure extraction features without requiring heavy machine learning models. This tool operates entirely on CPU, making it accessible for use in various environments including browser applications, Python scripts, and REST APIs. Notably, it ensures that files are not stored, so users maintain complete control over their documents.
The tool reconstructs various elements from PDF documents:
5/12 is output as \frac{5}{12}.To utilize the extraction functionalities:
pip install papero-extract
Then, in Python:
from pdf_text_api import extract
doc = extract("paper.pdf")
print(doc.to_markdown())
Alternatively, for a hands-on experience, users can directly access the application in the browser: Try it now.
Upon extraction, users receive well-structured JSON data with detailed formatting, ensuring that every block of text maintains its identity and location. For instance:
{
"type": "table",
"bbox": [56.7, 294.8, 481.9, 374.2],
"rows": [["Model", "Accuracy"], ["Base", "0.81"]],
"caption": "Table 1: Comparison between models."
}
In benchmark tests with dense academic papers, papero showcased impressive performance metrics, particularly in maintaining reading order and structural accuracy while operating 0 failures on extraction attempts for 54 papers.
The extraction process employs two main engines that operate simultaneously:
Users are encouraged to report any issues regarding PDF extraction by submitting relevant files, which greatly aids in improving the tool's functionality.
Some related topics include: PDF to Markdown, PDF parsing, layout analysis, formula extraction, bounding boxes, and document preprocessing for LLM applications.
Comments
0Start the conversation
Share the first comment.