Document structure extraction without the heavyweight stack.
PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.
CPU only. No ML models. Runs in your browser, in Python, or as an API.
▶ Try it in your browser · Quick start · Benchmarks
30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (MP4)
Getting the text out of a PDF is easy. Getting its structure back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.
|
📖 Reading order Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside. |
▦ Real tables Ruled, borderless and LaTeX booktabs tables come back as rows and columns — multi-line cells included. Export to CSV or Excel. |
∑ Formulas Superscripts, subscripts and math symbols become LaTeX ( E = mc^{2}), plus a cropped image of the formula.
|
|
📍 Position of everything Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it. |
🖼 Figures & charts Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together. |
📝 Back to Word Alignment, indents, line spacing, bold runs and fonts are kept, so a .docx export looks like the original page.
|
Also: accents drawn as separate glyphs in LaTeX PDFs (Computa¸ca˜o → Computação), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.
from pdf_text_api import extract
doc = extract("paper.pdf")
print(doc.to_markdown())Or skip the install: open the browser app, drop a PDF, export to the format you need.
More Python — tables, formulas, positions, images, options
from pdf_text_api import extract, extract_text
doc = extract("paper.pdf", images=True)
doc.tables[0].rows # [["Model", "Accuracy"], ["Base", "0.81"], ...]
doc.formulas[0].latex # "E = mc^{2}"
doc.figures[0].image.data # PNG bytes
for block in doc.pages[0].blocks: # reading order, with positions
print(block.type, block.bbox, block.text[:60])
doc.to_html() # keeps alignment and indents
doc.to_dict() # the full JSON
extract("slides.pptx").to_markdown() # any format Apache Tika reads
extract_text("contract.pdf").text # fastest: clean text only| Option | Default | |
|---|---|---|
pages |
all | "1-3,5,10-" |
images |
False |
crop figures, tables and formulas to PNG |
tables / formulas |
True |
detection on/off |
ocr |
"auto" |
"auto" (scanned pages only), "force", "off" |
ocr_language |
"por+eng" |
Tesseract languages |
tika |
True |
False runs the layout engine alone (no Java) |
workers |
1 |
processes for long documents |
CLI
pdf-text-api extract paper.pdf -o paper.md --images # Markdown + images/ folder
pdf-text-api extract paper.pdf -o paper.json # format from the extension
pdf-text-api extract paper.pdf -f csv -o tables.csv # tables only
pdf-text-api extract paper.pdf -p 1-5 -f html
pdf-text-api extract paper.pdf --fast # clean text only
pdf-text-api serve --port 8000 # API + browser appREST API & Docker
docker compose up # API + Apache Tika + Tesseract + browser app on :8000curl -F "[email protected]" "localhost:8000/v1/extract?format=markdown"
curl -F "[email protected]" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
curl -F "[email protected]" "localhost:8000/v1/extract?per_page=true" # blocks + positionsOne endpoint, POST /v1/extract; interactive docs at /docs.
| Parameter | Default | |
|---|---|---|
mode |
structured |
structured (layout + Tika) or fast (text only) |
format |
json |
json, markdown, text, html, csv, zip |
pages |
all | 1-3,5,10- |
per_page |
false |
include pages, blocks and positions in the JSON |
images |
false |
crop figures, tables and formulas |
ocr |
auto |
auto, force, off |
Configuration through environment variables — see .env.example.
Every block knows what it is and where it was:
{
"type": "table",
"bbox": [56.7, 294.8, 481.9, 374.2],
"rows": [["Model", "Accuracy"], ["Base", "0.81"]],
"caption": "Table 1: Comparison between models."
}| Output | Python · CLI · API | Browser app |
|---|---|---|
| Markdown, plain text, JSON | ✓ | ✓ |
| HTML (keeps alignment and indents) | ✓ | ✓ |
| CSV of the tables, ZIP with images | ✓ | ✓ |
Word .docx that keeps the page's look |
— | ✓ |
Excel .xlsx, one sheet per table |
— | ✓ |
Full JSON schema and block types
{
"schema": "pdf-text-api/document@1",
"engine": "tika+pdfium",
"page_count": 12,
"metadata": { "title": "...", "author": "...", "language": "en" },
"pages": [{
"number": 1, "width": 595.3, "height": 841.9,
"blocks": [{
"id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
"text": "Atestamos que a estudante ...",
"style": { "pt": 11.0, "font": "Arial", "bold": false },
"format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
"runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
}]
}]
}Block types: heading (with level), paragraph, list_item (with marker), table (with rows), figure, formula (with latex), caption, code, and — kept apart from the text — header, footer, page_number. Bounding boxes are [x0, y0, x1, y1] in points, origin at the top-left of the page.
Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. papero · fast returns clean text; papero · structured also rebuilds reading order, tables, formulas and figures — 0 failures on 54 papers, 39 ms per page (median). Reproduce with benchmarks/.
| papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker | |
|---|---|---|---|---|---|---|
| License | MIT | AGPL | MIT | BSD | MIT | GPL |
| Needs ML models / PyTorch | no | no | no | no | yes | yes |
| Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ |
| Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ |
| Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial |
| Runs entirely in the browser | ✓ | — | — | — | — | — |
ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX and as an image so nothing is lost.
Two engines run on the same file at the same time:
- A layout engine on PDFium reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
- Apache Tika adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.
The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.
Limitations
- Math: LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
- Borderless tables with very narrow gaps between columns can read as text.
- Scanned PDFs need OCR, which runs on the server path (Tesseract is in the Docker image).
- Word/Excel export is in the browser app for now.
Development
git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
pip install -e ".[dev]"
pytest -q # includes real-world regressions
ruff check src tests && ruff format --check src tests
npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
python -m http.server -d web # browser app at http://localhost:8000src/pdf_text_api/ is the Python engine, API and CLI · web/ is the browser app (GitHub Pages) · tests/js/ checks the two engines agree · benchmarks/ downloads the dataset and draws the chart.
Found a PDF papero gets wrong? That's the most useful issue you can open — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.
If papero saves you time, a ⭐ helps other people find it.
Keywords: PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado
MIT © Beatriz Almeida · package and imports keep the name pdf-text-api / pdf_text_api for compatibility.