轻量级 PDF 解析器,支持版面、表格、公式和边界框。
Lightweight PDF parser with layout, tables, formulas and bounding boxes

原始链接: https://github.com/beatrizalmeidaf/papero-pdf-text-extractor

**papero** 是一款采用 MIT 许可证、基于几何的文档提取工具,可将 PDF 及其他受支持的文件转换为结构化数据,无需机器学习模型或 GPU。它可在浏览器、Python、CLI 或 REST API 中本地运行,使 PDF 始终保留在用户的设备上。 其基于 PDFium 的引擎能够重建多栏阅读顺序、分离页眉和页脚、检测有边框及无边框表格、将公式转换为近似 LaTeX、裁剪图表,并保留边界框、字体、对齐方式、缩进和文本样式。Apache Tika 提供元数据、带标签的文档结构、OCR,以及对 DOCX、PPTX、XLSX、EPUB 和 HTML 的支持。 输出包括 Markdown、JSON、文本、保留布局信息的 HTML、CSV 表格、包含图像的 ZIP 压缩包、Word、Excel,以及按页面划分的块数据。Python API 提供对页面、块、表格、公式、图表和导出选项的访问;同时还提供 Docker 和交互式 API 文档。 其限制包括:复杂数学公式的线性化、紧密排列的无边框表格可能识别错误、服务器端 OCR,以及 Word/Excel 导出目前仅限于浏览器应用。报告的基准测试显示,在 54 篇测试论文中,每页的中位处理时间为 39 毫秒,且未出现失败。

Hacker News 最新 | past | 评论 | 提问 | 展示 | 职位 | 提交 登录 支持版面、表格、公式和边界框的轻量级 PDF 解析器 (github.com/beatrizalmeidaf) 7 分 由 beatrizalmeidaf 提交 1 小时前 隐藏 | 历史 | 收藏 | 1 条评论 帮助 beatrizalmeidaf 1 小时前 [–] 我构建这个项目,是因为将 PDF 提取为 Markdown/JSON 时,通常会丢失阅读顺序、表格、公式、图表及其原始位置。 目标是提供一个轻量级的文档提取流程,在导出为 Markdown、JSON、Excel 和 Word 的同时,保留文档结构和边界框。 我还在开发用于 RAG 的结构感知语义分块功能,使检索到的文本块能够保留其所在章节、页码以及在 PDF 中的精确视觉位置。 该项目是开源的,希望收到关于架构、提取质量以及实用场景的反馈。 回复 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请加入 YC | 联系方式 搜索:
相关文章

原文

papero

Document structure extraction without the heavyweight stack.

PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.
CPU only. No ML models. Runs in your browser, in Python, or as an API.

CI Python 3.10–3.13 MIT license No ML models

▶ Try it in your browser  ·  Quick start  ·  Benchmarks

papero demo: load a PDF, every block outlined on the page, click a table to inspect it, see tables and LaTeX formulas, export to Word
30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (MP4)


Getting the text out of a PDF is easy. Getting its structure back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.

📖 Reading order
Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside.
▦ Real tables
Ruled, borderless and LaTeX booktabs tables come back as rows and columns — multi-line cells included. Export to CSV or Excel.
∑ Formulas
Superscripts, subscripts and math symbols become LaTeX (E = mc^{2}), plus a cropped image of the formula.
📍 Position of everything
Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it.
🖼 Figures & charts
Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together.
📝 Back to Word
Alignment, indents, line spacing, bold runs and fonts are kept, so a .docx export looks like the original page.

Also: accents drawn as separate glyphs in LaTeX PDFs (Computa¸ca˜o → Computação), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.

from pdf_text_api import extract

doc = extract("paper.pdf")
print(doc.to_markdown())

Or skip the install: open the browser app, drop a PDF, export to the format you need.

More Python — tables, formulas, positions, images, options
from pdf_text_api import extract, extract_text

doc = extract("paper.pdf", images=True)

doc.tables[0].rows  # [["Model", "Accuracy"], ["Base", "0.81"], ...]
doc.formulas[0].latex  # "E = mc^{2}"
doc.figures[0].image.data  # PNG bytes

for block in doc.pages[0].blocks:  # reading order, with positions
    print(block.type, block.bbox, block.text[:60])

doc.to_html()  # keeps alignment and indents
doc.to_dict()  # the full JSON

extract("slides.pptx").to_markdown()  # any format Apache Tika reads
extract_text("contract.pdf").text  # fastest: clean text only
Option Default
pages all "1-3,5,10-"
images False crop figures, tables and formulas to PNG
tables / formulas True detection on/off
ocr "auto" "auto" (scanned pages only), "force", "off"
ocr_language "por+eng" Tesseract languages
tika True False runs the layout engine alone (no Java)
workers 1 processes for long documents
CLI
pdf-text-api extract paper.pdf -o paper.md --images   # Markdown + images/ folder
pdf-text-api extract paper.pdf -o paper.json          # format from the extension
pdf-text-api extract paper.pdf -f csv -o tables.csv   # tables only
pdf-text-api extract paper.pdf -p 1-5 -f html
pdf-text-api extract paper.pdf --fast                 # clean text only
pdf-text-api serve --port 8000                        # API + browser app
REST API & Docker
docker compose up        # API + Apache Tika + Tesseract + browser app on :8000
curl -F "[email protected]" "localhost:8000/v1/extract?format=markdown"
curl -F "[email protected]" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
curl -F "[email protected]" "localhost:8000/v1/extract?per_page=true"     # blocks + positions

One endpoint, POST /v1/extract; interactive docs at /docs.

Parameter Default
mode structured structured (layout + Tika) or fast (text only)
format json json, markdown, text, html, csv, zip
pages all 1-3,5,10-
per_page false include pages, blocks and positions in the JSON
images false crop figures, tables and formulas
ocr auto auto, force, off

Configuration through environment variables — see .env.example.

Every block knows what it is and where it was:

{
  "type": "table",
  "bbox": [56.7, 294.8, 481.9, 374.2],
  "rows": [["Model", "Accuracy"], ["Base", "0.81"]],
  "caption": "Table 1: Comparison between models."
}
Output Python · CLI · API Browser app
Markdown, plain text, JSON ✓ ✓
HTML (keeps alignment and indents) ✓ ✓
CSV of the tables, ZIP with images ✓ ✓
Word .docx that keeps the page's look — ✓
Excel .xlsx, one sheet per table — ✓
Full JSON schema and block types
{
  "schema": "pdf-text-api/document@1",
  "engine": "tika+pdfium",
  "page_count": 12,
  "metadata": { "title": "...", "author": "...", "language": "en" },
  "pages": [{
    "number": 1, "width": 595.3, "height": 841.9,
    "blocks": [{
      "id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
      "text": "Atestamos que a estudante ...",
      "style": { "pt": 11.0, "font": "Arial", "bold": false },
      "format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
      "runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
    }]
  }]
}

Block types: heading (with level), paragraph, list_item (with marker), table (with rows), figure, formula (with latex), caption, code, and — kept apart from the text — header, footer, page_number. Bounding boxes are [x0, y0, x1, y1] in points, origin at the top-left of the page.

Average extraction time per PDF on a log scale: PyMuPDF 97 ms, papero fast 136 ms, papero structured 543 ms, pypdf 1.52 s, pdfplumber 3.56 s, Docling 82.9 s

Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. papero · fast returns clean text; papero · structured also rebuilds reading order, tables, formulas and figures — 0 failures on 54 papers, 39 ms per page (median). Reproduce with benchmarks/.

papero PyMuPDF pdfplumber pypdf Docling Marker
License MIT AGPL MIT BSD MIT GPL
Needs ML models / PyTorch no no no no yes yes
Multi-column reading order ✓ partial — — ✓ ✓
Structured tables ✓ ✓ ✓ — ✓ ✓
Formulas LaTeX from glyphs + image — — — ✓ ✓
Bounding boxes ✓ ✓ ✓ — ✓ ✓
DOCX / PPTX / XLSX / EPUB ✓ partial — — ✓ partial
Runs entirely in the browser ✓ — — — — —

ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX and as an image so nothing is lost.

Two engines run on the same file at the same time:

  • A layout engine on PDFium reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
  • Apache Tika adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.

The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.

Limitations
  • Math: LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
  • Borderless tables with very narrow gaps between columns can read as text.
  • Scanned PDFs need OCR, which runs on the server path (Tesseract is in the Docker image).
  • Word/Excel export is in the browser app for now.
Development
git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
pip install -e ".[dev]"
pytest -q                                   # includes real-world regressions
ruff check src tests && ruff format --check src tests
npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
python -m http.server -d web                # browser app at http://localhost:8000

src/pdf_text_api/ is the Python engine, API and CLI · web/ is the browser app (GitHub Pages) · tests/js/ checks the two engines agree · benchmarks/ downloads the dataset and draws the chart.

Found a PDF papero gets wrong? That's the most useful issue you can open — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.

If papero saves you time, a ⭐ helps other people find it.

Keywords: PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado

MIT © Beatriz Almeida · package and imports keep the name pdf-text-api / pdf_text_api for compatibility.

联系我们 contact @ memedata.com