paddleocr-doc-parsing
Use this skill to extract structured Markdown/JSON from PDFs and document images—tables with cell-level precision, formulas as LaTeX, figures, seals, charts, headers/footers, multi-column layout and correct reading order. Trigger terms: 文档解析, 版面分析, 版面还原, 表格提取, 公式识别, 多栏排版, 扫描件结构化, 发票, 财报, 复杂 PDF, PDF
By aidenwu0209 · 383 installs
npx skills add aidenwu0209/paddleocr-skills --skill paddleocr-doc-parsing
Source repository · Upstream listing
PaddleOCR Document Parsing Skill
When to Use This Skill
Trigger keywords (routing) : Bilingual trigger terms (Chinese and English) are listed in the YAML description above—use that field for discovery and routing.
Use this skill for :
Documents with tables (invoices, financial reports, spreadsheets)
Documents with mathematical formulas (academic papers, scientific documents)
Documents with charts and diagrams
Multi column layouts (newspapers, magazines, brochures)
Complex document structures requiring layout analysis
Any document requiring structured understanding
Do not use for :
Simple text only extraction
Quick OCR tasks where speed is critical
Screenshots or simple images with clear text
Installation
Scripts declare their dependencies inline ([PEP 723](https://peps.python.org/pep 0723/)). No separate install step is needed — [uv](https://docs.astral.sh/uv/) resolves dependencies automatically:
How to Use This Skill
Working directory : All uv run scripts/... commands below should be run from this skill's root directory (the directory containing this SKILL.md file).
Basic Workflow
1. Identify the input source :
User provides URL: Use the file url parameter
User provides local file path: Use the file path parameter
2. Execute document parsing :
Or for local files:
Optional: explicitly set file type :
file type 0 : PDF
file type 1 : image
If omitted, the type is auto detected from the file extension. For local files, a recognized extension ( .pdf , .png , .jpg , .jpeg , .bmp , .tiff , .tif , .webp ) is required; otherwise pass file type explicitly. For URLs with unrecognized extensions, the service attempts inference.
Performance note : Parsing time scales with document complexity. Single page images typically complete in 1 5 seconds; large PDFs (50+ pages) may take several minutes. Allow adequate time before assuming a timeout.
Default behavior: save raw JSON to a temp file :
If output is omitted, the script saves automatically under the system temp directory
Default path pattern: <system temp /paddleocr/doc parsing/results/result <timestamp <id .json
If output is provided, it overrides the default temp file destination
If stdout is provided, JSON is printed to stdout and no file is saved
In save mode, the script prints the absolute saved path on stderr: Result saved to: /absolute/path/...
In default/custom save mode, read and parse the saved JSON file before responding
Use stdout only when you explicitly want to skip file persistence
3. Parse JSON response :
Check the ok field: true means success, false means error
The output contains complete document data: text, tables, formulas (LaTeX), figures, seals, headers/footers, and reading order
Use the appropriate field based on what the user needs:
text — full document text across all pages
result.result.layoutParsingResults[n].markdown.text — page level markdown
result.result.layoutParsingResults[n].prunedResult — structured layout data with positions and confidence
Handle errors: If ok is false, display error.message
4. Present results to user :
Display content based on what the user requested (see "Complete Output Display" below)
If the content is empty, the document may contain no extractable text
In save mode, always tell the user the saved file path and that full raw JSON is available there
What to Do After Parsing
Common next steps once you have the structured output:
Save as Markdown : Write the text field to a .md file — tables, headings, and formulas are preserved
Extract specific tables : Navigate result.result.layoutParsingResults[n].prunedResult to access individual layout elements with position and confidence data
Feed to RAG / search pipeline : The text field is structured markdown, ready for chunking and indexing
Poor results : See "Tips for Better Results" below before retrying
Complete Output Display
Display the COMPLETE extracted content based on what the user asked for. The parsed output is only useful if the user receives all of it — truncation silently drops data.
If user asks for "all text", show the entire text field
If user asks for "tables", show ALL tables in the document
If user asks for "main content", filter out headers/footers but show ALL body text
Do not truncate with "..." unless content is excessively long ( 10,000 chars)
Do not say "Here's a preview" when user expects complete output
Example Correct :
Example Incorrect :
Understanding the Output
The script returns an envelope with ok , text , result , and error . Use text for the full document content; navigate result.result.layoutParsingResults[n] for per page structured data.
For the complete schema and field level details, see references/output schema.md .
Raw result location (default): the temp file path printed by the script on stderr
Alternative: paddleocr CLI
This mirror keeps the bundled scripts/layout caller.py as the default path. The upstream PaddleOCR project (since [PR 18090](https://github.com/PaddlePaddle/PaddleOCR/pull/18090), 2026 06 03) also ships an official CLI that calls the same API directly. If the paddleocr package is installed, you can use it as a drop in alternative — no uv run or local scripts required.
Install (one time):
Environment : the CLI only needs PADDLEOCR ACCESS TOKEN . It resolves the API endpoint internally, so PADDLEOCR DOC PARSING API URL is not required when using the CLI (the URL is still required by the script).
Basic document parsing :
Common options :
CLI output format — different from the script envelope :
The CLI prints {jobId, pages:[...]} to stdout. It does not wrap the response in the script's {ok, text, result, error} envelope, does not auto save to a temp file, and does not concatenate a top level text field for you. Field names also differ from the script's raw result.result.layoutParsingResults[n].markdown.text path. If you switch paths, update your parsing logic accordingly.
Scripts vs CLI — at a glance:
Scripts (default) paddleocr CLI (alternative)
Install uv resolves PEP 723 inline deps pip install "paddleocr =3.7.0"
Required env PADDLEOCR DOC PARSING API URL + PADDLEOCR ACCESS TOKEN PADDLEOCR ACCESS TOKEN only
Entry uv run scripts/layout caller.py ... paddleocr api model type doc parsing ...
Page selection Pre split with scripts/split pdf.py Native page ranges "1 5,10"
Image optimization scripts/optimize file.py before upload Same, or rely on CLI defaults
Output {ok, text, result, error} envelope, auto saved to temp file {jobId, pages:[...]} to stdout
Result location Path printed on stderr (or output / stdout ) stdout (or output )
Best for Skills runtimes, offline friendly, no extra install Already have paddleocr installed, want the upstream canonical flow
Run paddleocr api help for the full option list.
Usage Examples
Example 1: Extract Full Document Text
Then use:
Top level text for quick full text output
result.result.layoutParsingResults[n].markdown when page level output is needed
Example 2: Extract Structured Page Data
Then use:
result.result.layoutParsingResults[n].prunedResult for structured parsing data (layout/content/confidence)
Example 3: Print JSON to stdout (without saving to file)
By default the script writes JSON to a temp file and prints the path to stderr. Add stdout to print the full JSON directly to stdout instead. Use this when you need to inspect the result inline or pipe it to another tool.
First Time Configuration
When API is not configured , the script outputs:
Configuration workflow :
1. Show the exact error message to the user.
2. Guide the user to obtain credentials : Visit the [PaddleOCR website](https://www.paddleocr.com), click API , select a model ( PP StructureV3 , PaddleOCR VL , or PaddleOCR VL 1.5 ), then copy the API URL and Token . They map to these environment variables:
PADDLEOCR DOC PARSING API URL — full endpoint URL ending with /layout parsing
PADDLEOCR ACCESS TOKEN — 40 character alphanumeric string
Optionally configure PADDLEOCR DOC PARSING TIMEOUT for request timeout. Recommend using the host application's standard configuration method rather than pasting credentials in chat.
3. Apply credentials — one of:
User configured via the host UI : ask the user to confirm, then retry.
User pastes credentials in chat : warn that they may be stored in conversation history, help the user persist them using the host's standard configuration method, then retry.
Handling Large Files
For PDFs, the maximum is 100 pages per request.
Optimize Large Images Before Parsing
For large image files, compress before uploading — this reduces upload time and can improve processing stability:
quality controls JPEG/WebP lossy compression (1 100, default 85); it has no effect on PNG output. Use target size (in MB, default 20) to set the max file size — the script iteratively downscales until the target is met.
Use URL for Large Local Files (Recommended)
For very large local files, prefer file url over file path to avoid base64 encoding overhead:
Process Specific Pages (PDF Only)
If you only need certain pages from a large PDF, extract them first:
Error Handling
All errors return JSON with ok: false . Show the error message and stop — do not fall back to your own vision capabilities. Identify the issue from error.code and error.message :
Authentication failed (403) — error.message contains "Authentication failed"
Token is invalid, reconfigure with correct credentials
Quota exceeded (429) — error.message contains "API rate limit exceeded"
Daily API quota exhausted, inform user to wait or upgrade
Unsupported format — error.message contains "Unsupported file format"
File format not supported, convert to PDF/PNG/JPG
No content detected :
text field is empty
Document may be blank, image only, or contain no extractable text
Tips for Better Results
If parsing quality is poor:
Large or high resolution images : Compress with optimize file.py before parsing — oversized inputs can degrade layout detection:
Check confidence : result.result.layoutParsingResults[n].prunedResult includes confidence scores per layout element — low values indicate regions worth reviewing
Reference Documentation
references/output schema.md — Full output schema, field descriptions, and command examples
Note : Model version and capabilities are determined by your API endpoint ( PADDLEOCR DOC PARSING API URL ).
Testing the Skill
To verify the skill is working properly:
The first form tests configuration and API connectivity. skip api test checks configuration only. test url overrides the default sample document URL.