paddleocr-doc-parsing

Advanced document parsing with PaddleOCR. Returns complete document structure including text, tables, formulas, charts, and layout information. The AI agent extracts relevant content based on user needs.

By freestylefly · 360 installs

npx skills add freestylefly/canghe-skills --skill paddleocr-doc-parsing

Source repository · Upstream listing

PaddleOCR Document Parsing Skill When to Use This Skill Use Document Parsing for : Documents with tables (invoices, financial reports, spreadsheets) Documents with mathematical formulas (academic papers, scientific documents) Documents with charts and diagrams Multi column layouts (newspapers, magazines, brochures) Complex document structures requiring layout analysis Any document requiring structured understanding Use Text Recognition instead for : Simple text only extraction Quick OCR tasks where speed is critical Screenshots or simple images with clear text How to Use This Skill ⛔ MANDATORY RESTRICTIONS DO NOT VIOLATE ⛔ 1. ONLY use PaddleOCR Document Parsing API Execute the script python scripts/vl caller.py 2. NEVER parse documents directly Do NOT parse documents yourself 3. NEVER offer alternatives Do NOT suggest "I can try to analyze it" or similar 4. IF API fails Display the error message and STOP immediately 5. NO fallback methods Do NOT attempt document parsing any other way If the script execution fails (API not configured, network error, etc.): Show the error message to the user Do NOT offer to help using your vision capabilities Do NOT ask "Would you like me to try parsing it?" Simply stop and wait for user to fix the configuration Basic Workflow 1. Execute document parsing : Or for local files: Optional: explicitly set file type : file type 0 : PDF file type 1 : image If omitted, the service can infer file type from input. Default behavior: save raw JSON to a temp file : If output is omitted, the script saves automatically under the system temp directory Default path pattern: <system temp /paddleocr/doc parsing/results/result <timestamp <id .json If output is provided, it overrides the default temp file destination If stdout is provided, JSON is printed to stdout and no file is saved In save mode, the script prints the absolute saved path on stderr: Result saved to: /absolute/path/... In default/custom save mode, read and parse the saved JSON file before responding In save mode, always tell the user the saved file path and that full raw JSON is available there Use stdout only when you explicitly want to skip file persistence 2. The output JSON contains COMPLETE content with all document data: Headers, footers, page numbers Main text content Tables with structure Formulas (with LaTeX) Figures and charts Footnotes and references Seals and stamps Layout and reading order Input type note : Supported file types depend on the model and endpoint configuration. Always follow the file type constraints documented by your endpoint API. 3. Extract what the user needs from the output JSON using these fields: Top level text result[n].markdown result[n].prunedResult IMPORTANT: Complete Content Display CRITICAL : You must display the COMPLETE extracted content to the user based on their needs. The output JSON contains ALL document content in a structured format In save mode, the raw provider result can be inspected in the saved JSON file Display the full content requested by the user , do NOT truncate or summarize If user asks for "all text", show the entire text field If user asks for "tables", show ALL tables in the document If user asks for "main content", filter out headers/footers but show ALL body text What this means : DO : Display complete text, all tables, all formulas as requested DO : Present content using these fields: top level text , result[n].markdown , and result[n].prunedResult DON'T : Truncate with "..." unless content is excessively long ( 10,000 chars) DON'T : Summarize or provide excerpts when user asks for full content DON'T : Say "Here's a preview" when user expects complete output Example Correct : Example Incorrect : Understanding the JSON Response The output JSON uses an envelope wrapping the raw API result: Key fields : text — extracted markdown text from all pages (use this for quick text display) result raw provider response object result[n].prunedResult structured parsing output for each page (layout/content/confidence and related metadata) result[n].markdown — full rendered page output in markdown/HTML Raw result location (default): the temp file path printed by the script on stderr Usage Examples Example 1: Extract Full Document Text Then use: Top level text for quick full text output result[n].markdown when page level output is needed Example 2: Extract Structured Page Data Then use: result[n].prunedResult for structured parsing data (layout/content/confidence) result[n].markdown for rendered page content Example 3: Print JSON Without Saving Then return: Full text when user asks for full document content result[n].prunedResult and result[n].markdown when user needs complete structured page data First Time Configuration When API is not configured : The error will show: Configuration workflow : 1. Show the exact error message to the user (including the URL). 2. Guide the user to configure securely : Recommend configuring through the host application's standard method (e.g., settings file, environment variable UI) rather than pasting credentials in chat. List the required environment variables: 3. If the user provides credentials in chat anyway (accept any reasonable format): PADDLEOCR DOC PARSING API URL=https://xxx.paddleocr.com/layout parsing, PADDLEOCR ACCESS TOKEN=abc123... Here's my API: https://xxx and token: abc123 Copy pasted code format Any other reasonable format Security note : Warn the user that credentials shared in chat may be stored in conversation history. Recommend setting them through the host application's configuration instead when possible. 4. Parse and validate the values : Extract PADDLEOCR DOC PARSING API URL (look for URLs with paddleocr.com or similar) Confirm PADDLEOCR DOC PARSING API URL is a full endpoint ending with /layout parsing Extract PADDLEOCR ACCESS TOKEN (long alphanumeric string, usually 40+ chars) Tell the user exactly which environment variables to set 5. Ask the user to confirm the environment is configured : Wait for the user to confirm these values have been set in their host application, runtime environment, or appropriate config file For security reasons, do not run configure.py or create a local .env file by default if the skill is installed under a host application directory (for example, ~/.claude/skills ) 6. Retry only after confirmation : Once the user confirms the environment variables are available, retry the original parsing task IMPORTANT : The error message format is STRICT and must be shown exactly as provided by the script. Do not modify or paraphrase it. Handling Large Files There is no file size limit for the API. For PDFs, the maximum is 100 pages per request. Tips for large files : Use URL for Large Local Files (Recommended) For very large local files, prefer file url over file path to avoid base64 encoding overhead: Process Specific Pages (PDF Only) If you only need certain pages from a large PDF, extract them first: Error Handling Authentication failed (403) : → Token is invalid, reconfigure with correct credentials API quota exceeded (429) : → Daily API quota exhausted, inform user to wait or upgrade Unsupported format : → File format not supported, convert to PDF/PNG/JPG Important Notes The script NEVER filters content It always returns complete data The AI agent decides what to present Based on user's specific request All data is always available Can be re interpreted for different needs No information is lost Complete document structure preserved Reference Documentation references/output schema.md Output format specification Note : Model version and capabilities are determined by your API endpoint ( PADDLEOCR DOC PARSING API URL ). Load these reference documents into context when: Debugging complex parsing issues Need to understand output format Working with provider API details Testing the Skill To verify the skill is working properly: This tests configuration and optionally API connectivity.