ocr-super-surya
GPU-optimized OCR using Surya. Use when extracting text from images/screenshots, recovering image-only content in PDFs or slides, or inspecting multilingual document layouts. Preserve source locations and report recognition gaps separately from verified text.
By aktsmm · 551 installs
npx skills add aktsmm/agent-skills --skill ocr-super-surya
Source repository · Upstream listing
OCR Super Surya
GPU optimized OCR using [Surya](https://github.com/datalab to/surya).
When to Use
OCR , extract text from image , text recognition , 画像から文字
Extracting text from screenshots, photos, or scanned images
Processing PDFs with embedded images
Multi language document OCR (90+ languages including Japanese)
Extraction and Verification
1. Confirm the input scope and a non public output location for sensitive material; preserve source files and do not bypass document protection. Text redaction does not anonymize retained images or authorize redistribution.
2. Extract native PDF/slide text first, including slide notes where present. A nonempty text layer can still omit screenshots, diagrams, or image only pages; inventory those gaps before selecting OCR targets. For image targets, resolve the highest resolution original before recognizing: chat attachments and thumbnails are downscaled re encodes, so locate the original capture by timestamp, then crop the region of interest and upscale 4 7x (LANCZOS). Dense CJK glyphs collapse at screen scale.
3. Record source relative path, source hash, page/slide number, and image hash. Deduplicate identical images without dropping their source locations; distinguish pending, recognized, no text, and failed outcomes.
4. Run a small sample and compare it with the original image before scaling. Reuse predictors within a batch and checkpoint each completed item so interruption does not require starting over.
5. Keep raw OCR separate from corrected notes. Flag low or missing confidence and inspect multi column order, code symbols, and tables; high confidence is not proof of correctness, and no text is not proof of an empty image. Confirm ambiguous CJK characters line by line on the zoomed crop and mark the unresolved ones; a plausible homoglyph is the default failure mode, not a blank.
6. Report extracted documents, processed OCR targets, unrecognized targets, and manually reviewed scope separately. Verify saved files and hashes; a saved hash alone proves neither source freshness nor semantic accuracy. Recheck the source hash when claiming the same source version.
Quick Start
Installation
Check the selected interpreter and existing dedicated environments before installing or declaring a PDF/OCR dependency unavailable. Set OcrPython to the resolved environment's interpreter; do not assume the active workspace environment has the same packages.
After imports succeed, check package versions and torch.cuda.is available() with that same interpreter. GPU unavailability alone is not a reason to uninstall a working environment. If dependencies are missing, use an isolated environment; stop repeated equivalent TLS failures instead of disabling certificate validation.
Windows + uv 環境(OneDrive配下でのインストール)
OneDrive 配下のフォルダでは uv のハードリンクが失敗するため、以下の手順を使う:
Usage
Python API
API変更履歴 (v0.17.x) :
RecognitionPredictor(foundation predictor) FoundationPredictor が必須引数に変更
call () から langs 引数が削除(自動検出に変更)
GPU Configuration
Variable Default Description
RECOGNITION BATCH SIZE 512 Reduce for lower VRAM
DETECTOR BATCH SIZE 36 Reduce if OOM
Scripts
Script Description
scripts/ocr helper.py Helper with OOM auto retry, batch support
Troubleshooting
エラー 原因 対処
RecognitionPredictor. init () missing 1 required positional argument: 'foundation predictor' v0.13+ でAPIが変更 found pred = FoundationPredictor() を作成して引数に渡す
TypeError: call () got an unexpected keyword argument 'langs' v0.17.x で langs 引数廃止 langs 引数を削除する
AttributeError: 'SuryaDecoderConfig' object has no attribute 'pad token id' transformers 5.x との非互換 pip install "transformers<5.0" でダウングレード
failed to hardlink file ... OneDrive (uv, os error 396) OneDrive のハードリンク制限 link mode=copy を付けてインストール+ UV CACHE DIR をOneDrive外に設定
UnicodeEncodeError: 'cp932' codec can't encode character Windows のCP932デフォルトエンコード sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf 8') を先頭に追加
PdfDocument does not support the context manager protocol Installed pypdfium2 API differs Use explicit try/finally cleanup and close bitmap, page, and document after copying the rendered image; verify against the installed version.
License Note
Surya : GPL 3.0 (code), commercial license required for $2M revenue