Each engine has its own UI — files stay local; weights download on opt-in
Browser OCR
A suite of client-side OCR / document engines. Pick a model below — dedicated routes, honest download sizes, no silent gigabyte fetches. (Still not full Python Docling.)
Ready now
- Granite-Docling — DocTags → HTML/Markdown/crops via WebGPU. Opt-in ~1 GB ONNX download.
- RapidOCR English — PP-OCRv4 mobile EN via ONNX Runtime Web (~14 MB). Boxes + text, no WebGPU.
- RapidOCR Chinese — PP-OCRv4 mobile 中文 pack (~16 MB). Same ORT pipeline as English.
- PaddleOCR.js — Official @paddleocr/paddleocr-js PP-OCRv5 mobile (~22 MB+). ORT WASM + OpenCV.
- Tesseract.js — WASM eng fallback (~23 MB tessdata). No WebGPU. Honest classic baseline.
- TrOCR + detect — TrOCR small-printed (~64 MB) + RapidOCR EN detect→crop. Transformers.js.
- Florence-2 — Florence-2-base-ft OCR task via WebGPU (~320 MB). Not DocTags.
LightOnOCR / GLM-OCR stay blocked (download wall / immature Hub card). Florence, TrOCR, and PaddleOCR.js shipped with honest opt-in gates.
Layout VLM
Plain OCR
Classic OCR
VLM OCR
VLM OCR
TrOCR + detect
TrOCR small-printed (~64 MB) + RapidOCR EN detect→crop. Transformers.js.
VLM OCR
Florence-2
Florence-2-base-ft OCR task via WebGPU (~320 MB). Not DocTags.
VLM OCR · blocked
LightOnOCR
Blocked: ~0.7–0.8 GB q4 overlaps Granite download wall.
VLM OCR · blocked
GLM-OCR
Blocked: Hub card immature (size/docs TBD). Prefer Granite.