TeleOCR: Local 1.2B Model for Camera-Captured Document Parsing
Clip title: This Local 1.2B OCR Model Beats GPT-5.2 (and Runs on 8GB GPU) — TeleOCR Author / channel: Prompt Engineer 48 URL: https://www.youtube.com/watch?v=6TnE5pMVbCQ
Summary
The video introduces TeleOCR, a 1.2 billion-parameter document parser developed by China Telecom’s AI research group, designed to accurately extract structured data from various document types, especially “camera-captured” documents. Unlike traditional parsers that struggle with distortions, shadows, or angles inherent in photos, TeleOCR demonstrates impressive capabilities, converting images of receipts, handwritten notes, tables, formulas, and code snippets into clean, structured markdown, HTML tables, or LaTeX. A key highlight is its ability to run efficiently on a local system, specifically an RTX 4060 laptop with just 8GB of VRAM, making it accessible for personal or private offline use, and it is released under the Apache 2.0 license.
TeleOCR’s exceptional performance is underscored by its top rankings across several major document parsing benchmarks, including OmniDocBench v1.6 and Wild-OmniDocBench (specifically for camera photos), often outperforming much larger and more complex models like Gemini 3 Pro. The model addresses a fundamental challenge in document parsing: distinguishing between “digital documents” (flat, straight, clean PDFs/scans) and “camera-captured documents” (bent, folded, tilted, shadowed photos). While most parsers excel at the former, they often fail on the latter. TeleOCR achieves this by adopting a decoupled, two-stage approach that uniquely employs polygons instead of restrictive rectangular bounding boxes for layout segmentation, allowing it to dewarp pages internally and precisely capture the curved shapes of text blocks.
The video elaborates on six core ideas behind TeleOCR’s success. These include Multi-node Consensus Voting for efficient data labeling, Geometry-aware Document Modelling that synthetically bends digital pages to create vast amounts of distorted training data, and Curvature-Guided Douglas-Peucker Sampling for intelligent polygon generation. Other innovations cover Image-to-Image Self-Verification to validate output quality, a Progressive Four-Stage Training pipeline, and Content-Structure Decoupled Learning, which first predicts the structural skeleton of elements like tables (using an OTSL format) and then fills in the content. These sophisticated techniques enable TeleOCR to maintain high accuracy despite challenging real-world photographic conditions.
In conclusion, TeleOCR stands out as a powerful and practical document parsing solution. Its ability to accurately process complex, distorted documents, extract various data types (text, tables, formulas, code, charts), and its local deployability make it highly versatile. It is particularly well-suited for applications such as building Retrieval-Augmented Generation (RAG) systems over scanned documents, digitizing receipts and invoices from phone photos, converting academic papers with formulas into editable Markdown, and enabling private, offline document parsing where sending sensitive files to cloud APIs is undesirable. While it has limitations like occasional character errors on heavily warped small text and current training limited to Chinese and English, its strengths offer significant utility for a wide range of tasks.
Video Description & Links
Description
TeleOCR is a 1.2B-parameter open-source (Apache 2.0) document parser from China Telecom’s AI team. It is #1 on OmniDocBench v1.6 (96.87 overall — above Gemini 3 Pro, GPT-5.2 and Qwen3-VL-235B on that benchmark), #1 on Wild-OmniDocBench (camera photos), and won the ICDAR 2026 Sci-ImageMiner challenge.
In this video I download it, run it on my RTX 4060 laptop (8GB VRAM, ~2.5GB used), explain every idea in the paper (consensus voting, polygon layout, CGDP sampling, self-judgement, OTSL tables), run all official examples plus my own test set (tilted receipt photo, bent table, handwriting, dark code screenshot, Chinese+English, crumpled math page), port the two-stage pipeline to Windows, and wrap it in a Gradio app. Honest failures included.
🔗 Links Model → https://huggingface.co/XingChen-AGI/TeleOCR Code → https://github.com/caipeng328/TeleOCR Paper → https://arxiv.org/abs/2608.12898 GGUF (community) → https://huggingface.co/nandraj/NaviDC-OCR-GGUF
⏱ Chapters 0:00 Intro - a tilted receipt, parsed perfectly 1:10 What is TeleOCR 1:30 The problem: digital vs camera documents 2:48 Architecture (Qwen2.5-VL encoder + Qwen3-0.6B) 3:13 The 6 key ideas 6:21 Benchmarks 7:37 Install + download 7:57 Gotcha: torchvision 9:06 Official examples: text, table, formula, code, chart 10:40 Layout detection 11:22 My own tests 13:06 Honest failure: whole-photo prompt 13:25 My two-stage pipeline on Windows 13:53 Receipt photo: 100% correct 15:07 Gradio app demo 15:45 Speed on an 8GB laptop 16:28 Honest limits 16:52 The good 17:19 Use cases + wrap-up
Get GPUs Runpod: https://get.runpod.io/pe48 Get Hostinger: https://www.hostg.xyz/SHJiu
🔗 Connect with me: Email → prompt.engineer48.alerts@gmail.com GitHub → https://github.com/PromptEngineer48
Tags
ai, TeleOCR, OCR, document parsing, local OCR, OmniDocBench, vision language model, VLM, Qwen2.5-VL, Qwen3, PDF to markdown, receipt OCR, table extraction, LaTeX OCR, RTX 4060, 8GB GPU, open source AI, Hugging Face, Gradio, RAG, MinerU, PaddleOCR-VL
URLs
- https://huggingface.co/XingChen-AGI/TeleOCR
- https://github.com/caipeng328/TeleOCR
- https://arxiv.org/abs/2608.12898
- https://huggingface.co/nandraj/NaviDC-OCR-GGUF
- https://get.runpod.io/pe48
- https://www.hostg.xyz/SHJiu
- https://github.com/PromptEngineer48
Related Concepts
- TeleOCR
- optical character recognition — Wikipedia
- document parsing
- structured data extraction
- camera-captured documents
- markdown — Wikipedia
- HTML tables — Wikipedia
- LaTeX — Wikipedia
- 1.2B parameter model
- local AI model
- RAG systems
Related Entities
- TeleOCR
- China Telecom — Wikipedia
- Prompt Engineer 48
- Gemini 3 Pro — Wikipedia
- GPT-5.2 — Wikipedia
- Apache 2.0 — Wikipedia
- RTX 4060 — Wikipedia