TeleOCR

TeleOCR is a 1.2 billion-parameter document parser developed by China Telecom’s AI research group. It is designed to accurately extract structured data from various document types, with a specific focus on “camera-captured” documents.

Key Features

  • High Accuracy on Distorted Inputs: Specifically optimized to handle distortions, shadows, and perspective angles inherent in photos, addressing common failures in traditional parsers.
  • Efficiency: Runs on consumer-grade hardware, specifically requiring only 8GB of GPU VRAM.
  • Performance: Claims to outperform larger commercial models like GPT-5.2 in specific document parsing tasks.

Technical Details

References