Firecrawl pdf-inspector: Fast PDF Classification and Content Extraction for AI

Clip title: Firecrawl Made PDF Parsing 100x Faster For AI Agents Author / channel: Firecrawl URL: https://www.youtube.com/watch?v=qXYuhmGW524

Summary

This video introduces Firecrawl’s pdf-inspector, an open-source, Rust-powered tool designed for rapid and local processing of PDF documents. Its primary function is to classify PDFs and extract native text, tables, code blocks, and other elements into clean, position-aware Markdown, effectively bypassing the often-slow Optical Character Recognition (OCR) process for documents that don’t require it. The speaker highlights its speed and efficiency, demonstrating it processing a 24-page PDF in 185ms and a benchmark of 200 PDFs in under 3 seconds, significantly outperforming other tools on the opendataloader-bench corpus.

The core advantage of pdf-inspector lies in its intelligent four-step process. First, it parses the PDF by reading its raw bytes and converting it into a navigatable object graph or content stream. Second, and most crucially, it performs classification. Instead of immediately processing every page, it samples a few (typically the first, last, and six evenly spaced pages in between) to determine if the PDF is “TextBased,” “Scanned/Image Based,” or “Mixed.” This classification is based on signals like image content and font decodability, without actually reading the text itself, allowing it to quickly identify documents that contain native, selectable text and avoid unnecessary full-page processing or OCR.

If the document is classified as “TextBased,” pdf-inspector can proceed directly to extracting the content. This involves decoding each glyph (the vector representation of a character) into its corresponding character and recording its precise position on the page. The final step is layout and tables, where the tool reorders the extracted text to determine the intended reading flow, identifies columns, lines, and tables (using a dual-mode approach combining rectangle-based detection and heuristic text alignment), and then converts all this structured information into clean Markdown, ready for consumption by AI agents. This entire process is performed locally on the user’s machine and is notable for not utilizing AI models for its core parsing, contributing to its speed and efficiency.

For scenarios where documents are not PDFs, or if specific pages within a PDF are truly image-based and require OCR, Firecrawl offers its parse endpoint. This API integrates pdf-inspector for native text extraction and intelligently routes image-only pages to a vision model for OCR. Additionally, Firecrawl has released anydoc, another Rust-powered open-source tool that can convert a wide array of document formats, including Word, Excel, PowerPoint, and more, into clean Markdown suitable for Large Language Models (LLMs), further enhancing its document processing ecosystem.

Description

pdf-inspector is an open source PDF parser from Firecrawl that lets coding agents process PDFs without waiting on OCR, classifying any PDF in around 20ms and extracting clean markdown locally.

Written in Rust with no AI models needed, pdf-inspector reads the content stream, samples pages to classify a document as text-based, scanned, or mixed, then handles glyph extraction, layout ordering, and table detection before converting everything to clean markdown. It processes 200 PDFs in 2.8 seconds and ranks among the fastest runs on the open data loader benchmark. In this video I walk through how PDF parsing and PDF to markdown conversion actually work under the hood, run the WASM build in the browser, and show how to wire it into agents like Claude Code through the CLI or one of the language SDKs.

Useful if you’re building RAG pipelines, document extraction workflows, or local PDF processing for AI agents.

pdf-inspector - https://firecrawl.github.io/pdf-inspector/ Firecrawl anydoc - https://github.com/firecrawl/anydoc

👇 Learn more about Firecrawl 👇

📚 Docs: https://docs.firecrawl.dev/features/parse

💻 GitHub: https://github.com/firecrawl/firecrawl

Tags

web scraping, ai agents, ai automation, ai, claude code

URLs