pdf-inspector is a fast Rust library for classifying PDFs and extracting structured text without OCR. It distinguishes text-based, scanned, image-based, and mixed documents while returning confidence scores and pages that may need OCR. Its position-aware parser preserves font data, coordinates, reading order, and multi-column layouts. The converter produces clean Markdown with headings, lists, code blocks, links, page breaks, and formatted tables. It supports CID fonts, several encodings, right-to-left text, and automatic detection of broken font mappings. The same core is available through Rust, Python, Node.js, browser WebAssembly, and command-line interfaces. It is designed for local document pipelines that need fast routing before using slower OCR services.

Features

  • Smart PDF classification with confidence scores
  • Per-page OCR routing recommendations
  • Position-aware text and layout extraction
  • Structured Markdown and table conversion
  • CID font and multiple encoding support
  • Rust, Python, Node.js, WebAssembly, and CLI access

Project Samples

Project Activity

See All Activity >

Categories

Libraries

License

MIT License

Follow pdf-inspector

pdf-inspector Web Site

Other Useful Business Software
MongoDB Atlas runs apps anywhere Icon
MongoDB Atlas runs apps anywhere

Deploy in 115+ regions with the modern database for every enterprise.

MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of pdf-inspector!

Additional Project Details

Programming Language

Rust

Related Categories

Rust Libraries

Registered

2026-08-04