Reducto Launches r-1 to Tackle AI’s Hardest Document Parsing Problems

Reducto r-1 Brings AI Document Parsing to 1¢

AI systems are getting better at reasoning, but they remain dependent on a less glamorous layer of infrastructure: turning messy documents into reliable machine-readable data. Reducto is targeting that bottleneck with r-1, a new document parsing model designed to process complex PDFs, scans, spreadsheets and other files while preserving the layout and context that conventional extraction pipelines can lose.

For enterprise AI teams, getting information out of a document can be nearly as important as what the AI model does with it afterward. A financial statement is not simply a collection of words; a contract can depend on a strikethrough; and a scanned form may encode meaning through handwriting, checkboxes or the position of individual fields.

Reducto says its new r-1 document parsing model, now available in preview, is designed to address those problems in a single processing layer.

The company claims r-1 reduces parsing errors by as much as 20% compared with its previous agentic parsing pipelines, while lowering the cost of a complete parse to 1 cent per page. Reducto also says the model improves latency for high-volume workloads.

The significance goes beyond another OCR product launch. As enterprises move from demonstrations of generative AI toward production systems, document ingestion has become an increasingly important part of the AI infrastructure stack. Retrieval-augmented generation (RAG), enterprise search and AI agents all depend on information being extracted from source material without losing the relationships that give that information meaning.

Gartner has increasingly emphasized this data-readiness problem. Its research identifies the lack of GenAI-ready data as a major reason AI deployments fail and argues that organizations need stronger capabilities for managing unstructured information.

Reducto’s approach is to make document parsing more unified. According to the company, r-1 combines layout detection, reading-order analysis, table reconstruction, formatting recognition, grounding and granular citations within one model rather than requiring organizations to assemble separate components.

That distinction matters because traditional OCR is primarily concerned with recognizing characters. Modern document AI has a harder job: it must determine what those characters mean in relation to everything around them.

A table, for example, contains relationships between rows and columns. A two-column PDF has an intended reading order. A contract’s formatting can distinguish deleted language from active provisions. A checkbox can represent a binary decision. Losing those relationships can introduce errors downstream even when the underlying words have been recognized correctly.

Reducto says r-1 is designed for precisely these difficult cases, including dense tables, unusual layouts, poor-quality scans, handwriting, watermarked documents and files without predictable templates.

That puts the company in an increasingly competitive part of the enterprise AI stack.

Cloud providers already offer sophisticated document-analysis services. Amazon Textract, for example, can extract text, tables, forms, selection elements and signatures, while its layout capabilities return information about document structure and element positioning. Microsoft’s Azure AI Document Intelligence similarly competes in document extraction and analysis, while large language models from companies such as Google, Microsoft and Amazon can also be incorporated into document-processing pipelines.

The difference is increasingly less about whether a system can perform OCR and more about how much orchestration is required to reach production-grade results.

Reducto’s pitch is that organizations should not need separate tools for OCR, layout analysis, table extraction and downstream cleanup. Its Parse API converts documents into structured JSON and includes typed blocks, page positions and confidence information.

The company is also making price predictability part of the proposition. r-1 costs 1 cent per page, according to Reducto, with the complete parsing process included rather than separate charges for individual capabilities. That could matter for enterprises processing millions of pages, where seemingly small per-page differences can become significant infrastructure costs.

The economics are relevant because document processing is increasingly becoming a prerequisite for AI automation rather than an isolated back-office function. McKinsey estimates generative AI could ultimately create $2.6 trillion to $4.4 trillion in annual economic value across 63 use cases, with the benefits dependent heavily on organizations having usable underlying data.

Reducto is planning to extend the r-1 family rather than treating the new model as a single endpoint. The company says r-1 mini will target workloads where speed and cost take priority, while automatic routing is planned to select the appropriate model for individual pages.

That suggests a familiar direction for AI infrastructure: enterprises may increasingly use multiple specialized models behind a single abstraction layer rather than choosing one model for every workload.

For enterprise teams, the more important question will be whether r-1’s claimed accuracy improvements hold up on their own document collections. Benchmark performance on difficult PDFs is useful, but production environments contain years of inconsistent scans, proprietary templates, handwritten annotations and edge cases that generic evaluations may not capture.

Reducto is offering organizations using other parsers up to $5,000 in credits for testing and side-by-side benchmarking. The model is available in preview through a configuration flag in the company’s Parse API.

The broader trend is clear: as AI moves deeper into enterprise workflows, document parsing is becoming foundational infrastructure. The winners may not simply be the models that understand language best, but the systems that can reliably convert the messy information businesses already possess into structured, traceable data that those models can safely use.

Market Landscape

The document AI market is shifting from conventional OCR toward AI-native document understanding. Earlier generations focused on extracting text from scans; newer systems increasingly need to preserve tables, layout, visual relationships, citations and semantic context for downstream AI applications.

This creates a competitive landscape spanning hyperscalers, enterprise software vendors and specialized AI infrastructure companies.

Amazon Textract already supports text, handwriting, forms, tables, queries, signatures and layout analysis. Microsoft competes through its document-intelligence capabilities, while AI platforms from Google and other hyperscalers can be combined with document-processing workflows.

Reducto is positioning r-1 around a narrower proposition: high-fidelity parsing of the long tail of difficult documents through a unified model and predictable per-page pricing.

For enterprise buyers, the choice therefore is not simply “which OCR is best?” It is whether to build and maintain a multi-stage document pipeline or adopt a specialized parsing layer that can feed RAG systems, AI agents, knowledge bases and enterprise search.

That market should continue expanding as enterprises attempt to make previously inaccessible unstructured information usable by AI. Gartner says organizations increasingly need to structure unstructured content so it can be discovered, retrieved and used beyond the applications where it originally resides.

Top Insights

  • Reducto r-1 combines OCR, layout, tables and citations into one parsing model, targeting enterprise AI teams processing complex unstructured documents at scale.
  • The 1-cent-per-page pricing model could simplify document AI budgeting for companies replacing multi-tool pipelines with unified parsing infrastructure.
  • RAG and AI-agent deployments increasingly depend on reliable document ingestion, making parsing accuracy a foundational concern rather than a peripheral preprocessing task.
  • Hyperscalers including Amazon and Microsoft already provide document intelligence, raising the bar for specialized vendors competing on accuracy, latency and economics.
  • Reducto’s planned r-1 mini and automatic routing point toward model-selection infrastructure that balances accuracy, speed and cost across enterprise workloads.

Power Tomorrow’s Intelligence — Build It with TechEdgeAI

Grow Your
Brand Visibility

Looking to publish a press release, guest article, interview or podcast? Connect with us.

GET FEATURED
Subscribe

Sign up today for exclusive insights and updates.

Newsletter Signup