This is a vendor post, and it sells two products: Azure Content Understanding and Azure Document Intelligence, both now part of Foundry Tools. It is still worth reading, because the technical argument holds up and the roadmap it describes affects anyone building document pipelines on Azure. The part that needs the most care is the line between what is production-ready and what is still preview.
The core argument: better models don't fix messy inputs
Microsoft's thesis is that enterprise AI depends less on which model you pick and more on whether agents can trust the content they're given. The post puts it this way: better models raise the ceiling but "do not remove the floor of what has to be true."
The post describes a pattern many IT teams will know. A team puts a chat interface on top of a document repository and the early results look good. Then the collection grows from hundreds of pages to millions, and new requirements show up:
- Predictable costs at volume
- Consistent latency
- Reliable handling of complex layouts, tables and scans
- Answers traceable to a source, which is what the audit team asks for first
The team's point is that the usual fix, "we need a smarter model," misses the problem. Extraction is the layer most production generative AI systems still get wrong.
What "just use the LLM" actually costs
The most useful part of the post lists the work that piles up when a team builds extraction directly on an LLM. The first prompt is easy. Then come:
- Format sprawl. PDFs, TIFFs, DOCX files and scans all parse differently.
- Context limits. Page counts and context windows force you to build a chunker and decide how to handle pages full of images without running up token costs.
- Tables. When the model loses table structure, you add a layout parser.
- Relationships across pages. A total on page 12 that refers to a line item on page 4 needs another pass.
- Audit needs. Per-field confidence, grounding (showing where each value came from), and normalization of dates, currencies and party names.
- Operations. Error handling, token optimization, evaluation, scaling, security and compliance, all of which have to be benchmarked again every time a new model ships.
The post doesn't claim that building your own is always wrong. It says building makes sense when the problem is narrow and the document formats are stable. Microsoft's own tool-selection guidance on Microsoft Learn says the same: building directly on Foundry models can be appropriate when a team needs complete control over model choice, prompts or its own infrastructure. The same guidance adds a practical warning. Without confidence scores, a custom pipeline leaves you either accepting every result or reviewing every result.
Summary: The trade-off is real. Microsoft's list is a good checklist whichever way you go. It is not proof that a managed service will cost less for your particular documents.
Two services, two jobs
The post traces a long history. Azure Form Recognizer became Azure Document Intelligence. Its patented Custom Template technology used random forests to learn repeatable form layouts from a few labeled examples. Custom Neural models followed, drawing on multimodal research such as Microsoft Research Asia's LayoutXLM, which models text, layout and visual information together. The Microsoft Research paper on LayoutXLM describes a multilingual benchmark covering seven languages. That explains why document extraction depends on where things sit on the page, not just on the transcribed text. It does not show which model runs inside any current commercial analyzer.
Microsoft now positions the two services like this:
| Azure Document Intelligence | Azure Content Understanding | |
|---|---|---|
| Approach | Models trained specifically for document tasks | Content extraction plus generative AI (Foundry models) |
| Sweet spot | Known, structured types: invoices, receipts, IDs, tax forms | Variable, unstructured or multimodal content; fields that must be inferred |
| Inputs | Documents | Documents, images, audio, video |
| Setup | Prebuilt or labeled custom models | Schema-based; can start without labeled examples |
Microsoft's architecture guidance says plainly that for standard document types with an existing Document Intelligence prebuilt model, Document Intelligence is more cost-effective and deterministic. Its decision guide also tells existing Document Intelligence customers that their APIs, endpoints, SDKs and billing are unchanged and no migration is required. The Foundry post itself admits that Document Intelligence "continues to excel" on highly structured documents. So the generative service is not automatically the upgrade.
The roadmap, and what's actually shippable
The post lists five areas for coming investment: Advanced Contextualization for prebuilt analyzers, an agentic mode, synchronous Read and Layout APIs, integrations (Foundry IQ, Microsoft Agent Framework, LangChain, MarkItDown and a Content Understanding CLI), and governance features such as grounding, confidence scores and human-in-the-loop review.
Before you plan around that list, check the version boundaries. Microsoft Learn says to use API version 2025-11-01 (GA) for production and 2026-06-01-preview to evaluate newer capabilities. The preview comes without a service-level agreement and isn't recommended for production workloads.
Here is where the headline features stand:
- Agentic mode (preview). According to Microsoft's release notes, Content Understanding now supports agentic mode for document analyzers, available with API version 2026-06-01-preview. Set config.workflow to "agentic" to enable agentic mode. It is meant for cases where the result isn't a single extractable value, such as multistep calculations, validation against a set of conditions, and analysis that requires visual inspection of tables or figures. The initial preview supports one input file per analysis request. Microsoft's analyzer reference also warns that agentic workflows are billed at the advanced contextualization rate and can use more tokens and take longer.
- Synchronous operations (preview). A Microsoft Tech Community announcement says synchronous Read, Layout, and Digital Parse operations in Azure Content Understanding, part of Foundry Tools, deliver real-time content extraction without temporary service-side storage. It also cautions that the features, limits, pricing, API versions, and SDK details described in this post apply to a public preview and may change before general availability. The release notes say these operations handle small documents and images in memory.
- SDKs. The Content Understanding client libraries for Python, .NET, Java, and JavaScript/TypeScript are now available for the 2026-06-01-preview API version. These SDKs support the latest preview features, including agentic mode, synchronous Read and Layout operations, and the newest analyzer improvements. SDKs targeting the 2025-11-01 GA API were released earlier.
- Integrations. The release notes say new articles explain how to use Content Understanding with the Microsoft Agent Framework, LangChain, Logic Apps, and MarkItDown. The September 2026 notes add the CU CLI, which is installable from PyPI and supports both the GA and preview API versions.
For context, Microsoft's August 2026 update post calls this preview wave "CU 2.0." It says the release adds synchronous Read and Layout APIs, advanced contextualization, semantic chunking, improved classification, new prebuilt analyzers, and agentic document reasoning. The same post recommends that customers evaluate accuracy, latency, and cost on representative data before deciding which model deployment and API version best fit their production needs.
Summary: Agentic mode and synchronous Read/Layout are real and you can try them now, but they are preview. Test them; don't put them into production yet.
Grounding and confidence: useful, but not a guarantee
The post leans on per-field grounding and confidence scores as the answer to the auditor's question. Here's what the documentation actually supports:
- In document analyzers, the
estimateFieldSourceAndConfidencesetting returns source locations (page numbers and bounding boxes) and confidence scores from 0 to 1. - Microsoft's transparency note says grounding and confidence are currently available only for the document modality, with expansion planned. Don't assume audio, image and video outputs come with the same audit trail.
- Microsoft's best-practices page gives example thresholds (0.90 for critical fields, 0.80 for important ones, 0.70 for non-critical ones), then says thresholds should be set experimentally for each use case.
In practice: build a labeled test set, measure real error rates per field, and send low-confidence or high-stakes results to a human. A confidence score is an estimate, not a signed guarantee.
The bill is two bills
"Managed" doesn't mean flat-rate. According to Microsoft's pricing explainer, when AI-powered features call LLMs you pay contextualization charges: prepares context, generates confidence scores, provides source grounding, and formats output. On top of that come token-based costs from Microsoft Foundry model deployments. That is in addition to per-page or per-minute content extraction. Because you bring your own model deployments, model choice and deployment type (pay-as-you-go or provisioned throughput) directly affect your costs. The Foundry post names predictable cost as a production requirement but gives no benchmark or pricing comparison to back up its quality and cost claims.
Eleven scenarios, handle with care
The post ends with eleven "less obvious" scenarios: layout-aware contract translation, multimodal search, extracting data from charts and tables, audit-grade extraction for claims and healthcare intake, security monitoring of video clips and images, insurance claim triage, mortgage underwriting, manufacturing incident analysis, vendor onboarding, contract review, and call-center analytics.
They're reasonable ideas. They are also ideas, not deployments with published results. Several involve regulated decisions, and Microsoft's own transparency note says to avoid using the service where a human in the loop or a secondary verification method isn't available. For underwriting or healthcare intake, keep that guidance in front of any rollout.
The takeaway for Azure shops
Microsoft's case for a dedicated extraction layer is sound, even if the post reads partly like a sales pitch. The practical steps:
- Structured forms with a prebuilt model? Start with Document Intelligence.
- Variable, unstructured or mixed-media content? Evaluate Content Understanding on the GA 2025-11-01 API.
- Curious about agentic mode or sync APIs? Test them against 2026-06-01-preview in a non-production environment.
- Whatever you choose, measure accuracy, latency and total cost on your own documents, and calibrate confidence thresholds yourself.
Microsoft says the next post in the series will compare both services with calling LLMs directly. If it includes real numbers, that post will be the more useful one.
References
- Why content extraction still matters in the GenAI era Microsoft Foundry Blog · 2026-09-29T21:30:42+00:00
- From Sync APIs to support for the GPT-5 model series and agentic workflows: What's new in Azure Content Understanding - August 2026 | Microsoft Foundry Blog devblogs.microsoft.com
- What is Azure Content Understanding in Foundry Tools? - Foundry Tools | Microsoft Learn learn.microsoft.com