[{"data":1,"prerenderedAt":92},["ShallowReactive",2],{"blog-tag-human-in-the-loop":3},[4,24,36,48,59,70,79],{"id":5,"slug":6,"body":7,"html":8,"title":9,"description":10,"category":11,"tags":12,"author":17,"date":18,"year":19,"month":20,"quarter":21,"status":22,"featured":23},"2026\u002F08\u002Findustry-applications\u002Foperations-control-and-disruption-management","operations-control-and-disruption-management","\nIn aviation and logistics, disruption is normal: weather, technical faults, crew limits, port congestion, customs holds, missed connections. What separates a good day from a bad one is how quickly the operation understands the impact, agrees a recovery and executes it.\n\nIn many operations that coordination still happens over phone, radio, chat groups and whiteboards. Decisions are made well, but they aren't recorded well. Downstream teams learn about changes late.\n\n## What the application does\n\nThe **operations control** family in the Atlas provides a shared workflow for disruption:\n\n1. **Detect:** events arrive from operational systems (flight or shipment status, maintenance, crew, weather, partner messages).\n2. **Assess impact:** affected flights, shipments, crews, passengers or customers, and downstream connections.\n3. **Generate options:** recovery options such as swap, delay, cancel, reroute or re-book, with their consequences.\n4. **Decide:** the controller selects an option, with the rationale recorded.\n5. **Execute:** tasks go to the affected teams (ground handling, crew control, customer service, partners), each with an owner.\n6. **Communicate:** updates to customers and partners.\n7. **Log and learn:** an operational log of events, decisions and outcomes, available for post-event review and regulatory records.\n\n## Where AI helps\n\n- **Impact summarization:** “what does this delay break?” answered in seconds.\n- **Recovery option generation:** candidate plans scored against cost, delay minutes, crew legality and customer impact. The controller chooses.\n- **Forecasting:** disruption risk from weather and schedule patterns, so teams prepare early.\n- **Drafting communications:** customer and partner messages for review.\n- **Post-event analysis:** timelines and contributing factors compiled from the log.\n\n## Human authority stays explicit\n\nOperational decisions carry safety, regulatory and commercial consequences. The application frames AI outputs as options, never actions. It records who decided and keeps deterministic rules, such as crew duty limits or dangerous-goods constraints, as hard constraints rather than model suggestions.\n\n## Integrations\n\nOperations and scheduling systems, crew management, maintenance and technical records, passenger service or TMS\u002FWMS, partner messaging (such as airline industry message formats or EDI), weather and airport data, and customer communication platforms.\n\n## Who uses it\n\nOperations controllers and duty managers, crew and maintenance control, ground and hub operations, customer service leads and operations leadership.\n\n## First scope\n\nOne disruption type that recurs weekly, where the recovery decision and downstream tasks are currently coordinated by phone. Measure recovery time, communication lag and log completeness. Scope it in a [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint).\n\nSee [logistics, transport and aviation](\u002Findustries\u002Flogistics-transport-aviation), explore the [Atlas](\u002Fatlas), or [bring us your disruption playbook](\u002Fcontact).\n","\u003Cp>In aviation and logistics, disruption is normal: weather, technical faults, crew limits, port congestion, customs holds, missed connections. What separates a good day from a bad one is how quickly the operation understands the impact, agrees a recovery and executes it.\u003C\u002Fp>\n\u003Cp>In many operations that coordination still happens over phone, radio, chat groups and whiteboards. Decisions are made well, but they aren&#39;t recorded well. Downstream teams learn about changes late.\u003C\u002Fp>\n\u003Ch2>What the application does\u003C\u002Fh2>\n\u003Cp>The \u003Cstrong>operations control\u003C\u002Fstrong> family in the Atlas provides a shared workflow for disruption:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Detect:\u003C\u002Fstrong> events arrive from operational systems (flight or shipment status, maintenance, crew, weather, partner messages).\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Assess impact:\u003C\u002Fstrong> affected flights, shipments, crews, passengers or customers, and downstream connections.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Generate options:\u003C\u002Fstrong> recovery options such as swap, delay, cancel, reroute or re-book, with their consequences.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Decide:\u003C\u002Fstrong> the controller selects an option, with the rationale recorded.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Execute:\u003C\u002Fstrong> tasks go to the affected teams (ground handling, crew control, customer service, partners), each with an owner.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Communicate:\u003C\u002Fstrong> updates to customers and partners.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Log and learn:\u003C\u002Fstrong> an operational log of events, decisions and outcomes, available for post-event review and regulatory records.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2>Where AI helps\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Impact summarization:\u003C\u002Fstrong> “what does this delay break?” answered in seconds.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Recovery option generation:\u003C\u002Fstrong> candidate plans scored against cost, delay minutes, crew legality and customer impact. The controller chooses.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Forecasting:\u003C\u002Fstrong> disruption risk from weather and schedule patterns, so teams prepare early.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Drafting communications:\u003C\u002Fstrong> customer and partner messages for review.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Post-event analysis:\u003C\u002Fstrong> timelines and contributing factors compiled from the log.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Human authority stays explicit\u003C\u002Fh2>\n\u003Cp>Operational decisions carry safety, regulatory and commercial consequences. The application frames AI outputs as options, never actions. It records who decided and keeps deterministic rules, such as crew duty limits or dangerous-goods constraints, as hard constraints rather than model suggestions.\u003C\u002Fp>\n\u003Ch2>Integrations\u003C\u002Fh2>\n\u003Cp>Operations and scheduling systems, crew management, maintenance and technical records, passenger service or TMS\u002FWMS, partner messaging (such as airline industry message formats or EDI), weather and airport data, and customer communication platforms.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cp>Operations controllers and duty managers, crew and maintenance control, ground and hub operations, customer service leads and operations leadership.\u003C\u002Fp>\n\u003Ch2>First scope\u003C\u002Fh2>\n\u003Cp>One disruption type that recurs weekly, where the recovery decision and downstream tasks are currently coordinated by phone. Measure recovery time, communication lag and log completeness. Scope it in a \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>See \u003Ca href=\"\u002Findustries\u002Flogistics-transport-aviation\">logistics, transport and aviation\u003C\u002Fa>, explore the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your disruption playbook\u003C\u002Fa>.\u003C\u002Fp>\n","Operations control in aviation and logistics: managing disruption as a workflow","Operations-control applications that turn disruption handling into a shared, auditable workflow with AI-assisted recovery options and human decisions.","industry-applications",[13,14,15,16],"logistics-aviation","operations","human-in-the-loop","agents","fazezero-editorial","2026-08-20T00:00:00.000Z",2026,8,3,"published",false,{"id":25,"slug":26,"body":27,"html":28,"title":29,"description":30,"category":11,"tags":31,"author":17,"date":35,"year":19,"month":20,"quarter":21,"status":22,"featured":23},"2026\u002F08\u002Findustry-applications\u002Freferrals-and-care-coordination","referrals-and-care-coordination","\nClinicians spend a significant part of their day on work that isn't clinical: referral letters, pre-authorization requests, follow-up coordination, chasing results and scheduling across providers. Patients experience that work as waiting.\n\nThe **care coordination** family in the Atlas focuses on these administrative and coordination workflows. It doesn't touch clinical decision-making, and it's designed so it cannot drift into it.\n\n## Workflows covered\n\n- **Referral intake:** referrals arrive from primary care, other hospitals or payers, and are checked for completeness.\n- **Triage and routing:** referrals go to the right service and are prioritized according to clinical rules defined by the provider.\n- **Pre-authorization:** requests are assembled with the required documentation, submitted to payers and tracked.\n- **Scheduling coordination:** appointments are linked across departments and providers.\n- **Care pathway tasks:** follow-ups, results, patient communication and hand-offs, each with an owner and due date.\n- **Closure and feedback:** outcomes communicated back to the referring provider.\n- **Reporting:** waiting times, bottlenecks and service-level performance.\n\n## Where AI helps\n\n- **Document extraction:** pull structured data from referral letters and attachments.\n- **Completeness checks:** identify missing information before a referral reaches a coordinator.\n- **Summaries:** a concise case summary for coordinators, drawn from the documents.\n- **Drafting:** pre-authorization justifications and patient communications, for staff to review.\n- **Queue prioritization:** suggestions based on the provider's own rules, never the model's opinion of clinical urgency.\n\n## Where it must not\n\nAI output in this family never replaces clinical judgement. Clinical triage rules are configured by the provider and applied deterministically, and any AI suggestion that touches clinical content is shown to a qualified person before it has effect. Each AI output is labelled and its acceptance recorded.\n\n## Privacy and hosting\n\nHealth data demands strict handling:\n\n- in-country hosting where regulations require it\n- role-based access down to record level\n- full access logging\n- a data-minimization default for AI features: models see only what the task needs\n- a documented choice of AI provider, including private or self-hosted models where required\n\n## Integrations\n\nEHR and HIS systems (typically via HL7 or FHIR interfaces), payer portals and APIs, scheduling systems, patient messaging, and the identity provider.\n\n## Who uses it\n\nReferral coordinators, care coordinators, pre-authorization teams, department administrators, clinicians (for review and sign-off) and operations leadership.\n\n## First scope\n\nOne referral pathway with a visible waiting-time problem. Measure time from referral to first appointment and the share of referrals returned incomplete. Scope it in a [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint).\n\nSee [healthcare](\u002Findustries\u002Fhealthcare), explore the [Atlas](\u002Fatlas), or [bring us your pathway](\u002Fcontact).\n","\u003Cp>Clinicians spend a significant part of their day on work that isn&#39;t clinical: referral letters, pre-authorization requests, follow-up coordination, chasing results and scheduling across providers. Patients experience that work as waiting.\u003C\u002Fp>\n\u003Cp>The \u003Cstrong>care coordination\u003C\u002Fstrong> family in the Atlas focuses on these administrative and coordination workflows. It doesn&#39;t touch clinical decision-making, and it&#39;s designed so it cannot drift into it.\u003C\u002Fp>\n\u003Ch2>Workflows covered\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Referral intake:\u003C\u002Fstrong> referrals arrive from primary care, other hospitals or payers, and are checked for completeness.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Triage and routing:\u003C\u002Fstrong> referrals go to the right service and are prioritized according to clinical rules defined by the provider.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Pre-authorization:\u003C\u002Fstrong> requests are assembled with the required documentation, submitted to payers and tracked.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Scheduling coordination:\u003C\u002Fstrong> appointments are linked across departments and providers.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Care pathway tasks:\u003C\u002Fstrong> follow-ups, results, patient communication and hand-offs, each with an owner and due date.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Closure and feedback:\u003C\u002Fstrong> outcomes communicated back to the referring provider.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Reporting:\u003C\u002Fstrong> waiting times, bottlenecks and service-level performance.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Where AI helps\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Document extraction:\u003C\u002Fstrong> pull structured data from referral letters and attachments.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Completeness checks:\u003C\u002Fstrong> identify missing information before a referral reaches a coordinator.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Summaries:\u003C\u002Fstrong> a concise case summary for coordinators, drawn from the documents.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Drafting:\u003C\u002Fstrong> pre-authorization justifications and patient communications, for staff to review.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Queue prioritization:\u003C\u002Fstrong> suggestions based on the provider&#39;s own rules, never the model&#39;s opinion of clinical urgency.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Where it must not\u003C\u002Fh2>\n\u003Cp>AI output in this family never replaces clinical judgement. Clinical triage rules are configured by the provider and applied deterministically, and any AI suggestion that touches clinical content is shown to a qualified person before it has effect. Each AI output is labelled and its acceptance recorded.\u003C\u002Fp>\n\u003Ch2>Privacy and hosting\u003C\u002Fh2>\n\u003Cp>Health data demands strict handling:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>in-country hosting where regulations require it\u003C\u002Fli>\n\u003Cli>role-based access down to record level\u003C\u002Fli>\n\u003Cli>full access logging\u003C\u002Fli>\n\u003Cli>a data-minimization default for AI features: models see only what the task needs\u003C\u002Fli>\n\u003Cli>a documented choice of AI provider, including private or self-hosted models where required\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Integrations\u003C\u002Fh2>\n\u003Cp>EHR and HIS systems (typically via HL7 or FHIR interfaces), payer portals and APIs, scheduling systems, patient messaging, and the identity provider.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cp>Referral coordinators, care coordinators, pre-authorization teams, department administrators, clinicians (for review and sign-off) and operations leadership.\u003C\u002Fp>\n\u003Ch2>First scope\u003C\u002Fh2>\n\u003Cp>One referral pathway with a visible waiting-time problem. Measure time from referral to first appointment and the share of referrals returned incomplete. Scope it in a \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>See \u003Ca href=\"\u002Findustries\u002Fhealthcare\">healthcare\u003C\u002Fa>, explore the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your pathway\u003C\u002Fa>.\u003C\u002Fp>\n","Referrals and care coordination: the administrative workflows around care","Healthcare operations applications for referrals, pre-authorization and care coordination, with AI on paperwork and humans on every clinical decision.",[32,33,34,15],"healthcare","document-intelligence","case-management","2026-08-13T00:00:00.000Z",{"id":37,"slug":38,"body":39,"html":40,"title":41,"description":42,"category":11,"tags":43,"author":17,"date":46,"year":19,"month":47,"quarter":21,"status":22,"featured":23},"2026\u002F07\u002Findustry-applications\u002Fai-in-the-soc-triage-and-investigation","ai-in-the-soc-triage-and-investigation","\nSecurity operations centres don't lack alerts. They lack analyst time. Every tool in the stack produces detections, and many are duplicates, benign or low value. Real incidents compete for attention with noise, and analysts spend a large share of their day gathering context rather than making judgements.\n\n## What the application does\n\nThe **security operations** family in the Atlas focuses on the workflow between detection and response:\n\n1. **Ingest:** alerts from SIEM, EDR, email security, identity and cloud security tools, normalized into one model.\n2. **Enrich:** asset ownership, user context, threat intelligence and related alerts attached automatically.\n3. **Correlate:** group related alerts into a single investigation.\n4. **Triage:** prioritize by severity, asset criticality and confidence.\n5. **Investigate:** a case with a timeline, evidence, notes and tasks.\n6. **Respond:** response actions through the organization's tools, with approvals for high-impact steps.\n7. **Close and learn:** a disposition, lessons learned and tuning feedback to the detection owners.\n8. **Report:** metrics for SOC leadership and control evidence for audit.\n\n## Where AI helps\n\n- **Summarization:** a plain-language summary of what happened, affected assets and the evidence so far.\n- **Triage support:** a suggested priority and likely disposition, with the reasoning shown.\n- **Investigation assistance:** suggested next queries and pivots, and drafted incident timelines.\n- **Agentic enrichment:** bounded, read-only lookups across tools to assemble context before an analyst opens the case.\n- **Reporting:** draft incident reports and management summaries.\n\n## Guardrails that matter here\n\nSecurity is where uncontrolled automation does the most damage. The application enforces:\n\n- **Read-only by default.** Enrichment agents can look, not act.\n- **Human approval for containment.** Isolating hosts, disabling accounts and blocking traffic require an analyst, and a second approver for high-impact actions.\n- **Prompt-injection awareness.** Alert content is treated as untrusted data, never as instructions.\n- **A full audit trail** of every AI suggestion, every action and who approved it.\n\nWe cover the general pattern in [agentic automation with human checkpoints](\u002Fblog\u002Fagentic-automation-with-human-checkpoints).\n\n## Who uses it\n\nSOC analysts (tier 1 to 3), incident responders, SOC managers, CISOs, and control owners who need evidence for audits.\n\n## Integrations\n\nSIEM and log platforms, EDR\u002FXDR, identity providers, email security, cloud security posture tools, ticketing and ITSM, threat intelligence feeds, and asset inventories or CMDBs.\n\n## Measuring it honestly\n\nTrack time to triage, time to contain, the share of alerts closed as benign and analyst hours per incident. Agree the baseline first. Improvements should show up in your own metrics, not in vendor claims.\n\n## Where it applies\n\nEnterprise SOCs, managed security providers, financial institutions with regulatory incident-reporting obligations, and government security operations.\n\nExplore the [Atlas](\u002Fatlas), or [bring us your triage queue](\u002Fcontact).\n","\u003Cp>Security operations centres don&#39;t lack alerts. They lack analyst time. Every tool in the stack produces detections, and many are duplicates, benign or low value. Real incidents compete for attention with noise, and analysts spend a large share of their day gathering context rather than making judgements.\u003C\u002Fp>\n\u003Ch2>What the application does\u003C\u002Fh2>\n\u003Cp>The \u003Cstrong>security operations\u003C\u002Fstrong> family in the Atlas focuses on the workflow between detection and response:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Ingest:\u003C\u002Fstrong> alerts from SIEM, EDR, email security, identity and cloud security tools, normalized into one model.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Enrich:\u003C\u002Fstrong> asset ownership, user context, threat intelligence and related alerts attached automatically.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Correlate:\u003C\u002Fstrong> group related alerts into a single investigation.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Triage:\u003C\u002Fstrong> prioritize by severity, asset criticality and confidence.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Investigate:\u003C\u002Fstrong> a case with a timeline, evidence, notes and tasks.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Respond:\u003C\u002Fstrong> response actions through the organization&#39;s tools, with approvals for high-impact steps.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Close and learn:\u003C\u002Fstrong> a disposition, lessons learned and tuning feedback to the detection owners.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Report:\u003C\u002Fstrong> metrics for SOC leadership and control evidence for audit.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2>Where AI helps\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Summarization:\u003C\u002Fstrong> a plain-language summary of what happened, affected assets and the evidence so far.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Triage support:\u003C\u002Fstrong> a suggested priority and likely disposition, with the reasoning shown.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Investigation assistance:\u003C\u002Fstrong> suggested next queries and pivots, and drafted incident timelines.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Agentic enrichment:\u003C\u002Fstrong> bounded, read-only lookups across tools to assemble context before an analyst opens the case.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Reporting:\u003C\u002Fstrong> draft incident reports and management summaries.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Guardrails that matter here\u003C\u002Fh2>\n\u003Cp>Security is where uncontrolled automation does the most damage. The application enforces:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Read-only by default.\u003C\u002Fstrong> Enrichment agents can look, not act.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Human approval for containment.\u003C\u002Fstrong> Isolating hosts, disabling accounts and blocking traffic require an analyst, and a second approver for high-impact actions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Prompt-injection awareness.\u003C\u002Fstrong> Alert content is treated as untrusted data, never as instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>A full audit trail\u003C\u002Fstrong> of every AI suggestion, every action and who approved it.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>We cover the general pattern in \u003Ca href=\"\u002Fblog\u002Fagentic-automation-with-human-checkpoints\">agentic automation with human checkpoints\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cp>SOC analysts (tier 1 to 3), incident responders, SOC managers, CISOs, and control owners who need evidence for audits.\u003C\u002Fp>\n\u003Ch2>Integrations\u003C\u002Fh2>\n\u003Cp>SIEM and log platforms, EDR\u002FXDR, identity providers, email security, cloud security posture tools, ticketing and ITSM, threat intelligence feeds, and asset inventories or CMDBs.\u003C\u002Fp>\n\u003Ch2>Measuring it honestly\u003C\u002Fh2>\n\u003Cp>Track time to triage, time to contain, the share of alerts closed as benign and analyst hours per incident. Agree the baseline first. Improvements should show up in your own metrics, not in vendor claims.\u003C\u002Fp>\n\u003Ch2>Where it applies\u003C\u002Fh2>\n\u003Cp>Enterprise SOCs, managed security providers, financial institutions with regulatory incident-reporting obligations, and government security operations.\u003C\u002Fp>\n\u003Cp>Explore the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your triage queue\u003C\u002Fa>.\u003C\u002Fp>\n","AI in the SOC: alert triage and investigation with evidence","Security operations applications that use AI to enrich, summarize and prioritize alerts while analysts keep the decisions and the evidence trail.",[44,34,45,15,16],"cybersecurity","evidence","2026-07-28T00:00:00.000Z",7,{"id":49,"slug":50,"body":51,"html":52,"title":53,"description":54,"category":55,"tags":56,"author":17,"date":58,"year":19,"month":47,"quarter":21,"status":22,"featured":23},"2026\u002F07\u002Fai-in-production\u002Fdocument-intelligence-in-regulated-workflows","document-intelligence-in-regulated-workflows","\nRegulated workflows run on documents: identity documents, company registries, financial statements, invoices, contracts, permits, medical referrals, supplier certificates, audit reports. Extracting data from them is among the most valuable uses of AI, and among the easiest to get subtly wrong.\n\nA demo extracts ten fields from a clean PDF perfectly. Production brings scans, photos, handwriting, multiple languages, unusual layouts and documents that are simply the wrong document.\n\n## The production pattern\n\n**1. Classify first.** Before extracting anything, determine what the document is. A bank statement sent where a trade licence was expected should be caught at the door.\n\n**2. Extract to a schema.** Every document type has a defined schema of fields, types and formats. The model's output is validated against it, and anything that doesn't conform is rejected.\n\n**3. Validate against rules and sources.** Cross-check extracted values: totals that should add up, dates that should be in order, registration numbers that should exist in a registry, names that should match the application.\n\n**4. Carry confidence and provenance.** Every extracted field records where it came from on the page and how confident the extraction is. Reviewers see the source next to the value.\n\n**5. Route by confidence and risk.** High-confidence, low-risk fields flow straight through. Low-confidence or high-risk fields go to a human review queue. The thresholds are business decisions, not model defaults.\n\n**6. Learn from corrections.** Every human correction is recorded and becomes evaluation data for the next model or prompt change.\n\n## Where it appears across the Atlas\n\nDocument intelligence isn't a product on its own. It's a capability inside many application families:\n\n- **Onboarding and KYC\u002FKYB:** identity and company documents\n- **Case management:** evidence submitted by applicants ([AI-assisted case management](\u002Fblog\u002Fai-assisted-case-management))\n- **Referrals and pre-authorization** in healthcare ([care coordination](\u002Fblog\u002Freferrals-and-care-coordination))\n- **Permits** in the built environment ([permitting and inspections](\u002Fblog\u002Fpermitting-and-inspections-for-the-built-environment))\n- **Supplier assurance:** SOC reports and certificates ([third-party risk](\u002Fblog\u002Fthird-party-and-supplier-risk-reviews))\n- **Finance:** remittances and statements ([reconciliation](\u002Fblog\u002Freconciliation-and-exception-workbenches))\n\n## Controls designed in\n\n- Original documents retained, unaltered, with hashes\n- Extracted values linked to their source location\n- Every human override recorded, with the reviewer and reason\n- Access to sensitive documents restricted by role and logged\n- The model provider and hosting chosen to meet data-residency requirements\n\n## Measuring it honestly\n\nField-level accuracy on a held-out test set per document type, straight-through processing rate, review queue volume and correction rate. Agree the thresholds with the business and compliance owners before launch. See [evaluation and guardrails](\u002Fblog\u002Fevaluation-and-guardrails-before-production).\n\n## Arabic and bilingual documents\n\nIn the GCC, many documents are Arabic, English or both, and include stamps, signatures and handwriting. Test sets must reflect that mix from day one. Performance on English samples says little about performance on the documents you'll actually receive.\n\n[Bring us the document types](\u002Fcontact) that slow your workflow down.\n","\u003Cp>Regulated workflows run on documents: identity documents, company registries, financial statements, invoices, contracts, permits, medical referrals, supplier certificates, audit reports. Extracting data from them is among the most valuable uses of AI, and among the easiest to get subtly wrong.\u003C\u002Fp>\n\u003Cp>A demo extracts ten fields from a clean PDF perfectly. Production brings scans, photos, handwriting, multiple languages, unusual layouts and documents that are simply the wrong document.\u003C\u002Fp>\n\u003Ch2>The production pattern\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. Classify first.\u003C\u002Fstrong> Before extracting anything, determine what the document is. A bank statement sent where a trade licence was expected should be caught at the door.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Extract to a schema.\u003C\u002Fstrong> Every document type has a defined schema of fields, types and formats. The model&#39;s output is validated against it, and anything that doesn&#39;t conform is rejected.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Validate against rules and sources.\u003C\u002Fstrong> Cross-check extracted values: totals that should add up, dates that should be in order, registration numbers that should exist in a registry, names that should match the application.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Carry confidence and provenance.\u003C\u002Fstrong> Every extracted field records where it came from on the page and how confident the extraction is. Reviewers see the source next to the value.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>5. Route by confidence and risk.\u003C\u002Fstrong> High-confidence, low-risk fields flow straight through. Low-confidence or high-risk fields go to a human review queue. The thresholds are business decisions, not model defaults.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>6. Learn from corrections.\u003C\u002Fstrong> Every human correction is recorded and becomes evaluation data for the next model or prompt change.\u003C\u002Fp>\n\u003Ch2>Where it appears across the Atlas\u003C\u002Fh2>\n\u003Cp>Document intelligence isn&#39;t a product on its own. It&#39;s a capability inside many application families:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Onboarding and KYC\u002FKYB:\u003C\u002Fstrong> identity and company documents\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Case management:\u003C\u002Fstrong> evidence submitted by applicants (\u003Ca href=\"\u002Fblog\u002Fai-assisted-case-management\">AI-assisted case management\u003C\u002Fa>)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Referrals and pre-authorization\u003C\u002Fstrong> in healthcare (\u003Ca href=\"\u002Fblog\u002Freferrals-and-care-coordination\">care coordination\u003C\u002Fa>)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Permits\u003C\u002Fstrong> in the built environment (\u003Ca href=\"\u002Fblog\u002Fpermitting-and-inspections-for-the-built-environment\">permitting and inspections\u003C\u002Fa>)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Supplier assurance:\u003C\u002Fstrong> SOC reports and certificates (\u003Ca href=\"\u002Fblog\u002Fthird-party-and-supplier-risk-reviews\">third-party risk\u003C\u002Fa>)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Finance:\u003C\u002Fstrong> remittances and statements (\u003Ca href=\"\u002Fblog\u002Freconciliation-and-exception-workbenches\">reconciliation\u003C\u002Fa>)\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Controls designed in\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>Original documents retained, unaltered, with hashes\u003C\u002Fli>\n\u003Cli>Extracted values linked to their source location\u003C\u002Fli>\n\u003Cli>Every human override recorded, with the reviewer and reason\u003C\u002Fli>\n\u003Cli>Access to sensitive documents restricted by role and logged\u003C\u002Fli>\n\u003Cli>The model provider and hosting chosen to meet data-residency requirements\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Measuring it honestly\u003C\u002Fh2>\n\u003Cp>Field-level accuracy on a held-out test set per document type, straight-through processing rate, review queue volume and correction rate. Agree the thresholds with the business and compliance owners before launch. See \u003Ca href=\"\u002Fblog\u002Fevaluation-and-guardrails-before-production\">evaluation and guardrails\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch2>Arabic and bilingual documents\u003C\u002Fh2>\n\u003Cp>In the GCC, many documents are Arabic, English or both, and include stamps, signatures and handwriting. Test sets must reflect that mix from day one. Performance on English samples says little about performance on the documents you&#39;ll actually receive.\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"\u002Fcontact\">Bring us the document types\u003C\u002Fa> that slow your workflow down.\u003C\u002Fp>\n","Document intelligence in regulated workflows: extraction with verification","Extracting data from documents with AI is easy to demo and hard to trust. How to build extraction with validation, confidence and human review.","ai-in-production",[33,15,45,57],"production","2026-07-23T00:00:00.000Z",{"id":60,"slug":61,"body":62,"html":63,"title":64,"description":65,"category":11,"tags":66,"author":17,"date":69,"year":19,"month":47,"quarter":21,"status":22,"featured":23},"2026\u002F07\u002Findustry-applications\u002Fai-assisted-case-management","ai-assisted-case-management","\nCase management is everywhere once you look for it: benefit applications, licensing requests, complaints, investigations, customer disputes, employee cases, service requests. The shape is the same each time. Something arrives, it's triaged, someone works it, a decision is made and it may be appealed. Backlogs grow when intake outpaces the people who decide.\n\nThat common shape is why case management is one of the most reusable application families in the Atlas, and one of the best places to apply AI safely.\n\n## The core workflow\n\n1. **Intake:** cases arrive through portals, email, APIs or other systems, with documents attached.\n2. **Triage:** each case is classified by type, urgency and complexity, and routed to the right queue.\n3. **Assignment:** workload-aware allocation to case workers, with skills and conflicts respected.\n4. **Work:** information requests, internal consultations, notes and deadlines.\n5. **Decision:** a structured decision with its rationale, approved where policy requires.\n6. **Communication:** notifications and letters to the applicant or customer.\n7. **Appeal or reopen:** a linked case with its full history.\n8. **Reporting:** backlog, ageing, service levels and outcomes.\n\n## Where AI helps\n\n- **Document intelligence:** extract fields from submitted documents and check completeness before a case reaches a person.\n- **Classification and routing:** suggest case type and priority, with the suggestion recorded.\n- **Case summaries:** a short, current summary at the top of every case, so a new case worker doesn't reread forty pages.\n- **Similar-case retrieval:** find precedents and relevant policy passages with citations.\n- **Drafting:** propose decision letters and information requests for the case worker to edit.\n\n## Where it must not\n\nAI never makes the decision in consequential cases. It doesn't deny, approve or close on its own. The workflow puts human checkpoints at every decision, records who decided, and keeps AI-generated text visibly marked until a person accepts it. In the public sector, this is about legitimacy as much as risk: citizens are entitled to an accountable decision-maker.\n\n## Controls designed in\n\n- Role-based access to sensitive case data\n- Conflict-of-interest checks on assignment\n- A complete audit history of every change, view and decision\n- Retention and disclosure rules configured per case type\n\n## Integrations\n\nCitizen or customer portals, national identity and SSO, document management, CRM or registry systems, payment systems for fees, and messaging services.\n\n## Where it applies\n\nGovernment and public services, financial services complaints and disputes, insurance claims triage, HR case management and enterprise service teams. The foundation is the same, and the domain vocabulary and policies are configured.\n\n## First scope\n\nOne case type with a real backlog. Measure time to first touch, time to decision and backlog ageing before and after. Scope it in a [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint).\n\nSee [government and public sector](\u002Findustries\u002Fgovernment-public-sector), explore the [Atlas](\u002Fatlas), or [bring us your backlog](\u002Fcontact).\n","\u003Cp>Case management is everywhere once you look for it: benefit applications, licensing requests, complaints, investigations, customer disputes, employee cases, service requests. The shape is the same each time. Something arrives, it&#39;s triaged, someone works it, a decision is made and it may be appealed. Backlogs grow when intake outpaces the people who decide.\u003C\u002Fp>\n\u003Cp>That common shape is why case management is one of the most reusable application families in the Atlas, and one of the best places to apply AI safely.\u003C\u002Fp>\n\u003Ch2>The core workflow\u003C\u002Fh2>\n\u003Col>\n\u003Cli>\u003Cstrong>Intake:\u003C\u002Fstrong> cases arrive through portals, email, APIs or other systems, with documents attached.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Triage:\u003C\u002Fstrong> each case is classified by type, urgency and complexity, and routed to the right queue.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Assignment:\u003C\u002Fstrong> workload-aware allocation to case workers, with skills and conflicts respected.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Work:\u003C\u002Fstrong> information requests, internal consultations, notes and deadlines.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Decision:\u003C\u002Fstrong> a structured decision with its rationale, approved where policy requires.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Communication:\u003C\u002Fstrong> notifications and letters to the applicant or customer.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Appeal or reopen:\u003C\u002Fstrong> a linked case with its full history.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Reporting:\u003C\u002Fstrong> backlog, ageing, service levels and outcomes.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2>Where AI helps\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Document intelligence:\u003C\u002Fstrong> extract fields from submitted documents and check completeness before a case reaches a person.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Classification and routing:\u003C\u002Fstrong> suggest case type and priority, with the suggestion recorded.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Case summaries:\u003C\u002Fstrong> a short, current summary at the top of every case, so a new case worker doesn&#39;t reread forty pages.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Similar-case retrieval:\u003C\u002Fstrong> find precedents and relevant policy passages with citations.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Drafting:\u003C\u002Fstrong> propose decision letters and information requests for the case worker to edit.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Where it must not\u003C\u002Fh2>\n\u003Cp>AI never makes the decision in consequential cases. It doesn&#39;t deny, approve or close on its own. The workflow puts human checkpoints at every decision, records who decided, and keeps AI-generated text visibly marked until a person accepts it. In the public sector, this is about legitimacy as much as risk: citizens are entitled to an accountable decision-maker.\u003C\u002Fp>\n\u003Ch2>Controls designed in\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>Role-based access to sensitive case data\u003C\u002Fli>\n\u003Cli>Conflict-of-interest checks on assignment\u003C\u002Fli>\n\u003Cli>A complete audit history of every change, view and decision\u003C\u002Fli>\n\u003Cli>Retention and disclosure rules configured per case type\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Integrations\u003C\u002Fh2>\n\u003Cp>Citizen or customer portals, national identity and SSO, document management, CRM or registry systems, payment systems for fees, and messaging services.\u003C\u002Fp>\n\u003Ch2>Where it applies\u003C\u002Fh2>\n\u003Cp>Government and public services, financial services complaints and disputes, insurance claims triage, HR case management and enterprise service teams. The foundation is the same, and the domain vocabulary and policies are configured.\u003C\u002Fp>\n\u003Ch2>First scope\u003C\u002Fh2>\n\u003Cp>One case type with a real backlog. Measure time to first touch, time to decision and backlog ageing before and after. Scope it in a \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>See \u003Ca href=\"\u002Findustries\u002Fgovernment-public-sector\">government and public sector\u003C\u002Fa>, explore the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your backlog\u003C\u002Fa>.\u003C\u002Fp>\n","AI-assisted case management: summaries, triage and human decisions","Case management across government services and enterprise operations: intake, triage, assignment, decisions and appeals, with AI assisting.",[34,67,68,15,33],"government","enterprise-operations","2026-07-16T00:00:00.000Z",{"id":71,"slug":72,"body":73,"html":74,"title":75,"description":76,"category":55,"tags":77,"author":17,"date":78,"year":19,"month":47,"quarter":21,"status":22,"featured":23},"2026\u002F07\u002Fai-in-production\u002Fagentic-automation-with-human-checkpoints","agentic-automation-with-human-checkpoints","\nAgents, meaning AI systems that plan and take multi-step actions with tools, are the most exciting and the most dangerous AI capability in the enterprise. An agent that gathers context from five systems before an analyst opens a case saves real time. An agent that closes accounts, moves money or emails customers on its own is a governance incident waiting to happen.\n\nThe answer isn't to avoid agents. It's to put them inside a workflow with **checkpoints**.\n\n## Design principles\n\n**1. Bounded tools.** An agent can only call tools that the application explicitly exposes to it, each with a narrow purpose and validated inputs. No general shell, no arbitrary API access.\n\n**2. Read before write.** Most value comes from read-only work: gathering context, correlating records, drafting. Make read-only the default and treat every write as a separate, higher-risk capability.\n\n**3. Explicit approval for consequential actions.** Anything that changes a record of consequence, contacts a customer, moves value or changes access requires a person to approve. Some actions require two people.\n\n**4. Identity and least privilege.** The agent acts with its own service identity or on behalf of a user, never with broader permissions than the user who invoked it.\n\n**5. Deterministic workflow state.** The workflow engine, not the model, decides what state a case is in and what happens next. The agent proposes, and the workflow disposes.\n\n**6. Untrusted input.** Content the agent reads (emails, documents, alerts, web pages) is data. Instructions embedded in it are ignored, and attempts are logged.\n\n**7. Full traceability.** Every plan, tool call, input, output, approval and rejection is recorded, so reviewers can reconstruct why something happened.\n\n## Where agents earn their keep\n\n- **Case preparation:** assemble customer, transaction and history context before a human opens the case. See [AI-assisted case management](\u002Fblog\u002Fai-assisted-case-management).\n- **Security enrichment:** read-only lookups across security tools. See [AI in the SOC](\u002Fblog\u002Fai-in-the-soc-triage-and-investigation).\n- **Document workflows:** extract, validate and route documents, and escalate what fails validation.\n- **Operations recovery:** generate and score recovery options for a controller to choose from. See [operations control](\u002Fblog\u002Foperations-control-and-disruption-management).\n- **Reconciliation:** propose matches and classify breaks for an analyst to confirm.\n\nIn each case, the agent compresses the time *before* a human decision. It doesn't replace the decision.\n\n## What to measure\n\nTime saved before the decision point, how often agent proposals are accepted unchanged, the rejection reasons, how often approval gates fire, and incidents caused by agent actions. That last number should be zero, and the design should make it hard to be anything else.\n\n## How it fits the architecture\n\nIn our application foundations, agent tools are ordinary application services with contracts, authorization and tests. That's the same discipline as any other API. This is the practical meaning of [AI accelerates the implementation, architecture governs it](\u002Fblog\u002Fai-accelerates-architecture-governs).\n\n[Bring us a workflow](\u002Fcontact) where an agent could prepare the decision, and we'll scope the checkpoints with you.\n","\u003Cp>Agents, meaning AI systems that plan and take multi-step actions with tools, are the most exciting and the most dangerous AI capability in the enterprise. An agent that gathers context from five systems before an analyst opens a case saves real time. An agent that closes accounts, moves money or emails customers on its own is a governance incident waiting to happen.\u003C\u002Fp>\n\u003Cp>The answer isn&#39;t to avoid agents. It&#39;s to put them inside a workflow with \u003Cstrong>checkpoints\u003C\u002Fstrong>.\u003C\u002Fp>\n\u003Ch2>Design principles\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. Bounded tools.\u003C\u002Fstrong> An agent can only call tools that the application explicitly exposes to it, each with a narrow purpose and validated inputs. No general shell, no arbitrary API access.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Read before write.\u003C\u002Fstrong> Most value comes from read-only work: gathering context, correlating records, drafting. Make read-only the default and treat every write as a separate, higher-risk capability.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Explicit approval for consequential actions.\u003C\u002Fstrong> Anything that changes a record of consequence, contacts a customer, moves value or changes access requires a person to approve. Some actions require two people.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Identity and least privilege.\u003C\u002Fstrong> The agent acts with its own service identity or on behalf of a user, never with broader permissions than the user who invoked it.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>5. Deterministic workflow state.\u003C\u002Fstrong> The workflow engine, not the model, decides what state a case is in and what happens next. The agent proposes, and the workflow disposes.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>6. Untrusted input.\u003C\u002Fstrong> Content the agent reads (emails, documents, alerts, web pages) is data. Instructions embedded in it are ignored, and attempts are logged.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>7. Full traceability.\u003C\u002Fstrong> Every plan, tool call, input, output, approval and rejection is recorded, so reviewers can reconstruct why something happened.\u003C\u002Fp>\n\u003Ch2>Where agents earn their keep\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Case preparation:\u003C\u002Fstrong> assemble customer, transaction and history context before a human opens the case. See \u003Ca href=\"\u002Fblog\u002Fai-assisted-case-management\">AI-assisted case management\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Security enrichment:\u003C\u002Fstrong> read-only lookups across security tools. See \u003Ca href=\"\u002Fblog\u002Fai-in-the-soc-triage-and-investigation\">AI in the SOC\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Document workflows:\u003C\u002Fstrong> extract, validate and route documents, and escalate what fails validation.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Operations recovery:\u003C\u002Fstrong> generate and score recovery options for a controller to choose from. See \u003Ca href=\"\u002Fblog\u002Foperations-control-and-disruption-management\">operations control\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Reconciliation:\u003C\u002Fstrong> propose matches and classify breaks for an analyst to confirm.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>In each case, the agent compresses the time \u003Cem>before\u003C\u002Fem> a human decision. It doesn&#39;t replace the decision.\u003C\u002Fp>\n\u003Ch2>What to measure\u003C\u002Fh2>\n\u003Cp>Time saved before the decision point, how often agent proposals are accepted unchanged, the rejection reasons, how often approval gates fire, and incidents caused by agent actions. That last number should be zero, and the design should make it hard to be anything else.\u003C\u002Fp>\n\u003Ch2>How it fits the architecture\u003C\u002Fh2>\n\u003Cp>In our application foundations, agent tools are ordinary application services with contracts, authorization and tests. That&#39;s the same discipline as any other API. This is the practical meaning of \u003Ca href=\"\u002Fblog\u002Fai-accelerates-architecture-governs\">AI accelerates the implementation, architecture governs it\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"\u002Fcontact\">Bring us a workflow\u003C\u002Fa> where an agent could prepare the decision, and we&#39;ll scope the checkpoints with you.\u003C\u002Fp>\n","Agentic automation with human checkpoints","How to use AI agents in enterprise workflows safely: bounded tools, read-before-write, explicit approvals and an audit trail of every step.",[16,15,57,45],"2026-07-14T00:00:00.000Z",{"id":80,"slug":81,"body":82,"html":83,"title":84,"description":85,"category":55,"tags":86,"author":17,"date":89,"year":19,"month":90,"quarter":91,"status":22,"featured":23},"2026\u002F06\u002Fai-in-production\u002Fevaluation-and-guardrails-before-production","evaluation-and-guardrails-before-production","\nTraditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don't behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\n\nSo AI features need their own form of testing, **evaluation**, and it has to be a delivery gate, not a one-off exercise before a demo.\n\n## Four layers of evaluation\n\n**1. Task quality.** Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\n\n**2. Groundedness.** For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\n\n**3. Safety and policy.** Does the feature refuse what it should: out-of-scope questions, requests for data the user can't access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\n\n**4. Regression.** Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\n\n## Building the test sets\n\nGood test sets come from the workflow, not from the vendor:\n\n- real questions and documents from the pilot, anonymized where needed\n- edge cases the business owner worries about\n- known-hard cases collected from production feedback\n- adversarial cases: injection attempts, ambiguous requests, missing data\n\nEvery case has an expected outcome defined by a person who owns the domain.\n\n## Guardrails in the application\n\nEvaluation tells you how the feature behaves. Guardrails constrain it in production:\n\n- **Grounding rules:** answer only from retrieved, authorized sources, or say you don't know.\n- **Output validation:** structured outputs checked against schemas and business rules before use.\n- **Allow-lists:** an AI can only reference entities that exist. It can't invent a product, a customer or a case number.\n- **Human checkpoints:** consequential outputs are drafts until a person accepts them.\n- **Untrusted-input handling:** document and user content is treated as data, never as instructions.\n- **Fallbacks:** if the model is unavailable or uncertain, the workflow continues deterministically.\n\n## Monitoring after launch\n\nIn production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\n\n## How the factory handles it\n\nIn our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn't close until the agreed evaluation thresholds are met. It's one of the [quality gates](\u002Fservices\u002Fai-production-sprint) we use to decide whether something is done.\n\nRelated: [from AI pilot to production application](\u002Fblog\u002Ffrom-ai-pilot-to-production-application) and [AI model governance as an application](\u002Fblog\u002Fai-model-governance-as-an-application).\n\nHave a pilot that's never been evaluated properly? [Bring it to us](\u002Fcontact).\n","\u003Cp>Traditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don&#39;t behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\u003C\u002Fp>\n\u003Cp>So AI features need their own form of testing, \u003Cstrong>evaluation\u003C\u002Fstrong>, and it has to be a delivery gate, not a one-off exercise before a demo.\u003C\u002Fp>\n\u003Ch2>Four layers of evaluation\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. Task quality.\u003C\u002Fstrong> Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Groundedness.\u003C\u002Fstrong> For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Safety and policy.\u003C\u002Fstrong> Does the feature refuse what it should: out-of-scope questions, requests for data the user can&#39;t access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Regression.\u003C\u002Fstrong> Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\u003C\u002Fp>\n\u003Ch2>Building the test sets\u003C\u002Fh2>\n\u003Cp>Good test sets come from the workflow, not from the vendor:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>real questions and documents from the pilot, anonymized where needed\u003C\u002Fli>\n\u003Cli>edge cases the business owner worries about\u003C\u002Fli>\n\u003Cli>known-hard cases collected from production feedback\u003C\u002Fli>\n\u003Cli>adversarial cases: injection attempts, ambiguous requests, missing data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every case has an expected outcome defined by a person who owns the domain.\u003C\u002Fp>\n\u003Ch2>Guardrails in the application\u003C\u002Fh2>\n\u003Cp>Evaluation tells you how the feature behaves. Guardrails constrain it in production:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Grounding rules:\u003C\u002Fstrong> answer only from retrieved, authorized sources, or say you don&#39;t know.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Output validation:\u003C\u002Fstrong> structured outputs checked against schemas and business rules before use.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Allow-lists:\u003C\u002Fstrong> an AI can only reference entities that exist. It can&#39;t invent a product, a customer or a case number.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Human checkpoints:\u003C\u002Fstrong> consequential outputs are drafts until a person accepts them.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Untrusted-input handling:\u003C\u002Fstrong> document and user content is treated as data, never as instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fallbacks:\u003C\u002Fstrong> if the model is unavailable or uncertain, the workflow continues deterministically.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Monitoring after launch\u003C\u002Fh2>\n\u003Cp>In production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\u003C\u002Fp>\n\u003Ch2>How the factory handles it\u003C\u002Fh2>\n\u003Cp>In our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn&#39;t close until the agreed evaluation thresholds are met. It&#39;s one of the \u003Ca href=\"\u002Fservices\u002Fai-production-sprint\">quality gates\u003C\u002Fa> we use to decide whether something is done.\u003C\u002Fp>\n\u003Cp>Related: \u003Ca href=\"\u002Fblog\u002Ffrom-ai-pilot-to-production-application\">from AI pilot to production application\u003C\u002Fa> and \u003Ca href=\"\u002Fblog\u002Fai-model-governance-as-an-application\">AI model governance as an application\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Have a pilot that&#39;s never been evaluated properly? \u003Ca href=\"\u002Fcontact\">Bring it to us\u003C\u002Fa>.\u003C\u002Fp>\n","Evaluation and guardrails: how to test AI features before production","AI features need evaluation as a delivery gate, just like tests: test sets, groundedness checks, safety checks and regression on every change.",[87,57,88,15],"evaluation","ai-governance","2026-06-30T00:00:00.000Z",6,2,1790080513408]