[{"data":1,"prerenderedAt":65},["ShallowReactive",2],{"blog-tag-evaluation":3},[4,24,38,52],{"id":5,"slug":6,"body":7,"html":8,"title":9,"description":10,"category":11,"tags":12,"author":17,"date":18,"year":19,"month":20,"quarter":21,"status":22,"featured":23},"2026\u002F08\u002Findustry-applications\u002Fgrounded-enterprise-knowledge-assistants","grounded-enterprise-knowledge-assistants","\nThe enterprise knowledge assistant is the most requested AI application and one of the most often abandoned. The pilot answers questions impressively. Then someone notices it confidently quoted a superseded policy, or showed a document the user shouldn't have seen, and trust evaporates.\n\nThose failures aren't model problems. They are **application** problems, and they have application solutions.\n\n## What a grounded assistant needs\n\nThe **knowledge and assistants** family in the Atlas is built around five requirements.\n\n**1. Approved sources only.** The assistant answers from a curated set of repositories (policies, procedures, product documentation, knowledge articles), each with an owner. Content has a lifecycle: draft, approved, superseded. Superseded content is excluded.\n\n**2. Retrieval with citations.** Every answer links to the passages it relies on. If the sources don't support an answer, the assistant says so rather than improvising.\n\n**3. Permission-aware retrieval.** Users only retrieve content they are allowed to see. Permissions come from the source systems and the identity provider, not from a separate copy that drifts.\n\n**4. Evaluation before and after launch.** A test set of real questions with expected answers and sources, run on every change to prompts, models or content. We describe the approach in [evaluation and guardrails](\u002Fblog\u002Fevaluation-and-guardrails-before-production).\n\n**5. Feedback and content ownership.** Users flag wrong or missing answers. Flags become tasks for content owners, so the knowledge base improves instead of the prompt getting longer.\n\n## Beyond Q&A\n\nOnce retrieval is trustworthy, the same foundation supports more useful workflows:\n\n- **Drafting:** first drafts of customer replies, reports or procedures, grounded in approved content\n- **Policy lookup inside other applications:** the case worker or operator sees relevant policy passages in context\n- **Onboarding:** role-specific guided learning over the procedures a new joiner needs\n- **Change impact:** when a policy changes, find the procedures and articles that reference it\n\n## Controls designed in\n\n- Answers restricted to what the user may access\n- Logging of questions, retrieved sources and answers for audit, with retention rules\n- No training on customer data by default, and a documented choice of model provider and hosting\n- Sensitive-content filters configured per deployment\n\n## Integrations\n\nDocument management and intranets, knowledge bases, ticketing systems (resolved tickets are valuable knowledge), the identity provider and directory groups, and the chat or collaboration tools where people already work.\n\n## Who uses it\n\nEveryone, which is why it needs owners: the business owner of each knowledge domain, the AI platform team, and IT for integration and access.\n\n## First scope\n\nOne domain with an owner and a clear audience, such as HR policies, IT support or a product line's procedures. Measure answer accuracy on the test set and the rate of cited answers. Scope it in a [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint).\n\nExplore the [Atlas](\u002Fatlas), or [bring us your knowledge domain](\u002Fcontact).\n","\u003Cp>The enterprise knowledge assistant is the most requested AI application and one of the most often abandoned. The pilot answers questions impressively. Then someone notices it confidently quoted a superseded policy, or showed a document the user shouldn&#39;t have seen, and trust evaporates.\u003C\u002Fp>\n\u003Cp>Those failures aren&#39;t model problems. They are \u003Cstrong>application\u003C\u002Fstrong> problems, and they have application solutions.\u003C\u002Fp>\n\u003Ch2>What a grounded assistant needs\u003C\u002Fh2>\n\u003Cp>The \u003Cstrong>knowledge and assistants\u003C\u002Fstrong> family in the Atlas is built around five requirements.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>1. Approved sources only.\u003C\u002Fstrong> The assistant answers from a curated set of repositories (policies, procedures, product documentation, knowledge articles), each with an owner. Content has a lifecycle: draft, approved, superseded. Superseded content is excluded.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Retrieval with citations.\u003C\u002Fstrong> Every answer links to the passages it relies on. If the sources don&#39;t support an answer, the assistant says so rather than improvising.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Permission-aware retrieval.\u003C\u002Fstrong> Users only retrieve content they are allowed to see. Permissions come from the source systems and the identity provider, not from a separate copy that drifts.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Evaluation before and after launch.\u003C\u002Fstrong> A test set of real questions with expected answers and sources, run on every change to prompts, models or content. We describe the approach in \u003Ca href=\"\u002Fblog\u002Fevaluation-and-guardrails-before-production\">evaluation and guardrails\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>5. Feedback and content ownership.\u003C\u002Fstrong> Users flag wrong or missing answers. Flags become tasks for content owners, so the knowledge base improves instead of the prompt getting longer.\u003C\u002Fp>\n\u003Ch2>Beyond Q&amp;A\u003C\u002Fh2>\n\u003Cp>Once retrieval is trustworthy, the same foundation supports more useful workflows:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Drafting:\u003C\u002Fstrong> first drafts of customer replies, reports or procedures, grounded in approved content\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Policy lookup inside other applications:\u003C\u002Fstrong> the case worker or operator sees relevant policy passages in context\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Onboarding:\u003C\u002Fstrong> role-specific guided learning over the procedures a new joiner needs\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Change impact:\u003C\u002Fstrong> when a policy changes, find the procedures and articles that reference it\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Controls designed in\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>Answers restricted to what the user may access\u003C\u002Fli>\n\u003Cli>Logging of questions, retrieved sources and answers for audit, with retention rules\u003C\u002Fli>\n\u003Cli>No training on customer data by default, and a documented choice of model provider and hosting\u003C\u002Fli>\n\u003Cli>Sensitive-content filters configured per deployment\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Integrations\u003C\u002Fh2>\n\u003Cp>Document management and intranets, knowledge bases, ticketing systems (resolved tickets are valuable knowledge), the identity provider and directory groups, and the chat or collaboration tools where people already work.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cp>Everyone, which is why it needs owners: the business owner of each knowledge domain, the AI platform team, and IT for integration and access.\u003C\u002Fp>\n\u003Ch2>First scope\u003C\u002Fh2>\n\u003Cp>One domain with an owner and a clear audience, such as HR policies, IT support or a product line&#39;s procedures. Measure answer accuracy on the test set and the rate of cited answers. Scope it in a \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Explore the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your knowledge domain\u003C\u002Fa>.\u003C\u002Fp>\n","Grounded enterprise knowledge assistants: retrieval, citations and permissions","How to build an internal knowledge assistant people trust: retrieval over approved sources, citations, permission-aware answers and evaluation.","industry-applications",[13,14,15,16],"knowledge-retrieval","enterprise-operations","evaluation","identity","fazezero-editorial","2026-08-06T00:00:00.000Z",2026,8,3,"published",false,{"id":25,"slug":26,"body":27,"html":28,"title":29,"description":30,"category":11,"tags":31,"author":17,"date":36,"year":19,"month":37,"quarter":21,"status":22,"featured":23},"2026\u002F07\u002Findustry-applications\u002Fai-model-governance-as-an-application","ai-model-governance-as-an-application","\nMost enterprises now have an AI policy. Far fewer have an AI governance **system**. The policy says every model must be inventoried, evaluated, approved and monitored. In practice, the inventory is a spreadsheet, the evaluations are in notebooks, approvals happen in email and monitoring depends on whoever built the model.\n\nThat works for five models. It fails at fifty, and it fails immediately when an auditor or supervisor asks for evidence.\n\n## The workflow behind “AI governance”\n\nThe **AI governance** family in the Atlas treats governance as an operational workflow with a system of record:\n\n1. **Register.** Every model and AI use case gets an owner, a purpose, a risk tier, its data sources and where it is deployed. That includes vendor models, LLM features and internal models.\n2. **Evaluate.** Structured evaluations against defined criteria: accuracy, robustness, bias and fairness, and for LLM features, groundedness and safety. Results are stored as evidence, not screenshots.\n3. **Approve.** Deployment requests route through the right reviewers, such as model risk, security, the business owner and compliance, based on the risk tier. Every decision is recorded.\n4. **Monitor.** Production behaviour is tracked against thresholds. Drift and incidents raise cases with owners.\n5. **Evidence.** Packs for internal audit, the board or supervisors are generated from the record.\n\n## Where AI helps inside the governance application\n\nIt sounds recursive, but it's useful:\n\n- **Summarization** of model documentation and evaluation results for reviewers\n- **Classification** of new use cases into risk tiers, as a suggestion for a human to confirm\n- **Evaluation assistance**, generating test cases and red-team prompts for LLM features\n- **Drafting** evidence-pack narratives from structured records\n\nEvery one of these is a draft for a human. The approval decision is never automated.\n\n## Who uses it\n\n- **Head of AI and the AI platform team:** keep the portfolio visible and deployable.\n- **Model risk managers:** run reviews with consistent criteria.\n- **Risk and compliance officers:** answer supervisors and auditors from one record.\n- **CIO, CDO and CDAO:** see where AI is used, by whom, and at what risk.\n\n## Integrations that matter\n\nModel registries and ML platforms, CI\u002FCD pipelines (so deployment approval is a real gate rather than a formality), the identity provider for reviewer roles, ticketing, and data catalogues for lineage.\n\n## Controls designed in\n\n- Segregation between model owner and approver\n- An immutable decision history\n- Required evidence before approval can proceed\n- Periodic re-review based on risk tier and staleness\n- Role-based access to sensitive evaluation data\n\n## Why it belongs in financial services first\n\nBanks and insurers already run model risk management for credit and pricing models. Generative AI has multiplied the number of “models” and blurred their edges. A governance application extends existing discipline to the new portfolio instead of creating a parallel process.\n\nThe same foundation applies across enterprise operations, government and any organization preparing for AI-specific regulation.\n\n## Starting point\n\nThe fastest start is to take one line of business's AI inventory and move it into the application, with the approval workflow switched on for new deployments only. The [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint) scopes the delta: your risk tiers, reviewers, evaluation criteria and integrations.\n\nSee the [financial services](\u002Findustries\u002Ffinancial-services) page, search the [Atlas](\u002Fatlas), or [bring us your AI inventory](\u002Fcontact).\n","\u003Cp>Most enterprises now have an AI policy. Far fewer have an AI governance \u003Cstrong>system\u003C\u002Fstrong>. The policy says every model must be inventoried, evaluated, approved and monitored. In practice, the inventory is a spreadsheet, the evaluations are in notebooks, approvals happen in email and monitoring depends on whoever built the model.\u003C\u002Fp>\n\u003Cp>That works for five models. It fails at fifty, and it fails immediately when an auditor or supervisor asks for evidence.\u003C\u002Fp>\n\u003Ch2>The workflow behind “AI governance”\u003C\u002Fh2>\n\u003Cp>The \u003Cstrong>AI governance\u003C\u002Fstrong> family in the Atlas treats governance as an operational workflow with a system of record:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Register.\u003C\u002Fstrong> Every model and AI use case gets an owner, a purpose, a risk tier, its data sources and where it is deployed. That includes vendor models, LLM features and internal models.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evaluate.\u003C\u002Fstrong> Structured evaluations against defined criteria: accuracy, robustness, bias and fairness, and for LLM features, groundedness and safety. Results are stored as evidence, not screenshots.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Approve.\u003C\u002Fstrong> Deployment requests route through the right reviewers, such as model risk, security, the business owner and compliance, based on the risk tier. Every decision is recorded.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Monitor.\u003C\u002Fstrong> Production behaviour is tracked against thresholds. Drift and incidents raise cases with owners.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evidence.\u003C\u002Fstrong> Packs for internal audit, the board or supervisors are generated from the record.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2>Where AI helps inside the governance application\u003C\u002Fh2>\n\u003Cp>It sounds recursive, but it&#39;s useful:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Summarization\u003C\u002Fstrong> of model documentation and evaluation results for reviewers\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Classification\u003C\u002Fstrong> of new use cases into risk tiers, as a suggestion for a human to confirm\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evaluation assistance\u003C\u002Fstrong>, generating test cases and red-team prompts for LLM features\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Drafting\u003C\u002Fstrong> evidence-pack narratives from structured records\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every one of these is a draft for a human. The approval decision is never automated.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Head of AI and the AI platform team:\u003C\u002Fstrong> keep the portfolio visible and deployable.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Model risk managers:\u003C\u002Fstrong> run reviews with consistent criteria.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Risk and compliance officers:\u003C\u002Fstrong> answer supervisors and auditors from one record.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>CIO, CDO and CDAO:\u003C\u002Fstrong> see where AI is used, by whom, and at what risk.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Integrations that matter\u003C\u002Fh2>\n\u003Cp>Model registries and ML platforms, CI\u002FCD pipelines (so deployment approval is a real gate rather than a formality), the identity provider for reviewer roles, ticketing, and data catalogues for lineage.\u003C\u002Fp>\n\u003Ch2>Controls designed in\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>Segregation between model owner and approver\u003C\u002Fli>\n\u003Cli>An immutable decision history\u003C\u002Fli>\n\u003Cli>Required evidence before approval can proceed\u003C\u002Fli>\n\u003Cli>Periodic re-review based on risk tier and staleness\u003C\u002Fli>\n\u003Cli>Role-based access to sensitive evaluation data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Why it belongs in financial services first\u003C\u002Fh2>\n\u003Cp>Banks and insurers already run model risk management for credit and pricing models. Generative AI has multiplied the number of “models” and blurred their edges. A governance application extends existing discipline to the new portfolio instead of creating a parallel process.\u003C\u002Fp>\n\u003Cp>The same foundation applies across enterprise operations, government and any organization preparing for AI-specific regulation.\u003C\u002Fp>\n\u003Ch2>Starting point\u003C\u002Fh2>\n\u003Cp>The fastest start is to take one line of business&#39;s AI inventory and move it into the application, with the approval workflow switched on for new deployments only. The \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa> scopes the delta: your risk tiers, reviewers, evaluation criteria and integrations.\u003C\u002Fp>\n\u003Cp>See the \u003Ca href=\"\u002Findustries\u002Ffinancial-services\">financial services\u003C\u002Fa> page, search the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your AI inventory\u003C\u002Fa>.\u003C\u002Fp>\n","AI model governance should be an application, not a policy document","Model inventory, evaluation, deployment approval and monitoring as one governed workflow, so AI governance produces evidence instead of meetings.",[32,33,15,34,35],"ai-governance","financial-services","evidence","risk","2026-07-02T00:00:00.000Z",7,{"id":39,"slug":40,"body":41,"html":42,"title":43,"description":44,"category":45,"tags":46,"author":17,"date":49,"year":19,"month":50,"quarter":51,"status":22,"featured":23},"2026\u002F06\u002Fai-in-production\u002Fevaluation-and-guardrails-before-production","evaluation-and-guardrails-before-production","\nTraditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don't behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\n\nSo AI features need their own form of testing, **evaluation**, and it has to be a delivery gate, not a one-off exercise before a demo.\n\n## Four layers of evaluation\n\n**1. Task quality.** Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\n\n**2. Groundedness.** For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\n\n**3. Safety and policy.** Does the feature refuse what it should: out-of-scope questions, requests for data the user can't access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\n\n**4. Regression.** Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\n\n## Building the test sets\n\nGood test sets come from the workflow, not from the vendor:\n\n- real questions and documents from the pilot, anonymized where needed\n- edge cases the business owner worries about\n- known-hard cases collected from production feedback\n- adversarial cases: injection attempts, ambiguous requests, missing data\n\nEvery case has an expected outcome defined by a person who owns the domain.\n\n## Guardrails in the application\n\nEvaluation tells you how the feature behaves. Guardrails constrain it in production:\n\n- **Grounding rules:** answer only from retrieved, authorized sources, or say you don't know.\n- **Output validation:** structured outputs checked against schemas and business rules before use.\n- **Allow-lists:** an AI can only reference entities that exist. It can't invent a product, a customer or a case number.\n- **Human checkpoints:** consequential outputs are drafts until a person accepts them.\n- **Untrusted-input handling:** document and user content is treated as data, never as instructions.\n- **Fallbacks:** if the model is unavailable or uncertain, the workflow continues deterministically.\n\n## Monitoring after launch\n\nIn production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\n\n## How the factory handles it\n\nIn our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn't close until the agreed evaluation thresholds are met. It's one of the [quality gates](\u002Fservices\u002Fai-production-sprint) we use to decide whether something is done.\n\nRelated: [from AI pilot to production application](\u002Fblog\u002Ffrom-ai-pilot-to-production-application) and [AI model governance as an application](\u002Fblog\u002Fai-model-governance-as-an-application).\n\nHave a pilot that's never been evaluated properly? [Bring it to us](\u002Fcontact).\n","\u003Cp>Traditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don&#39;t behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\u003C\u002Fp>\n\u003Cp>So AI features need their own form of testing, \u003Cstrong>evaluation\u003C\u002Fstrong>, and it has to be a delivery gate, not a one-off exercise before a demo.\u003C\u002Fp>\n\u003Ch2>Four layers of evaluation\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. Task quality.\u003C\u002Fstrong> Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Groundedness.\u003C\u002Fstrong> For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Safety and policy.\u003C\u002Fstrong> Does the feature refuse what it should: out-of-scope questions, requests for data the user can&#39;t access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Regression.\u003C\u002Fstrong> Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\u003C\u002Fp>\n\u003Ch2>Building the test sets\u003C\u002Fh2>\n\u003Cp>Good test sets come from the workflow, not from the vendor:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>real questions and documents from the pilot, anonymized where needed\u003C\u002Fli>\n\u003Cli>edge cases the business owner worries about\u003C\u002Fli>\n\u003Cli>known-hard cases collected from production feedback\u003C\u002Fli>\n\u003Cli>adversarial cases: injection attempts, ambiguous requests, missing data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every case has an expected outcome defined by a person who owns the domain.\u003C\u002Fp>\n\u003Ch2>Guardrails in the application\u003C\u002Fh2>\n\u003Cp>Evaluation tells you how the feature behaves. Guardrails constrain it in production:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Grounding rules:\u003C\u002Fstrong> answer only from retrieved, authorized sources, or say you don&#39;t know.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Output validation:\u003C\u002Fstrong> structured outputs checked against schemas and business rules before use.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Allow-lists:\u003C\u002Fstrong> an AI can only reference entities that exist. It can&#39;t invent a product, a customer or a case number.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Human checkpoints:\u003C\u002Fstrong> consequential outputs are drafts until a person accepts them.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Untrusted-input handling:\u003C\u002Fstrong> document and user content is treated as data, never as instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fallbacks:\u003C\u002Fstrong> if the model is unavailable or uncertain, the workflow continues deterministically.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Monitoring after launch\u003C\u002Fh2>\n\u003Cp>In production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\u003C\u002Fp>\n\u003Ch2>How the factory handles it\u003C\u002Fh2>\n\u003Cp>In our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn&#39;t close until the agreed evaluation thresholds are met. It&#39;s one of the \u003Ca href=\"\u002Fservices\u002Fai-production-sprint\">quality gates\u003C\u002Fa> we use to decide whether something is done.\u003C\u002Fp>\n\u003Cp>Related: \u003Ca href=\"\u002Fblog\u002Ffrom-ai-pilot-to-production-application\">from AI pilot to production application\u003C\u002Fa> and \u003Ca href=\"\u002Fblog\u002Fai-model-governance-as-an-application\">AI model governance as an application\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Have a pilot that&#39;s never been evaluated properly? \u003Ca href=\"\u002Fcontact\">Bring it to us\u003C\u002Fa>.\u003C\u002Fp>\n","Evaluation and guardrails: how to test AI features before production","AI features need evaluation as a delivery gate, just like tests: test sets, groundedness checks, safety checks and regression on every change.","ai-in-production",[15,47,32,48],"production","human-in-the-loop","2026-06-30T00:00:00.000Z",6,2,{"id":53,"slug":54,"body":55,"html":56,"title":57,"description":58,"category":45,"tags":59,"author":17,"date":62,"year":19,"month":50,"quarter":51,"status":22,"featured":23,"series":63,"seriesOrder":64},"2026\u002F06\u002Fai-in-production\u002Ffrom-ai-pilot-to-production-application","from-ai-pilot-to-production-application","\nThere is a question we ask early in almost every conversation:\n\n> **How many AI use cases have you identified or prototyped that are not yet operating as production applications?**\n\nThe answer is rarely zero. Often it's a backlog: copilots that impressed a steering committee, agents that worked on sample data, proofs of concept that never got a production owner. The models were fine. Everything *around* the models was missing.\n\n## What a pilot proves, and what it doesn't\n\nA pilot proves that a model can do a task on representative data in a controlled setting. That is worth knowing. It does **not** prove that:\n\n- real users can reach it through enterprise identity with the right permissions\n- it integrates with the systems of record the workflow depends on\n- someone owns the output and the decisions made with it\n- failures, exceptions and edge cases have somewhere to go\n- quality is measured continuously, not once\n- the organization can audit what happened last Tuesday\n- it can be deployed, monitored and supported in the customer's cloud\n\nThose are the things production is made of.\n\n## Seven changes between pilot and production\n\n**1. From a notebook to an application.** Production AI lives inside a workflow application with a domain model, APIs and a user interface, not in a script or a chat window.\n\n**2. From shared keys to enterprise identity.** Users authenticate through the organization's identity provider. Authorization decides who can see which data and approve which actions. Multi-tenant deployments isolate data by design.\n\n**3. From sample data to integrations.** The application reads and writes through adapters to core systems, document stores and data platforms, with error handling, retries and idempotency.\n\n**4. From a demo to a human-accountable workflow.** AI drafts, summarizes, classifies and recommends. People approve, decide and remain accountable, with clear checkpoints in the workflow.\n\n**5. From one-off testing to evaluation as a gate.** Automated tests cover the application. AI evaluations cover the model behaviour, and both run on every change. See [evaluation and guardrails](\u002Fblog\u002Fevaluation-and-guardrails-before-production).\n\n**6. From “it worked” to evidence.** Every significant action produces an audit event. Controls produce evidence as a by-product of the work.\n\n**7. From a laptop to a deployment baseline.** Compute, API gateway, data, secrets, identity, AI services, observability and CI\u002FCD are defined for the target cloud.\n\n## Why most pilots stall at step two\n\nPilots are usually built to answer a capability question quickly, which is the right way to run a pilot. The trouble starts when the pilot code becomes the starting point for production. Identity, integration and controls then have to be retrofitted into a structure that never expected them, and that is where timelines collapse.\n\nWe take the opposite route. The pilot's *learning* carries forward: the prompts, the evaluation data, the workflow insight. The pilot's *code* usually doesn't. The production application starts from a deployment-ready foundation that already has steps 1, 2, 5, 6 and 7 built in, so the engineering effort goes into integration and the customer-specific workflow.\n\n## A practical path\n\n1. Pick the pilot with a named business owner and a workflow that runs weekly.\n2. Map it to the closest application foundation in the [Atlas](\u002Fatlas).\n3. Define the delta in a [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint): integrations, identity, controls, evaluations and data.\n4. Build it in an [AI Production Sprint](\u002Fservices\u002Fai-production-sprint).\n5. Once it is running, add the adjacent workflows through an [Application Family Program](\u002Fservices\u002Fapplication-family-program).\n\nIf you have a backlog of pilots, [bring us the one that matters most](\u002Fcontact).\n","\u003Cp>There is a question we ask early in almost every conversation:\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>\u003Cstrong>How many AI use cases have you identified or prototyped that are not yet operating as production applications?\u003C\u002Fstrong>\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Cp>The answer is rarely zero. Often it&#39;s a backlog: copilots that impressed a steering committee, agents that worked on sample data, proofs of concept that never got a production owner. The models were fine. Everything \u003Cem>around\u003C\u002Fem> the models was missing.\u003C\u002Fp>\n\u003Ch2>What a pilot proves, and what it doesn&#39;t\u003C\u002Fh2>\n\u003Cp>A pilot proves that a model can do a task on representative data in a controlled setting. That is worth knowing. It does \u003Cstrong>not\u003C\u002Fstrong> prove that:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>real users can reach it through enterprise identity with the right permissions\u003C\u002Fli>\n\u003Cli>it integrates with the systems of record the workflow depends on\u003C\u002Fli>\n\u003Cli>someone owns the output and the decisions made with it\u003C\u002Fli>\n\u003Cli>failures, exceptions and edge cases have somewhere to go\u003C\u002Fli>\n\u003Cli>quality is measured continuously, not once\u003C\u002Fli>\n\u003Cli>the organization can audit what happened last Tuesday\u003C\u002Fli>\n\u003Cli>it can be deployed, monitored and supported in the customer&#39;s cloud\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Those are the things production is made of.\u003C\u002Fp>\n\u003Ch2>Seven changes between pilot and production\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. From a notebook to an application.\u003C\u002Fstrong> Production AI lives inside a workflow application with a domain model, APIs and a user interface, not in a script or a chat window.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. From shared keys to enterprise identity.\u003C\u002Fstrong> Users authenticate through the organization&#39;s identity provider. Authorization decides who can see which data and approve which actions. Multi-tenant deployments isolate data by design.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. From sample data to integrations.\u003C\u002Fstrong> The application reads and writes through adapters to core systems, document stores and data platforms, with error handling, retries and idempotency.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. From a demo to a human-accountable workflow.\u003C\u002Fstrong> AI drafts, summarizes, classifies and recommends. People approve, decide and remain accountable, with clear checkpoints in the workflow.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>5. From one-off testing to evaluation as a gate.\u003C\u002Fstrong> Automated tests cover the application. AI evaluations cover the model behaviour, and both run on every change. See \u003Ca href=\"\u002Fblog\u002Fevaluation-and-guardrails-before-production\">evaluation and guardrails\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>6. From “it worked” to evidence.\u003C\u002Fstrong> Every significant action produces an audit event. Controls produce evidence as a by-product of the work.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>7. From a laptop to a deployment baseline.\u003C\u002Fstrong> Compute, API gateway, data, secrets, identity, AI services, observability and CI\u002FCD are defined for the target cloud.\u003C\u002Fp>\n\u003Ch2>Why most pilots stall at step two\u003C\u002Fh2>\n\u003Cp>Pilots are usually built to answer a capability question quickly, which is the right way to run a pilot. The trouble starts when the pilot code becomes the starting point for production. Identity, integration and controls then have to be retrofitted into a structure that never expected them, and that is where timelines collapse.\u003C\u002Fp>\n\u003Cp>We take the opposite route. The pilot&#39;s \u003Cem>learning\u003C\u002Fem> carries forward: the prompts, the evaluation data, the workflow insight. The pilot&#39;s \u003Cem>code\u003C\u002Fem> usually doesn&#39;t. The production application starts from a deployment-ready foundation that already has steps 1, 2, 5, 6 and 7 built in, so the engineering effort goes into integration and the customer-specific workflow.\u003C\u002Fp>\n\u003Ch2>A practical path\u003C\u002Fh2>\n\u003Col>\n\u003Cli>Pick the pilot with a named business owner and a workflow that runs weekly.\u003C\u002Fli>\n\u003Cli>Map it to the closest application foundation in the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>Define the delta in a \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa>: integrations, identity, controls, evaluations and data.\u003C\u002Fli>\n\u003Cli>Build it in an \u003Ca href=\"\u002Fservices\u002Fai-production-sprint\">AI Production Sprint\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>Once it is running, add the adjacent workflows through an \u003Ca href=\"\u002Fservices\u002Fapplication-family-program\">Application Family Program\u003C\u002Fa>.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>If you have a backlog of pilots, \u003Ca href=\"\u002Fcontact\">bring us the one that matters most\u003C\u002Fa>.\u003C\u002Fp>\n","From AI pilot to production application: what actually changes","Most enterprise AI pilots never reach production. The gap is not the model. It is identity, integration, controls, evaluation and ownership.",[47,60,61,15],"implementation","enterprise","2026-06-25T00:00:00.000Z","the-application-factory",4,1790080513358]