[{"data":1,"prerenderedAt":39},["ShallowReactive",2],{"blog-tag-ai-governance":3},[4,25],{"id":5,"slug":6,"body":7,"html":8,"title":9,"description":10,"category":11,"tags":12,"author":18,"date":19,"year":20,"month":21,"quarter":22,"status":23,"featured":24},"2026\u002F07\u002Findustry-applications\u002Fai-model-governance-as-an-application","ai-model-governance-as-an-application","\nMost enterprises now have an AI policy. Far fewer have an AI governance **system**. The policy says every model must be inventoried, evaluated, approved and monitored. In practice, the inventory is a spreadsheet, the evaluations are in notebooks, approvals happen in email and monitoring depends on whoever built the model.\n\nThat works for five models. It fails at fifty, and it fails immediately when an auditor or supervisor asks for evidence.\n\n## The workflow behind “AI governance”\n\nThe **AI governance** family in the Atlas treats governance as an operational workflow with a system of record:\n\n1. **Register.** Every model and AI use case gets an owner, a purpose, a risk tier, its data sources and where it is deployed. That includes vendor models, LLM features and internal models.\n2. **Evaluate.** Structured evaluations against defined criteria: accuracy, robustness, bias and fairness, and for LLM features, groundedness and safety. Results are stored as evidence, not screenshots.\n3. **Approve.** Deployment requests route through the right reviewers, such as model risk, security, the business owner and compliance, based on the risk tier. Every decision is recorded.\n4. **Monitor.** Production behaviour is tracked against thresholds. Drift and incidents raise cases with owners.\n5. **Evidence.** Packs for internal audit, the board or supervisors are generated from the record.\n\n## Where AI helps inside the governance application\n\nIt sounds recursive, but it's useful:\n\n- **Summarization** of model documentation and evaluation results for reviewers\n- **Classification** of new use cases into risk tiers, as a suggestion for a human to confirm\n- **Evaluation assistance**, generating test cases and red-team prompts for LLM features\n- **Drafting** evidence-pack narratives from structured records\n\nEvery one of these is a draft for a human. The approval decision is never automated.\n\n## Who uses it\n\n- **Head of AI and the AI platform team:** keep the portfolio visible and deployable.\n- **Model risk managers:** run reviews with consistent criteria.\n- **Risk and compliance officers:** answer supervisors and auditors from one record.\n- **CIO, CDO and CDAO:** see where AI is used, by whom, and at what risk.\n\n## Integrations that matter\n\nModel registries and ML platforms, CI\u002FCD pipelines (so deployment approval is a real gate rather than a formality), the identity provider for reviewer roles, ticketing, and data catalogues for lineage.\n\n## Controls designed in\n\n- Segregation between model owner and approver\n- An immutable decision history\n- Required evidence before approval can proceed\n- Periodic re-review based on risk tier and staleness\n- Role-based access to sensitive evaluation data\n\n## Why it belongs in financial services first\n\nBanks and insurers already run model risk management for credit and pricing models. Generative AI has multiplied the number of “models” and blurred their edges. A governance application extends existing discipline to the new portfolio instead of creating a parallel process.\n\nThe same foundation applies across enterprise operations, government and any organization preparing for AI-specific regulation.\n\n## Starting point\n\nThe fastest start is to take one line of business's AI inventory and move it into the application, with the approval workflow switched on for new deployments only. The [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint) scopes the delta: your risk tiers, reviewers, evaluation criteria and integrations.\n\nSee the [financial services](\u002Findustries\u002Ffinancial-services) page, search the [Atlas](\u002Fatlas), or [bring us your AI inventory](\u002Fcontact).\n","\u003Cp>Most enterprises now have an AI policy. Far fewer have an AI governance \u003Cstrong>system\u003C\u002Fstrong>. The policy says every model must be inventoried, evaluated, approved and monitored. In practice, the inventory is a spreadsheet, the evaluations are in notebooks, approvals happen in email and monitoring depends on whoever built the model.\u003C\u002Fp>\n\u003Cp>That works for five models. It fails at fifty, and it fails immediately when an auditor or supervisor asks for evidence.\u003C\u002Fp>\n\u003Ch2>The workflow behind “AI governance”\u003C\u002Fh2>\n\u003Cp>The \u003Cstrong>AI governance\u003C\u002Fstrong> family in the Atlas treats governance as an operational workflow with a system of record:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Register.\u003C\u002Fstrong> Every model and AI use case gets an owner, a purpose, a risk tier, its data sources and where it is deployed. That includes vendor models, LLM features and internal models.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evaluate.\u003C\u002Fstrong> Structured evaluations against defined criteria: accuracy, robustness, bias and fairness, and for LLM features, groundedness and safety. Results are stored as evidence, not screenshots.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Approve.\u003C\u002Fstrong> Deployment requests route through the right reviewers, such as model risk, security, the business owner and compliance, based on the risk tier. Every decision is recorded.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Monitor.\u003C\u002Fstrong> Production behaviour is tracked against thresholds. Drift and incidents raise cases with owners.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evidence.\u003C\u002Fstrong> Packs for internal audit, the board or supervisors are generated from the record.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2>Where AI helps inside the governance application\u003C\u002Fh2>\n\u003Cp>It sounds recursive, but it&#39;s useful:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Summarization\u003C\u002Fstrong> of model documentation and evaluation results for reviewers\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Classification\u003C\u002Fstrong> of new use cases into risk tiers, as a suggestion for a human to confirm\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evaluation assistance\u003C\u002Fstrong>, generating test cases and red-team prompts for LLM features\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Drafting\u003C\u002Fstrong> evidence-pack narratives from structured records\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every one of these is a draft for a human. The approval decision is never automated.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Head of AI and the AI platform team:\u003C\u002Fstrong> keep the portfolio visible and deployable.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Model risk managers:\u003C\u002Fstrong> run reviews with consistent criteria.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Risk and compliance officers:\u003C\u002Fstrong> answer supervisors and auditors from one record.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>CIO, CDO and CDAO:\u003C\u002Fstrong> see where AI is used, by whom, and at what risk.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Integrations that matter\u003C\u002Fh2>\n\u003Cp>Model registries and ML platforms, CI\u002FCD pipelines (so deployment approval is a real gate rather than a formality), the identity provider for reviewer roles, ticketing, and data catalogues for lineage.\u003C\u002Fp>\n\u003Ch2>Controls designed in\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>Segregation between model owner and approver\u003C\u002Fli>\n\u003Cli>An immutable decision history\u003C\u002Fli>\n\u003Cli>Required evidence before approval can proceed\u003C\u002Fli>\n\u003Cli>Periodic re-review based on risk tier and staleness\u003C\u002Fli>\n\u003Cli>Role-based access to sensitive evaluation data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Why it belongs in financial services first\u003C\u002Fh2>\n\u003Cp>Banks and insurers already run model risk management for credit and pricing models. Generative AI has multiplied the number of “models” and blurred their edges. A governance application extends existing discipline to the new portfolio instead of creating a parallel process.\u003C\u002Fp>\n\u003Cp>The same foundation applies across enterprise operations, government and any organization preparing for AI-specific regulation.\u003C\u002Fp>\n\u003Ch2>Starting point\u003C\u002Fh2>\n\u003Cp>The fastest start is to take one line of business&#39;s AI inventory and move it into the application, with the approval workflow switched on for new deployments only. The \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa> scopes the delta: your risk tiers, reviewers, evaluation criteria and integrations.\u003C\u002Fp>\n\u003Cp>See the \u003Ca href=\"\u002Findustries\u002Ffinancial-services\">financial services\u003C\u002Fa> page, search the \u003Ca href=\"\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your AI inventory\u003C\u002Fa>.\u003C\u002Fp>\n","AI model governance should be an application, not a policy document","Model inventory, evaluation, deployment approval and monitoring as one governed workflow, so AI governance produces evidence instead of meetings.","industry-applications",[13,14,15,16,17],"ai-governance","financial-services","evaluation","evidence","risk","fazezero-editorial","2026-07-02T00:00:00.000Z",2026,7,3,"published",false,{"id":26,"slug":27,"body":28,"html":29,"title":30,"description":31,"category":32,"tags":33,"author":18,"date":36,"year":20,"month":37,"quarter":38,"status":23,"featured":24},"2026\u002F06\u002Fai-in-production\u002Fevaluation-and-guardrails-before-production","evaluation-and-guardrails-before-production","\nTraditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don't behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\n\nSo AI features need their own form of testing, **evaluation**, and it has to be a delivery gate, not a one-off exercise before a demo.\n\n## Four layers of evaluation\n\n**1. Task quality.** Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\n\n**2. Groundedness.** For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\n\n**3. Safety and policy.** Does the feature refuse what it should: out-of-scope questions, requests for data the user can't access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\n\n**4. Regression.** Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\n\n## Building the test sets\n\nGood test sets come from the workflow, not from the vendor:\n\n- real questions and documents from the pilot, anonymized where needed\n- edge cases the business owner worries about\n- known-hard cases collected from production feedback\n- adversarial cases: injection attempts, ambiguous requests, missing data\n\nEvery case has an expected outcome defined by a person who owns the domain.\n\n## Guardrails in the application\n\nEvaluation tells you how the feature behaves. Guardrails constrain it in production:\n\n- **Grounding rules:** answer only from retrieved, authorized sources, or say you don't know.\n- **Output validation:** structured outputs checked against schemas and business rules before use.\n- **Allow-lists:** an AI can only reference entities that exist. It can't invent a product, a customer or a case number.\n- **Human checkpoints:** consequential outputs are drafts until a person accepts them.\n- **Untrusted-input handling:** document and user content is treated as data, never as instructions.\n- **Fallbacks:** if the model is unavailable or uncertain, the workflow continues deterministically.\n\n## Monitoring after launch\n\nIn production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\n\n## How the factory handles it\n\nIn our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn't close until the agreed evaluation thresholds are met. It's one of the [quality gates](\u002Fservices\u002Fai-production-sprint) we use to decide whether something is done.\n\nRelated: [from AI pilot to production application](\u002Fblog\u002Ffrom-ai-pilot-to-production-application) and [AI model governance as an application](\u002Fblog\u002Fai-model-governance-as-an-application).\n\nHave a pilot that's never been evaluated properly? [Bring it to us](\u002Fcontact).\n","\u003Cp>Traditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don&#39;t behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\u003C\u002Fp>\n\u003Cp>So AI features need their own form of testing, \u003Cstrong>evaluation\u003C\u002Fstrong>, and it has to be a delivery gate, not a one-off exercise before a demo.\u003C\u002Fp>\n\u003Ch2>Four layers of evaluation\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. Task quality.\u003C\u002Fstrong> Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Groundedness.\u003C\u002Fstrong> For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Safety and policy.\u003C\u002Fstrong> Does the feature refuse what it should: out-of-scope questions, requests for data the user can&#39;t access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Regression.\u003C\u002Fstrong> Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\u003C\u002Fp>\n\u003Ch2>Building the test sets\u003C\u002Fh2>\n\u003Cp>Good test sets come from the workflow, not from the vendor:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>real questions and documents from the pilot, anonymized where needed\u003C\u002Fli>\n\u003Cli>edge cases the business owner worries about\u003C\u002Fli>\n\u003Cli>known-hard cases collected from production feedback\u003C\u002Fli>\n\u003Cli>adversarial cases: injection attempts, ambiguous requests, missing data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every case has an expected outcome defined by a person who owns the domain.\u003C\u002Fp>\n\u003Ch2>Guardrails in the application\u003C\u002Fh2>\n\u003Cp>Evaluation tells you how the feature behaves. Guardrails constrain it in production:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Grounding rules:\u003C\u002Fstrong> answer only from retrieved, authorized sources, or say you don&#39;t know.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Output validation:\u003C\u002Fstrong> structured outputs checked against schemas and business rules before use.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Allow-lists:\u003C\u002Fstrong> an AI can only reference entities that exist. It can&#39;t invent a product, a customer or a case number.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Human checkpoints:\u003C\u002Fstrong> consequential outputs are drafts until a person accepts them.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Untrusted-input handling:\u003C\u002Fstrong> document and user content is treated as data, never as instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fallbacks:\u003C\u002Fstrong> if the model is unavailable or uncertain, the workflow continues deterministically.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Monitoring after launch\u003C\u002Fh2>\n\u003Cp>In production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\u003C\u002Fp>\n\u003Ch2>How the factory handles it\u003C\u002Fh2>\n\u003Cp>In our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn&#39;t close until the agreed evaluation thresholds are met. It&#39;s one of the \u003Ca href=\"\u002Fservices\u002Fai-production-sprint\">quality gates\u003C\u002Fa> we use to decide whether something is done.\u003C\u002Fp>\n\u003Cp>Related: \u003Ca href=\"\u002Fblog\u002Ffrom-ai-pilot-to-production-application\">from AI pilot to production application\u003C\u002Fa> and \u003Ca href=\"\u002Fblog\u002Fai-model-governance-as-an-application\">AI model governance as an application\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Have a pilot that&#39;s never been evaluated properly? \u003Ca href=\"\u002Fcontact\">Bring it to us\u003C\u002Fa>.\u003C\u002Fp>\n","Evaluation and guardrails: how to test AI features before production","AI features need evaluation as a delivery gate, just like tests: test sets, groundedness checks, safety checks and regression on every change.","ai-in-production",[15,34,13,35],"production","human-in-the-loop","2026-06-30T00:00:00.000Z",6,2,1790080513388]