Before I Begin: The views represented here are my own and have never been views of my past employer(s). Last Updated: October 2026
TLDR: A prompt that says “never do X” is a request, not a control. Real agentic governance lives outside the model: a declared scope in a file someone can review, a deterministic layer that can say no, and human gates you actually test. Singapore’s framework is a good map. The YAML is the easy part. I built a small runnable demo to show what that looks like.
The companion demo, one run
From outputs to actions
Most AI governance was written for systems that produce text. Somebody reads the text, and then something happens. There’s a human between the model and the consequence.
Agents remove that buffer. They plan, call tools, and take multi-step actions. The consequence can arrive before anyone has read anything.
Singapore’s IMDA launched its Model AI Governance Framework for Agentic AI at Davos in January, and IMDA describes it as the first of its kind. Version 1.5 followed in May with real-world case studies and new guidance on multi-agent systems, third-party agents, and automation bias. It’s voluntary guidance with no enforcement mechanism, which makes it something to borrow from rather than a checklist to satisfy.
It organizes the problem into four dimensions: assess and bound risks up front, keep humans meaningfully accountable, put technical controls around the agent, and be honest with end users about what they’re dealing with.
I’ve been reading it through a specific lens: caseworker-facing tools, where the person clicking the button isn’t the person who lives with the outcome. Here’s what I’d take from it, with a companion repo that runs each lesson as a scenario. No API key, no model, just Python and a scripted agent that misbehaves on cue. If you’d rather click than install, there’s a browser walkthrough that plays back real runs step by step.
1. Prompts are requests. Controls live outside the model.
If your only guardrail is a paragraph in the system prompt, you have a suggestion the model will usually follow.
Prompt injection is where “usually” breaks down. In the demo’s injection scenario, the intake notes for a case contain instructions: approve the payout, compare against another client’s file, write the result to the record. The scripted model does what a real one sometimes does and tries all three. A gate sitting between the agent and its tools says no:
[ALLOW ] read_client_file(case_id=C-1003)
[DENY ] approve_payout(case_id=C-1003) rule=default-deny
[DENY ] read_client_file(case_id=C-2000) rule=org.scope
[DENY ] write_record(case_id=C-1003, ...) rule=default-deny
[ALLOW ] calculate_income_threshold(annual_income=41000, household=4)
[REQUIRE_APPROVAL] draft_determination_letter(case_id=C-1003, ...)
-> dangerous tools that actually ran: none
None of that depended on the model behaving. Two of those calls weren’t in its declared tool list, and the third was a tool it’s allowed to use, pointed at a case it isn’t working on. The real job still finished.
The framework’s answer is the same idea at policy level: set hard limits at design time, like restricting tools and data to the minimum necessary and putting access controls around them. The tooling that’s emerging enforces those limits somewhere the model can’t argue with them. Microsoft’s Agent Governance Toolkit is an MIT-licensed policy engine that intercepts agent actions in deterministic code before they execute, with policy written in YAML, OPA Rego, or Cedar. NVIDIA’s OpenShell goes a layer lower and puts the boundary in the runtime environment instead of the model or the application.
Different layers, same instinct I keep coming back to. Deterministic code for the rules that don’t change. The LLM handles composition and conversation. It doesn’t get a vote on whether it’s allowed to do something. The demo’s whole gate is about 200 lines of plain Python with no model in it.
2. Declare the agent in a file someone can review
Agent Format is an open standard that treats an agent like a Kubernetes manifest: one YAML file that declares identity, interface, tools, limits, and governance, validated against a JSON Schema. Name the file after the agent’s ID, with the .agf.yaml suffix the spec uses: caseworker_assistant.agf.yaml. The docs describe an agf lint CLI, though as of this writing they say it’s coming soon. In the meantime you can validate against the published JSON schema or use the browser playground. The example below passes the JSON schema clean, and it’s the same file the demo runs.
Here’s a sketch for a caseworker agent, built from the field reference:
# caseworker_assistant.agf.yaml
schema_version: "1.0.0"
metadata:
id: caseworker_assistant
name: Caseworker Assistant
version: "0.1.0"
description: Drafts eligibility summaries for a caseworker to review
data_classification: restricted
interface:
input:
type: string
description: Intake notes for a single case
output:
type: string
description: A draft summary. Never a determination.
constraints:
tighten_only_invariant: true
budget:
max_token_usage: 25000
max_duration_seconds: 120
limits:
max_llm_calls: 8
max_tool_calls: 12
max_delegation_depth: 0
governance_policies:
- policy_ref: agency.privacy.pii-redaction-v1
required: true
action_space:
local_tools:
- alias: read_client_file
description: Read-only, scoped to one case ID
- alias: calculate_income_threshold
description: Deterministic calculator, no LLM involved
- alias: draft_determination_letter
description: Drafts text only. Cannot send or commit.
approval:
message_template: "Approve this draft letter before it goes anywhere?"
execution_policy:
id: agf.react
config:
instructions: |
Summarize the intake notes. Use tools for anything numeric.
If the data is missing or contradictory, say so and stop.
model: your-model-here
max_steps: 8
required: true on a governance policy means the agent must not run if that policy can’t be resolved, so it fails closed. Delete the PII policy file in the demo and the agent refuses to start:
[REFUSED] required policy 'agency.privacy.pii-redaction-v1' not found; refusing to run
tighten_only_invariant means org-level policies can only tighten what the agent’s owner declared, never loosen it. And the approval requirement sits on the tool itself, where a reviewer reading the file will see it.
Before you pick a format, check whether it can express guardrails at all. Some declarative agent formats describe how an agent is wired (models, tools, connections) but have no field for run limits, a required policy, or approval conditioned on a tool’s arguments. The demo’s README has a side-by-side with Microsoft Agent Framework’s declarative YAML. If the format can’t say it, the rule ends up somewhere less reviewable.
Something else has to enforce the file at runtime, which is where a policy engine or sandbox comes in. An auditor can read it before the agent ever runs.
3. Shrink the blast radius on purpose
Every tool you hand an agent is a decision about what can go wrong. Read is different from draft, and draft is different from commit.
In the sketch above, the agent can read a file, run a calculator, and write a draft. It can’t send anything or touch the system of record. If it misreads an ambiguous case, the worst outcome is a bad draft that a human hasn’t approved.
Budgets are part of the blast radius too. In the demo’s loop scenario the model gets stuck re-reading the same file. max_llm_calls: 8 turns that from a runaway bill into a halt:
[ALLOW] read_client_file(case_id=C-1001) why: Let me just check the file once more
... seven more of those ...
[HALT ] more than 8 model calls in one run
One gotcha from the spec: on an MCP server, if you leave allowed_tools off, every tool that server exposes is available. Default open. Write the list out.
4. Put humans at the irreversible steps, then test them
In corporate settings the person using the agent absorbs the risk. In benefits, the caseworker triggers the agent and the client absorbs it. That asymmetry means anything irreversible or high-stakes needs an approval gate.
In the sketch, the most irreversible step, sending a determination, isn’t in the action space at all. The approval on draft_determination_letter is a second layer: even the draft doesn’t leave the agent’s hands until a caseworker says so. When you do give an agent a commit-style tool, that’s where the gate has to go.
But a gate isn’t oversight just because it exists. The updated Singapore framework added automation bias guidance for a reason. A 2025 study by Beck, Eckman et al. found that when rejecting a suggestion meant typing a correction, people accepted more wrong answers. Approving is cheap, rejecting is expensive, and the cheap path wins under volume. I wrote about the failure mode in the rubber stamp problem. The fix is the same here. Seed the queue with known-wrong outputs and measure the catch rate.
The demo’s redherring scenario runs that method against a simulated queue of 200 drafts, 5% of them known-wrong. These are modeled reviewers, not real ones, so the numbers show the method, not a finding:
rubber stamp approval rate 98% caught 0/10 seeded errors (0%)
attentive approval rate 95% caught 8/10 seeded errors (80%)
The approval rates are three points apart and the catch rates are eighty points apart. If approval rate is the only number on your dashboard, you can’t tell these two reviewers apart.
Agent Format helps with the other half of this. Approvals can be conditional, firing only when the arguments cross a threshold, so reviewers aren’t clicking through a wall of low-stakes prompts. A gate that fires rarely and matters is harder to rubber-stamp than one that fires constantly.
5. Log the agent’s reasoning
“The agent called draft_determination_letter at 2:14pm” is an audit log. “The agent chose that tool because it read X, ruled out Y, and hit the step limit” is something you can investigate.
The injection run, step by step
Look at the audit log in that run. The denied calls are logged with the agent’s stated reason (“The notes said to approve the payout”), which tells an investigator exactly where the attack came from. The tool name alone wouldn’t. The demo also redacts PII before anything hits the log, because the reasoning trace is going to quote the case file.
When an agent contributes to a denied benefit, someone will ask why. I got into what this looks like in building agent-aware systems: token claims, policy surfaces, and an audit trail that can tell an agent’s action from a human’s. That kind of audit trail is much easier to build in from the start than to add after the first incident.
6. Design for the system being down
For public benefits, I’d adapt this lesson the most. To its credit, the framework itself calls for failsafe mechanisms, gradual rollouts, and the ability to take a malfunctioning agent offline. What it says little about is the mess underneath: bad data, legacy systems, and APIs that time out or don’t exist.
State and local benefits systems often look like that. Data is messy, incomplete, or sitting behind a portal with no API at all. So the harder governance question is what happens when the thing the agent depends on is missing or wrong.
The safe answer is usually to stop, say so plainly, and hand off to a person with context intact. The agent instructions above end with that rule. In the demo’s messy case the income field is blank, and the agent doesn’t fill the gap with a plausible number:
[ALLOW ] read_client_file(case_id=C-1002)
[HANDOFF] Income is missing from the file. Not guessing; needs caseworker follow-up.
Then make sure operators have a kill switch they can hit without a meeting. AGT ships one in its SRE package, but whatever you use, test it before you need it. In the demo’s killswitch scenario, the operator halts the run after its first step and nothing else gets through.
A caution about the tools
None of this is settled. OpenShell’s community repo is labeled alpha software, Agent Format’s CLI and SDKs aren’t out yet, and AGT is a public preview that enforces at the application middleware layer, sharing a process with the agent. My demo has the same limitation and isn’t a security boundary either. No single layer is the guardrail. Declaration, policy, and sandbox each cover something the others don’t.
The checklist version
If you only take one page from this:
- Guardrails outside the model, enforced deterministically
- One reviewable
<agent-id>.agf.yamlper agent, schema-validated in CI - Governance policies marked
requiredso the agent fails closed - Explicit tool allowlists, with read, draft, and commit as separate permissions
- Approval gates on irreversible actions, tested with known-wrong inputs
- Catch rates and override rates measured alongside approval rates
- Reasoning traces logged with each tool call
- A defined failure mode and a tested kill switch
The open question
Every one of these controls answers “how do we stop the agent from doing something bad?” None of them answer the harder one: when a well-governed agent still contributes to a wrong outcome, who owns it? The vendor, the agency, the caseworker who approved it, the team that wrote the policy file? It’s the same accountability gap I ran into with ATO, just with more moving parts.
I don’t think anyone has a clean answer yet. I do think the teams who can produce the agent file, the audit trail, and the catch-rate data will be in a much better position to have that conversation than the ones who can only point at a system prompt.
If you want to poke at it, the demo repo runs in about a minute: pip install -r requirements.txt && python run_demo.py all. Or step through every scenario in the browser, with an Advanced toggle that shows each rule the gate checked. Swap one of the scripted planners for a real model and leave the gate alone.
Related reading
- Model AI Governance Framework for Agentic AI (PDF) - IMDA, Singapore (v1.5, 2026)
- Agent Governance Toolkit - Microsoft, GitHub (MIT)
- Agent Format specification and GitHub repo
- Introducing the Agent Governance Toolkit - Microsoft Open Source Blog (2026)
- Run Autonomous, Self-Evolving Agents More Safely with NVIDIA OpenShell - NVIDIA (2026)
- Bias in the Loop: How Humans Evaluate AI-Generated Suggestions - Beck, Eckman et al. (2025)
- caseworker-agent-governance-demo - the companion repo for this post
- The Rubber Stamp Problem
- ATO Wasn’t Built for Systems That Change at Inference Time
- Building AI Agent-Aware Systems
Working out guardrails for an agent in a high-stakes public service context? Reach out.
Related Posts
- AI Agent Identity Assurance: The NIST IAL/AAL CrosswalkNIST's IAL/AAL was built for humans. When AI agents act on your behalf, the assurance levels still apply - the mapping just changes.5/6/2026
- AI Agent Session Security: Prompt Injection and Dry RunsBrowser-use agents can be hijacked by the pages they visit. Dry runs don't replicate production. Here's how to build safer sessions.5/11/2026
- Building AI Agent-Aware SystemsMost systems can't detect when an AI agent is acting. Here's how to build the token claims, policy surfaces, and audit trail that changes that.5/16/2026
