← Back to Resources
What is agentic AI in compliance and how do you keep AI agents auditable?
Karthigeyan R J
Agentic AI is software that plans and executes multi-step tasks using tools, rather than answering single prompts. In compliance that means an agent can monitor a source, assess a change and draft the register update on its own. UK regulators apply existing frameworks (SYSC, SS1/23, the Consumer Duty, SM&CR), so auditability is the firm's problem: log every step, ground every claim in rule text and evaluate against a golden set.
Three things this article will leave you able to do
Explain agent, copilot and chatbot to your board in one slide.
Name the six controls that make an agent's work auditable.
Ask a vendor the evaluation questions that separate demos from systems.
"Agentic AI" is the most searched and least defined term in compliance technology this year. Every vendor claims it. Almost nobody says what it is, which suits vendors and fails buyers, because the difference between an agent and a chatbot with a new label is precisely the difference that determines whether your regulator will accept its output.
Here is the working definition, then the part that matters: what it takes to let one loose in a regulated firm.
The definitions, plainly
A chatbot answers a prompt from its training. Ask it a question, get a paragraph. It holds no tools and takes no actions.
A copilot assists a human inside a task: drafting while you review, suggesting while you decide. The human stays in the loop for every step.
An agent is given a goal, not a prompt. It plans the steps, calls tools (a search index, a rules database, a document store, an email system), reads the results, decides what to do next and produces an outcome. The human sets the goal and reviews the output; the steps in between are the agent's.
Two supporting terms. MCP (Model Context Protocol) is an open standard for connecting AI systems to tools and data sources; it is how an agent reaches your obligations register without bespoke plumbing. A regulatory taxonomy or ontology is a structured map of the rulebook (sourcebooks, chapters, rules, the relationships between them) that lets an agent cite SYSC 6.1.1R as a stable reference rather than a paraphrase. Agents without a taxonomy underneath produce fluent text about rules; agents with one produce claims you can check.
In compliance terms: a chatbot answers "what is CASS 15?". A copilot helps you draft the CASS 15 gap analysis. An agent watches the FCA's publications, notices the safeguarding consultation, maps it to your permissions, drafts the register delta and queues it for your sign-off.
Who this applies to
Any regulated firm deploying AI that takes multi-step actions, plus any compliance team evaluating vendors that claim it. The UK has no AI rulebook to point at: the FCA and PRA have stayed principles-based and technology-neutral, running AI Live Testing cohorts (the first from October 2025, the second with applications from 19 January to 24 March 2026) rather than writing rules. That is not a gap in your obligations. It means the existing ones apply: SYSC governance, the PRA's model risk principles in SS1/23, the Consumer Duty where retail outcomes are touched and SM&CR accountability for whichever senior manager owns the system. The Treasury Committee has pressed for guidance on how SM&CR maps to autonomous systems; until it arrives, the safe assumption is that a named human answers for every agent action.
The six controls that make an agent auditable
A decision log, per step. Every tool call, its inputs, its outputs and the agent's stated reason, timestamped and immutable. Not a chat transcript: a flight recorder. If the agent read a rule, the log shows which version, retrieved when.
Grounding in a taxonomy. Claims cite rule identifiers, not paraphrases. "CASS 15.4 requires daily reconciliation" with a link to the provision beats three fluent sentences about safeguarding. This is what makes review possible at all.
Evaluation against a golden set. A maintained test set of questions and tasks with verified answers, run before deployment and rerun on every model or prompt change. Track the pass rate; treat a drop as an incident. A vendor who cannot show you their evaluation results is showing you a demo.
Determinism boundaries. Defined actions the agent may take alone (draft, flag, classify), actions needing human sign-off (register changes, external communications) and actions it may never take. Written down, enforced in the system, not in the hope.
Version pinning and change control. The model, the prompts and the tools are versioned; changes go through the same change management as any other system affecting compliance. An agent that silently improved is also an agent that silently changed.
Data residency and access control. The agent sees what its task needs, in the jurisdiction the data must stay in. Smaller domain-tuned models help here: they can run inside your boundary rather than shipping rule queries to a general model elsewhere. (Why smaller models suit regulated firms is its own article, coming in this series.)
What you must evidence
| Requirement | Applied to an AI agent | What evidence proves it |
| SYSC 4/6 governance | The agent is a system within the control framework | System inventory entry; risk assessment; the six controls above, documented |
| SS1/23 model risk (PRA firms) | The agent is a model with lifecycle management | Validation reports; the golden set and its results over time; change records |
| SM&CR accountability | A named SMF answers for the agent's actions | Statement of Responsibilities covering the system; sign-off records at the determinism boundary |
| Consumer Duty (retail impact) | Agent outputs affecting customers deliver good outcomes | Outcome monitoring on agent-touched journeys; testing of customer-facing text |
| SYSC 9 records | The agent's work is retrievable | The decision log itself, exportable, retained on schedule |
How firms handle this
Most firms are running pilots with general-purpose models and no evaluation harness, which produces impressive demos and unreviewable output. A smaller group has learnt from model risk management and treats agents like models: validated before use, monitored in use, logged always. That is the group regulators will be comfortable with. We build our own compliance agents this way (every step logged against an FCA Handbook taxonomy, evaluated on a maintained golden set) because we expect to be asked to prove it. So should any vendor you buy from.
Primary sources
FCA, AI Lab and AI Live Testing. First cohort October 2025; second cohort applications 19 January to 24 March 2026.
PRA, SS1/23: Model risk management principles for banks. May 2023.
Volatile. Re-verify before each republish: any published FCA or PRA guidance on SM&CR and autonomous systems (called for by end of 2026) · AI Live Testing cohort outcomes · the Mills Review of AI in retail financial services (launched 27 January 2026) · MCP adoption and versioning.