Ceety Systems

Compliance AI

Governing AI agents with NIST AI RMF and ISO/IEC 42001

For AI and risk leaders: govern AI agents with NIST AI RMF, ISO/IEC 42001 and the EU AI Act, plus a practical control set from first agent to full program.

By the Ceety Systems teamUpdated 7 min read

Key takeaways

  • Agents act, not just answer, so governance has to cover tools, permissions and approvals as well as model output.
  • NIST AI RMF gives you the risk vocabulary; ISO/IEC 42001 gives you a management system an accredited body can certify.
  • As of September 2026, EU AI Act high-risk obligations apply from 2 December 2027 (Annex III) and 2 August 2028 (Annex I).
  • Five controls carry most of the weight: approval thresholds, least-privilege tools, evaluation, tracing and incident handling.
  • Start with one agent and a one-page register; the same controls grow into an enterprise program without rework.

To govern AI agents, treat each agent as a system that takes actions on your behalf, not as a chat window. Use the NIST AI Risk Management Framework to identify and measure the risks, ISO/IEC 42001 to run governance as a repeatable management system, and a small set of technical controls (human approval, least-privilege tools, evaluation, tracing and incident handling) to enforce it. The same approach works for a startup's first agent and an enterprise portfolio.

What makes AI agents different from chat?

A chat assistant produces text that a person reads and decides what to do with. An agent is given tools (APIs, databases, email, code execution, other systems) and a goal, then decides which tools to call and in what order. The person may only see the result.

That changes the risk in three ways:

  • Tools. Every tool is a permission. An agent that can read the CRM and send email can also leak CRM data by email if it is steered badly.
  • Autonomy. Agents run multi-step plans. An error in step two compounds through steps three to ten without a human reading each one.
  • Actions. Agents change things: they update records, issue refunds, merge code, book appointments. Some of those actions are hard or impossible to reverse.

Governance for chat is mostly about content: accuracy, harmful output, data in prompts. Governance for agents also has to answer who authorized this action, with what permission, and how you would know if it went wrong.

What does the NIST AI RMF ask you to do?

The NIST AI Risk Management Framework (AI RMF 1.0), released in January 2023, is voluntary and is not specific to any sector. It organizes AI risk work into four functions:

  1. Govern. Policies, roles, accountability and culture. NIST describes it as a cross-cutting function that runs through the other three.
  2. Map. Establish the context: what the system is for, who it affects, what could go wrong.
  3. Measure. Analyze, assess, benchmark and monitor the risks you mapped, using quantitative, qualitative or mixed methods.
  4. Manage. Prioritize and treat those risks, including plans to respond to, recover from and communicate about incidents.

For agents, "Map" is where teams often underinvest. Write down every tool the agent can call, what data each one exposes and what each action changes. That inventory drives everything else.

The Generative AI Profile (NIST AI 600-1)

In July 2024 NIST published the Generative AI Profile (NIST AI 600-1), a companion to AI RMF 1.0. It names 12 risks that generative AI creates or makes worse, including confabulation, data privacy, information security, human-AI configuration, and value chain and component integration. Those last two map directly onto agents: how much the human is in the loop, and how much you trust third-party models, tools and connectors. The profile also suggests actions for each risk, which makes it a useful starting checklist.

NIST has said AI RMF 1.0 is being revised, so check the framework page for the current version before you build your control mapping.

What is ISO/IEC 42001, and can you be certified?

ISO/IEC 42001, published in December 2023, specifies requirements for establishing, implementing, maintaining and continually improving an AI management system within an organization. If you know ISO/IEC 27001 for information security, the shape is familiar: scope, policy, risk assessment, controls, internal audit, management review and continual improvement.

The difference from NIST is purpose. The AI RMF tells you what to think about. ISO/IEC 42001 tells you how to run that thinking as an auditable system. And it is certifiable: ISO/IEC 42006, published in July 2025, sets the requirements for the bodies that audit and certify AI management systems against ISO/IEC 42001.

A practical reading: use the NIST functions and the Generative AI Profile to decide which risks and controls matter for your agents, then use ISO/IEC 42001 to prove to customers and auditors that you apply them consistently.

How does the EU AI Act relate?

The EU AI Act is law, not a voluntary framework, and it applies by risk category. According to the European Commission, as of September 2026:

  • It entered into force on 1 August 2024.
  • Prohibited practices and AI literacy obligations have applied since 2 February 2025.
  • Obligations for general-purpose AI models have applied since 2 August 2025.
  • The Act became generally applicable on 2 August 2026, with exceptions.
  • After the AI Omnibus amendment, which entered into force on 27 July 2026, rules for high-risk systems in the sensitive areas listed in Annex III apply from 2 December 2027, and for high-risk AI embedded in regulated products (Annex I) from 2 August 2028.

An agent is not high-risk because it is an agent. Its classification depends on what it is used for, such as employment decisions, credit or access to essential services. This is general information, not legal advice; your counsel should confirm how the Act applies to your use cases.

NIST AI RMF and ISO/IEC 42001 do not make you compliant with the Act on their own. They do give you the risk management, documentation, logging and human oversight practices the Act expects of high-risk systems, which makes them a sound operating layer underneath it.

A practical control set for AI agents

These five controls turn the frameworks into something an engineer can build and an auditor can test.

1. Human approval thresholds

Classify every action the agent can take by impact and reversibility:

  • Read-only (search, look up, summarize): no approval needed.
  • Reversible writes (draft an email, create a ticket): approval optional, logged.
  • Consequential or irreversible actions (send externally, pay, delete, change access, merge to production): a named human approves each one.

Put the threshold in code, not in the prompt. A model can be talked out of an instruction; it cannot talk its way past an approval gate in the tool layer.

2. Least-privilege tools

Give each agent its own identity and only the scopes it needs. Prefer narrow tools ("look up order status") over broad ones ("run any SQL"). Use short-lived tokens, separate read and write credentials, and keep production access out of development agents. If the agent connects through Model Context Protocol servers, allow-list the servers and review what each exposes.

3. Evaluation before and after release

Build a test set of real tasks, including adversarial ones: prompt injection in retrieved documents, requests to exceed scope, ambiguous instructions. Run it before every model, prompt or tool change. Keep evaluating in production with sampled reviews, because behavior drifts when models, data or tools change.

4. Tracing and logging

Record every step: the input, the plan, each tool call with its arguments and result, each approval and who gave it, and the final output. Tie traces to the agent's identity and the requesting user. This is your evidence for auditors and your only way to reconstruct what happened when something goes wrong. Treat traces as sensitive data, with retention and access rules to match.

5. Incident handling

Extend your existing incident process to agents. Define what counts as an AI incident (wrong action taken, data exposed, approval bypassed), who can switch an agent off, and how you roll back its actions. Test the kill switch. After an incident, feed the case back into the evaluation set so it cannot quietly recur.

How does this scale from a first agent to an enterprise program?

A startup with its first agent needs a one-page register (what the agent does, its tools, its data, its owner), approval gates on consequential actions, tracing switched on from day one, and a short evaluation set. That is enough to answer the AI questions in an enterprise customer's security review.

A growing company with several agents should standardize: one tool-permission model, one tracing pipeline, one evaluation harness, and a lightweight review before any new agent or tool goes live. Map those controls to the NIST functions so the next framework request is a mapping exercise, not a rebuild.

An enterprise runs AI governance as a management system. That means an AI policy, a risk assessment method, an inventory of every agent and model, internal audit, management review, and clear accountability across legal, security, data and the business. ISO/IEC 42001 certification is the natural endpoint if customers or regulators want independent assurance. EU AI Act classification belongs in the intake process for every new use case.

The controls do not change between stages. Only the formality around them does. That is the point of building them in from the first agent. If you want help putting this in place, see our compliance-first AI and automation practice or how we approach security and trust.

Frequently asked questions

Do we need both NIST AI RMF and ISO/IEC 42001?

Not necessarily. NIST AI RMF is a free, voluntary framework that works well for structuring risk decisions. ISO/IEC 42001 matters when you need a certificate or a formally auditable system, usually because customers or regulators ask for one.

Is an AI agent automatically high-risk under the EU AI Act?

No. Classification depends on the use case, not the technology. An agent that screens job applicants may fall into a high-risk category; one that drafts internal meeting notes generally will not. Confirm your specific case with legal counsel.

Where should human approval sit in an agent workflow?

At the tool layer, before any consequential or irreversible action runs. Approvals written only into the prompt can be bypassed by prompt injection or model error, while a gate in code cannot be argued with.

How long does ISO/IEC 42001 certification take?

It depends on your scope, how mature your existing management systems are and your certification body's schedule. Organizations that already run ISO/IEC 27001 can reuse much of that structure, which typically shortens the work.

What should we log for an AI agent?

Inputs, each tool call with arguments and results, approvals and approvers, outputs, and the identities of the agent and the requesting user. Keep enough to reconstruct any action after the fact, and protect the logs as sensitive data.

Tell us about your business.

Book a free consultation: a conversation about what you have and what you want. We’ll tell you honestly what you don’t need. Free, with no obligation.