Blog/ Buyer guides

Writing a Risk Register Entry for an AI Email Agent

Nafiul HasanNafiul Hasan· 14 min read
Abstract illustration of risk register rows for an AI email agent, each row pairing a failure mode with a control and a residual score

The short answer

Write one row per failure mode, not one row per vendor. Each entry names the risk, its cause, the business impact, an inherent likelihood-and-impact score, the specific control that reduces it, the residual score after that control, a named human owner, and a review date. Score the autonomy level, not the brand.

How to add AI tools to your risk register: six named risks for an AI email agent, with causes, controls, residual scores and a named owner.

On this page
  1. 01The short answer
  2. 02Score the autonomy level, not the vendor
  3. 03Criteria that actually matter in an AI entry
  4. 04How do I score AI risk likelihood and impact?
  5. 05What residual risk after controls really means
  6. 06A risk register example for an AI email tool
  7. 07The three that fail loudly
  8. 08The three that fail quietly
  9. 09A worked entry, written out in full
  10. 10Red flags that your entry is decorative
  11. 11Who owns AI risk in a company, and how often to review it
  12. 12What we would pick, and where we are the wrong answer

Most advice on how to add AI tools to your risk register stops at the vendor: read the security page, note the certification, score it medium, move on. That produces a row a committee can read but cannot act on.

An AI email agent does not fail as a vendor. It fails as a specific action, on a specific mailbox, at a specific autonomy setting. That is what the entry has to describe, and it is why one row per tool is almost always the wrong shape.

The short answer#

A register entry for an AI email agent is one row per failure mode, written so that someone who was not in the procurement meeting can challenge it. Four to six rows usually covers an email agent completely.

Every row carries the same nine fields. If a field is missing, the row is a note rather than a control.

  • Risk ID and a one-line risk statement, phrased as a thing that happens, not a technology
  • Cause: the specific mechanism, not 'AI is unpredictable'
  • Impact: what it costs the business, in business words
  • Inherent score: likelihood times impact, before any control
  • Control: the setting or process that reduces it, and where it is configured
  • Residual score: likelihood times impact assuming the control is operating
  • Evidence: how you would prove to an auditor that the control ran
  • Owner: one named person who can change the setting
  • Review date and the triggers that force an early re-score

Score the autonomy level, not the vendor#

This is the part generic AI risk guidance skips, and it is the part that decides whether the register stays true. The same product, on the same mailbox, has a different risk profile at 'suggests a draft' than at 'sends without asking'. Same vendor, same certification, same contract, different row.

So the entry must name the setting it was scored against. 'AI email assistant, Copilot mode, approval required on all sends' is a scoreable statement. 'AI email assistant' is not.

That has a second consequence people miss. If the score depends on a product toggle, then the owner of the row is whoever can flip that toggle, and a change to that toggle is a change to the register. Widening autonomy without a re-score is the single most common way an AI entry quietly becomes fiction.

One tool, several rows

If different teams run the same agent at different autonomy levels, that is not one entry with an average score. Support running on approval-required and sales running on autonomous send are two rows with two owners, because a control that exists in one place and not the other is not a control.

Criteria that actually matter in an AI entry#

Risk committees reject AI rows for predictable reasons. These are the criteria that separate an entry that survives review from one that gets sent back.

  • The risk statement is a failure, not a feature. 'Reply sent to the wrong recipient' beats 'use of generative AI'.
  • The cause is mechanical enough to argue with. An engineer or an ops lead should be able to say 'no, that is not how it goes wrong'.
  • The control is a setting, a queue or a scheduled task, with a named location. Policy sentences do not reduce residual risk on their own.
  • The residual score is honestly above zero. Every real control leaks.
  • The evidence field answers 'how would we show this ran last quarter?' before anyone asks.
  • The owner is a person, not a department. 'Security' has never once approved a draft.
  • The entry names its scope: which mailboxes, which teams, which autonomy mode.

How do I score AI risk likelihood and impact?#

Use whatever scale your register already uses. Most run one to five on each axis, multiplied to an inherent score out of twenty-five. The work is not the arithmetic; it is writing anchors specific enough that two people score the same row the same way.

These anchors are tuned for email volume rather than generic IT risk, which is why a once-a-month event scores as high as it does here.

ScoreLikelihood - how often you would expect itImpact - what it costs when it happens
1Not in the next two years at current volumesCaught internally, fixed in minutes, nobody outside notices
2About once a yearOne recipient inconvenienced; an apology closes it
3About once a quarterA client relationship needs repair; a manager loses a day to it
4About once a monthContractual or regulatory exposure; legal or the privacy owner is pulled in
5Weekly, or it is already happeningReportable breach, lost account, or an incident that goes public

What residual risk after controls really means#

Inherent score assumes no control at all: the agent is connected and nothing constrains it. Residual score assumes your control is in place and working as configured. The gap between them is the value of the control, and it is the number the committee is actually buying.

Score residual against the control as it is configured today, not as it is described in the vendor's marketing. If approval is required on sends but three people have permission to turn that off, the residual score should reflect three people, not zero.

Abstract balance weighing two values against each other, representing likelihood scored against impact to produce an inherent risk rating
Inherent score is likelihood times impact with nothing in the way. Residual is the same product after the control, and it is never zero.

A risk register example for an AI email tool#

Six rows cover an AI email agent for most organisations. The scores below are illustrative, using the one-to-five anchors above and assuming a mid-sized team with client-facing mail. Your own numbers will differ, and they should.

Abstract illustration of a magnifier over a stack of logged records, representing evidence that a risk control actually operated
Every row above needs an evidence field. A control nobody can show running is a control nobody can score.
RiskCauseImpactKey controlInherent to residualOwner
AI-01 Wrong recipientAgent resolves an ambiguous name, replies to a list address, or keeps a stale thread participant on the replyConfidential content reaches someone outside the intended party; possible notification dutyRecipient allowlist for any autonomous send; approval required for anything off the list; cancellable undo window16 to 6Head of IT
AI-02 Wrong contentDraft states a price, date or commitment the business cannot honour, phrased confidentlyContractual exposure, refunds, a client relationship to repairHuman approval on any first-party commitment; escalation keywords on price, contract and legal terms force review15 to 5Function head (sales or support)
AI-03 Prompt injectionInstructions hidden in an inbound message body are read by the agent as commandsData exfiltration, unauthorised forwarding, silent changes to filing rulesMessage bodies handled as data rather than instructions; fixed action allowlist; injection attempts logged and alerted12 to 4Security lead
AI-04 Over-automationAutonomy is widened faster than review capacity, and nobody reads the queue any moreErrors compound unnoticed; staff lose the habit of spotting themAutonomy changes go through change control; a sampled review of agent actions every week with a recorded result12 to 6Process owner
AI-05 Data exposureMail content leaves your control to a model provider, a subprocessor, or a logRegulatory exposure, breach of a client confidentiality clauseZero-retention terms in writing; encryption at rest; subprocessor list reviewed at every renewal12 to 4Privacy owner or DPO
AI-06 Vendor failureVendor is acquired, shuts down, or changes terms at renewalLoss of the workflow, a forced migration, stranded configurationMail stays in the underlying provider mailbox; documented exit path; configuration and audit log exported quarterly9 to 4Procurement or CIO

The three that fail loudly#

AI-01, AI-02 and AI-03 are the rows an executive already has an opinion about, which makes them easy to get on the register and easy to score badly.

Wrong recipient is scored high on inherent likelihood because email addressing is genuinely ambiguous and because reply-all exists. The control that moves it is not training. It is a constrained recipient set for anything the agent sends without a human, plus a delay you can cancel inside.

Wrong content is where most teams over-credit their control. A review step only reduces risk if the reviewer has the context to catch the error, which is why the useful control is narrower: force human review on the categories where a wrong answer is expensive, rather than claiming a person reads everything.

Prompt injection is the one to write carefully, because the naive version of this row is unfalsifiable. Ground it in a public taxonomy. Prompt injection is LLM01 in the OWASP Top 10 for Large Language Model Applications, with excessive agency and improper output handling nearby, and citing that gives your committee something to check the vendor against rather than a feeling.

The three that fail quietly#

AI-04, AI-05 and AI-06 are the rows that get dropped in the first draft and cause the trouble two years later.

Over-automation is a process risk, not a technology risk, and it is the only row on the list where the failure is caused by success. The agent performs well, trust grows, review becomes a formality, and the queue stops being read. Its control is a cadence with a recorded outcome, not a setting.

Data exposure is the row auditors read first. Score it against what the contract actually says about retention and training on your content, and record where you read it. A vendor page is evidence with a date on it; a salesperson's assurance is not.

Vendor failure is scored lower than people expect for one structural reason worth stating in the entry: with an AI client sitting on top of Gmail, Microsoft 365 or IMAP, the mail itself stays in the underlying mailbox. What you lose in a shutdown is the automation layer and its configuration, which is recoverable, rather than the archive, which would not be.

A worked entry, written out in full#

This is AI-03 expanded into every field, in the shape most registers accept. It is deliberately boring, which is the point: a committee should be able to challenge any single line of it.

AI-03 - Prompt injection via inbound email
Risk IDAI-03
Risk statementAn inbound email contains hidden instructions that the AI agent follows, causing it to take an action the sender chose
CategoryInformation security / third-party AI
ScopeAll mailboxes connected to the agent, Copilot and Autopilot modes
CauseMessage bodies are attacker-controlled text supplied to a model that also receives system instructions
ImpactData exfiltration, unauthorised forwarding, or altered filing rules; potential notifiable incident
Inherent likelihood3 - about once a quarter
Inherent impact4 - regulatory or contractual exposure
Inherent score12 (high)
ControlsBodies passed as data, not instructions; fixed action allowlist; injection classifier logs attempts; sends require approval outside the allowlist; remote content and tracking pixels blocked by default
Control ownerSecurity lead
Residual likelihood2 - about once a year
Residual impact2 - contained, one thread affected
Residual score4 (low)
Treatment decisionTreat and accept residual
EvidenceAudit log export showing flagged attempts and outcomes; vendor security page reviewed 2026-08-14
Risk ownerNamed security lead
Review dateQuarterly, next 2026-11-14
Early re-score triggersAutonomy widened; new mailbox scope; model or provider change; any flagged attempt that reached a send

Red flags that your entry is decorative#

These are the patterns that make an AI row fail its first serious review. Each one is fixable in a sentence.

  • The risk names a vendor instead of a failure. Vendors do not appear in incidents; actions do.
  • Residual risk is zero. That claims a control with no failure mode, which no control has.
  • The control is a policy line such as 'staff are trained to check AI output'. Write the setting or the scheduled task instead.
  • The owner is a team. Ask who can change the autonomy setting, and put that name in the field.
  • The score has not moved since procurement, although the autonomy setting has.
  • The impact is written in AI vocabulary. 'The model hallucinates' is a cause; the impact is the invoice you had to honour.
  • There is no evidence field, so nothing in the row can be tested at audit.
  • One row covers every team, at every autonomy level, on every mailbox.

A register that never changes is not a control

If your AI rows have carried the same scores for four consecutive quarters while usage grew, the review is a signature exercise. Add trigger-based re-scoring so the entry moves when the configuration moves, not only when the calendar does.

Who owns AI risk in a company, and how often to review it#

Ownership splits cleanly if you resist the urge to hand the whole thing to security. The business function whose work the agent does owns the output risks, because they are the only ones who can judge whether a draft was wrong. Security owns injection and exposure. The privacy owner or DPO owns the data terms. Procurement owns vendor continuity.

Above those sits one accountable executive for the AI inventory itself, which is the role structure ISO/IEC 42001 formalises for an AI management system. NIST's AI Risk Management Framework, which is voluntary, organises the same work under Govern, Map, Measure and Manage, and mapping your rows to those functions is an easy way to show the register is complete rather than opportunistic.

Quarterly review is a reasonable default for a stable configuration. It is also insufficient on its own, because the risk changes when the settings change, not when the quarter ends.

  1. 1

    Re-score on any autonomy change

    Moving a team from approval-required to autonomous send changes likelihood on at least three rows. Make it a change-control item, not a settings tweak.

  2. 2

    Re-score on scope change

    A new mailbox, a shared inbox, or a new team connected to the agent widens blast radius. Note the new scope in the entry.

  3. 3

    Re-score after any incident or near miss

    Even a caught error is evidence that your likelihood anchor was wrong. Adjust it while the detail is fresh.

  4. 4

    Re-score on model, provider or terms change

    A changed subprocessor list or retention clause moves the data-exposure row directly. Record the date you read the terms.

  5. 5

    Confirm evidence quarterly

    Pull the audit log export and the sampled-review record. If neither exists for the quarter, the control did not operate and the residual score is wrong.

What we would pick, and where we are the wrong answer#

If the controlling variable in your register is autonomy, then choose the tool that makes autonomy an explicit, per-rule setting with an exportable record, because that is the only version where your control has evidence behind it. A tool that quietly decides how much to do on your behalf gives you a row you cannot score.

AI Emaily is built around that shape, and we build it, so read this as an interested recommendation. Authority is three named modes: Manual, Copilot and Autopilot. Copilot holds every send for approval, so the wrong-recipient and wrong-content rows are controlled by design rather than by policy. Autopilot is gated behind a sender allowlist, a confidence floor you set, a cancellable undo window, escalation keywords, optional working hours and a global pause.

Message bodies are treated as untrusted input against a fixed action allowlist, which is what the prompt-injection row needs. The audit log is append-only, records the reasoning behind each action, and exports as JSON or CSV, which is what turns 'we have a control' into a line an auditor can check. That fits an operator or a small client-facing team that needs each register row to be provable.

Where we are the wrong answer: if the requirement behind the entry is tenant-wide retention, legal hold and eDiscovery across the whole organisation, a mail client is not the control layer. Microsoft Purview for Microsoft 365 and Google Vault for Workspace do that at tenant level, and we do not compete with them; check their current capabilities on their own documentation before you write the control. Two other limits worth knowing if your register includes endpoint standardisation: we have no native Linux build, and Android is a progressive web app rather than a native one.

Frequently asked

Nafiul Hasan

Written by

Nafiul Hasan

Nafiul Hasan is an entrepreneur and AI automation system builder with 10+ years of experience turning messy, manual workflows into reliable automated systems. He designs and ships AI enterprise solutions end-to-end — the agent logic, the data plumbing, and the product people actually use — and founded AI Emaily to give busy professionals their attention back. He writes here from the builder's seat: what works, what breaks, and how to put AI to work without giving up control.

EntrepreneurAI Automation System BuilderAI EnthusiastBuilds AI Enterprise Solutions10+ years experience
More from Nafiul
Ready when you are

Controls you can evidence, not just describe

Named authority modes, approve-before-send, a cancellable undo window and an append-only audit log you can export. See how AI Emaily makes each register row provable on a 7-day free trial.

  • 7-day free trial
  • Cancel anytime
  • Every provider