Writing a Risk Register Entry for an AI Email Agent

The short answer
Write one row per failure mode, not one row per vendor. Each entry names the risk, its cause, the business impact, an inherent likelihood-and-impact score, the specific control that reduces it, the residual score after that control, a named human owner, and a review date. Score the autonomy level, not the brand.
How to add AI tools to your risk register: six named risks for an AI email agent, with causes, controls, residual scores and a named owner.
On this page
- 01The short answer
- 02Score the autonomy level, not the vendor
- 03Criteria that actually matter in an AI entry
- 04How do I score AI risk likelihood and impact?
- 05What residual risk after controls really means
- 06A risk register example for an AI email tool
- 07The three that fail loudly
- 08The three that fail quietly
- 09A worked entry, written out in full
- 10Red flags that your entry is decorative
- 11Who owns AI risk in a company, and how often to review it
- 12What we would pick, and where we are the wrong answer
Most advice on how to add AI tools to your risk register stops at the vendor: read the security page, note the certification, score it medium, move on. That produces a row a committee can read but cannot act on.
An AI email agent does not fail as a vendor. It fails as a specific action, on a specific mailbox, at a specific autonomy setting. That is what the entry has to describe, and it is why one row per tool is almost always the wrong shape.
The short answer#
A register entry for an AI email agent is one row per failure mode, written so that someone who was not in the procurement meeting can challenge it. Four to six rows usually covers an email agent completely.
Every row carries the same nine fields. If a field is missing, the row is a note rather than a control.
- Risk ID and a one-line risk statement, phrased as a thing that happens, not a technology
- Cause: the specific mechanism, not 'AI is unpredictable'
- Impact: what it costs the business, in business words
- Inherent score: likelihood times impact, before any control
- Control: the setting or process that reduces it, and where it is configured
- Residual score: likelihood times impact assuming the control is operating
- Evidence: how you would prove to an auditor that the control ran
- Owner: one named person who can change the setting
- Review date and the triggers that force an early re-score
Score the autonomy level, not the vendor#
This is the part generic AI risk guidance skips, and it is the part that decides whether the register stays true. The same product, on the same mailbox, has a different risk profile at 'suggests a draft' than at 'sends without asking'. Same vendor, same certification, same contract, different row.
So the entry must name the setting it was scored against. 'AI email assistant, Copilot mode, approval required on all sends' is a scoreable statement. 'AI email assistant' is not.
That has a second consequence people miss. If the score depends on a product toggle, then the owner of the row is whoever can flip that toggle, and a change to that toggle is a change to the register. Widening autonomy without a re-score is the single most common way an AI entry quietly becomes fiction.
One tool, several rows
Criteria that actually matter in an AI entry#
Risk committees reject AI rows for predictable reasons. These are the criteria that separate an entry that survives review from one that gets sent back.
- The risk statement is a failure, not a feature. 'Reply sent to the wrong recipient' beats 'use of generative AI'.
- The cause is mechanical enough to argue with. An engineer or an ops lead should be able to say 'no, that is not how it goes wrong'.
- The control is a setting, a queue or a scheduled task, with a named location. Policy sentences do not reduce residual risk on their own.
- The residual score is honestly above zero. Every real control leaks.
- The evidence field answers 'how would we show this ran last quarter?' before anyone asks.
- The owner is a person, not a department. 'Security' has never once approved a draft.
- The entry names its scope: which mailboxes, which teams, which autonomy mode.
How do I score AI risk likelihood and impact?#
Use whatever scale your register already uses. Most run one to five on each axis, multiplied to an inherent score out of twenty-five. The work is not the arithmetic; it is writing anchors specific enough that two people score the same row the same way.
These anchors are tuned for email volume rather than generic IT risk, which is why a once-a-month event scores as high as it does here.
| Score | Likelihood - how often you would expect it | Impact - what it costs when it happens |
|---|---|---|
| 1 | Not in the next two years at current volumes | Caught internally, fixed in minutes, nobody outside notices |
| 2 | About once a year | One recipient inconvenienced; an apology closes it |
| 3 | About once a quarter | A client relationship needs repair; a manager loses a day to it |
| 4 | About once a month | Contractual or regulatory exposure; legal or the privacy owner is pulled in |
| 5 | Weekly, or it is already happening | Reportable breach, lost account, or an incident that goes public |
What residual risk after controls really means#
Inherent score assumes no control at all: the agent is connected and nothing constrains it. Residual score assumes your control is in place and working as configured. The gap between them is the value of the control, and it is the number the committee is actually buying.
Score residual against the control as it is configured today, not as it is described in the vendor's marketing. If approval is required on sends but three people have permission to turn that off, the residual score should reflect three people, not zero.

A risk register example for an AI email tool#
Six rows cover an AI email agent for most organisations. The scores below are illustrative, using the one-to-five anchors above and assuming a mid-sized team with client-facing mail. Your own numbers will differ, and they should.

| Risk | Cause | Impact | Key control | Inherent to residual | Owner |
|---|---|---|---|---|---|
| AI-01 Wrong recipient | Agent resolves an ambiguous name, replies to a list address, or keeps a stale thread participant on the reply | Confidential content reaches someone outside the intended party; possible notification duty | Recipient allowlist for any autonomous send; approval required for anything off the list; cancellable undo window | 16 to 6 | Head of IT |
| AI-02 Wrong content | Draft states a price, date or commitment the business cannot honour, phrased confidently | Contractual exposure, refunds, a client relationship to repair | Human approval on any first-party commitment; escalation keywords on price, contract and legal terms force review | 15 to 5 | Function head (sales or support) |
| AI-03 Prompt injection | Instructions hidden in an inbound message body are read by the agent as commands | Data exfiltration, unauthorised forwarding, silent changes to filing rules | Message bodies handled as data rather than instructions; fixed action allowlist; injection attempts logged and alerted | 12 to 4 | Security lead |
| AI-04 Over-automation | Autonomy is widened faster than review capacity, and nobody reads the queue any more | Errors compound unnoticed; staff lose the habit of spotting them | Autonomy changes go through change control; a sampled review of agent actions every week with a recorded result | 12 to 6 | Process owner |
| AI-05 Data exposure | Mail content leaves your control to a model provider, a subprocessor, or a log | Regulatory exposure, breach of a client confidentiality clause | Zero-retention terms in writing; encryption at rest; subprocessor list reviewed at every renewal | 12 to 4 | Privacy owner or DPO |
| AI-06 Vendor failure | Vendor is acquired, shuts down, or changes terms at renewal | Loss of the workflow, a forced migration, stranded configuration | Mail stays in the underlying provider mailbox; documented exit path; configuration and audit log exported quarterly | 9 to 4 | Procurement or CIO |
The three that fail loudly#
AI-01, AI-02 and AI-03 are the rows an executive already has an opinion about, which makes them easy to get on the register and easy to score badly.
Wrong recipient is scored high on inherent likelihood because email addressing is genuinely ambiguous and because reply-all exists. The control that moves it is not training. It is a constrained recipient set for anything the agent sends without a human, plus a delay you can cancel inside.
Wrong content is where most teams over-credit their control. A review step only reduces risk if the reviewer has the context to catch the error, which is why the useful control is narrower: force human review on the categories where a wrong answer is expensive, rather than claiming a person reads everything.
Prompt injection is the one to write carefully, because the naive version of this row is unfalsifiable. Ground it in a public taxonomy. Prompt injection is LLM01 in the OWASP Top 10 for Large Language Model Applications, with excessive agency and improper output handling nearby, and citing that gives your committee something to check the vendor against rather than a feeling.
The three that fail quietly#
AI-04, AI-05 and AI-06 are the rows that get dropped in the first draft and cause the trouble two years later.
Over-automation is a process risk, not a technology risk, and it is the only row on the list where the failure is caused by success. The agent performs well, trust grows, review becomes a formality, and the queue stops being read. Its control is a cadence with a recorded outcome, not a setting.
Data exposure is the row auditors read first. Score it against what the contract actually says about retention and training on your content, and record where you read it. A vendor page is evidence with a date on it; a salesperson's assurance is not.
Vendor failure is scored lower than people expect for one structural reason worth stating in the entry: with an AI client sitting on top of Gmail, Microsoft 365 or IMAP, the mail itself stays in the underlying mailbox. What you lose in a shutdown is the automation layer and its configuration, which is recoverable, rather than the archive, which would not be.
A worked entry, written out in full#
This is AI-03 expanded into every field, in the shape most registers accept. It is deliberately boring, which is the point: a committee should be able to challenge any single line of it.
Red flags that your entry is decorative#
These are the patterns that make an AI row fail its first serious review. Each one is fixable in a sentence.
- The risk names a vendor instead of a failure. Vendors do not appear in incidents; actions do.
- Residual risk is zero. That claims a control with no failure mode, which no control has.
- The control is a policy line such as 'staff are trained to check AI output'. Write the setting or the scheduled task instead.
- The owner is a team. Ask who can change the autonomy setting, and put that name in the field.
- The score has not moved since procurement, although the autonomy setting has.
- The impact is written in AI vocabulary. 'The model hallucinates' is a cause; the impact is the invoice you had to honour.
- There is no evidence field, so nothing in the row can be tested at audit.
- One row covers every team, at every autonomy level, on every mailbox.
A register that never changes is not a control
Who owns AI risk in a company, and how often to review it#
Ownership splits cleanly if you resist the urge to hand the whole thing to security. The business function whose work the agent does owns the output risks, because they are the only ones who can judge whether a draft was wrong. Security owns injection and exposure. The privacy owner or DPO owns the data terms. Procurement owns vendor continuity.
Above those sits one accountable executive for the AI inventory itself, which is the role structure ISO/IEC 42001 formalises for an AI management system. NIST's AI Risk Management Framework, which is voluntary, organises the same work under Govern, Map, Measure and Manage, and mapping your rows to those functions is an easy way to show the register is complete rather than opportunistic.
Quarterly review is a reasonable default for a stable configuration. It is also insufficient on its own, because the risk changes when the settings change, not when the quarter ends.
- 1
Re-score on any autonomy change
Moving a team from approval-required to autonomous send changes likelihood on at least three rows. Make it a change-control item, not a settings tweak.
- 2
Re-score on scope change
A new mailbox, a shared inbox, or a new team connected to the agent widens blast radius. Note the new scope in the entry.
- 3
Re-score after any incident or near miss
Even a caught error is evidence that your likelihood anchor was wrong. Adjust it while the detail is fresh.
- 4
Re-score on model, provider or terms change
A changed subprocessor list or retention clause moves the data-exposure row directly. Record the date you read the terms.
- 5
Confirm evidence quarterly
Pull the audit log export and the sampled-review record. If neither exists for the quarter, the control did not operate and the residual score is wrong.
What we would pick, and where we are the wrong answer#
If the controlling variable in your register is autonomy, then choose the tool that makes autonomy an explicit, per-rule setting with an exportable record, because that is the only version where your control has evidence behind it. A tool that quietly decides how much to do on your behalf gives you a row you cannot score.
AI Emaily is built around that shape, and we build it, so read this as an interested recommendation. Authority is three named modes: Manual, Copilot and Autopilot. Copilot holds every send for approval, so the wrong-recipient and wrong-content rows are controlled by design rather than by policy. Autopilot is gated behind a sender allowlist, a confidence floor you set, a cancellable undo window, escalation keywords, optional working hours and a global pause.
Message bodies are treated as untrusted input against a fixed action allowlist, which is what the prompt-injection row needs. The audit log is append-only, records the reasoning behind each action, and exports as JSON or CSV, which is what turns 'we have a control' into a line an auditor can check. That fits an operator or a small client-facing team that needs each register row to be provable.
Where we are the wrong answer: if the requirement behind the entry is tenant-wide retention, legal hold and eDiscovery across the whole organisation, a mail client is not the control layer. Microsoft Purview for Microsoft 365 and Google Vault for Workspace do that at tenant level, and we do not compete with them; check their current capabilities on their own documentation before you write the control. Two other limits worth knowing if your register includes endpoint standardisation: we have no native Linux build, and Android is a progressive web app rather than a native one.
Frequently asked
See it in AI Emaily
Keep reading
Sources

Written by
Nafiul HasanNafiul Hasan is an entrepreneur and AI automation system builder with 10+ years of experience turning messy, manual workflows into reliable automated systems. He designs and ships AI enterprise solutions end-to-end — the agent logic, the data plumbing, and the product people actually use — and founded AI Emaily to give busy professionals their attention back. He writes here from the builder's seat: what works, what breaks, and how to put AI to work without giving up control.