Blog/ Buyer guides

How to Evaluate AI Email Draft Quality: Test, Don't Guess

Nafiul HasanNafiul Hasan· 11 min read
AI Emaily blog cover for how to evaluate AI email draft quality, showing a scoring rubric for AI-generated email replies

The short answer

Freeze ten real threads from your own mailbox and score every candidate's reply on six dimensions: factual accuracy, invented commitments, tone match, length, awkward-case handling, and edit distance before you would send. Weight invented commitments heaviest. Then deploy the winner for a week and record how many drafts you sent untouched.

How to evaluate AI email draft quality: a ten-thread test, a six-dimension scoring rubric, and the send-without-edit rate that settles it.

On this page
  1. 01The short answer
  2. 02Why three drafts in a demo tell you nothing
  3. 03Before you start: freeze ten threads
  4. 04The six dimensions, and how to score them
  5. 05Hallucinated commitments get a veto, not a weight
  6. 06The protocol
  7. 07Turning the edit into a number
  8. 08The outcome measure: send-without-edit rate
  9. 09Where this protocol scores us down
  10. 10When two tools finish close

Nobody publishes a repeatable way to test AI email draft quality, so most buyers judge it on three drafts generated in a demo, on threads the vendor chose. That is a taste test.

This is the method instead: ten threads you pick from your own mailbox, six scoring dimensions, the same set run through every candidate, and one outcome number a week later.

We build AI Emaily, so we are one of the tools you would run it on. Every dimension below is something any drafting tool can be scored on, which means the protocol can fail us. Near the end we name the case where it does.

The short answer#

Freeze ten representative threads, generate one reply to each in every candidate, and score six dimensions: factual accuracy against the thread, invented commitments, tone match, length, how it handles the awkward case, and how much editing it needs before you would send.

Weight invented commitments heaviest. A draft that reads well and quietly agrees to a Friday deadline you never offered is worse than a clumsy one you rewrite, because you will approve it.

Then run the winner on real mail for a week and record the outcome measure: the share of drafts you sent without changing a word.

Why three drafts in a demo tell you nothing#

Demo threads are short, polite and unambiguous. The messages that consume your day are none of those: a buried question, one angry sentence, two people copied who should not be.

A tool that answers a scheduling request well has told you nothing about how it answers a scope dispute. The gap between easy and hard threads is wider than the gap between products, so easy threads measure the threads, not the tools.

Before you start: freeze ten threads#

Copy ten real threads into one document and give each an ID. This set is your instrument, and an instrument you adjust between measurements is not one.

Aim for a mix that looks like an ordinary week rather than your best day.

  • Three routine threads: a scheduling reply, an acknowledgement, a short factual answer. Every candidate should pass these.
  • Two long threads: forty or more messages with nested quoting, where the answer depends on something said near the top.
  • Two threads where an attachment carries the answer, such as a figure in a spreadsheet or a number in a quote PDF.
  • Two awkward threads: a complaint, a price pushback, or a request you have to decline. These separate the products.
  • One trap thread, where the correct reply commits to nothing because you have not decided yet.

Ten real threads, four vendors

This protocol means pasting confidential mail into every trial you open. Ask each vendor in writing how long message content is retained and whether any of it trains models. If you cannot get an answer, replace names and figures on the awkward threads. The drafting behaviour you are testing survives redaction; your obligations do not.

The six dimensions, and how to score them#

Each dimension scores 0 to 3 and carries a weight. The weights sum to 12, so a tool's ceiling is 36. Take each dimension's median across the ten drafts rather than its mean, so one disaster does not average away and one excellent draft does not carry the set.

Dimension (weight)What you are scoringScores 3Scores 0
Factual accuracy (x3)Whether every claim traces back to a message in the thread: dates, figures, names, who said what.Nothing asserted that the thread does not support, and unknowns left visibly open.One confident sentence about something the thread never said.
Invented commitments (x3)Promises made on your behalf: deadlines, discounts, scope, meetings, introductions, refunds.Zero across all ten. Where a commitment is expected, it leaves a blank or asks you.One invented commitment anywhere in the set. See the veto rule below.
Tone match (x2)Whether it reads like you writing to that recipient: formality, warmth, directness, sign-off.You would send it to that person without softening or stiffening it first.One register for everyone, or a brightness that would embarrass you on the complaint thread.
Length (x1)Whether the draft is as long as the reply needs, not as long as the model can write.Three lines back to a three-line question, and structure only where structure helps.Two paragraphs of preamble before the answer, or one line where a decision needed explaining.
The awkward case (x2)Scored on the two hard threads only: can it decline, disagree or apologise without giving ground.It holds the position, stays civil, and invents no concession to end the discomfort.It capitulates, or it dodges by proposing a call instead of answering.
Edit distance (x1)How much of the draft you change before you would actually send it.You adjust a sentence. Under roughly a tenth of the words move.You rewrite the middle, or delete it and start again.

Hallucinated commitments get a veto, not a weight#

Score invented commitments across the whole set rather than per draft, and treat a single one as a zero for that tool. Not a deduction. A zero, with a rule attached: a tool that invented one commitment does not send anything unattended, whatever the other five dimensions say.

The reason is the cost shape. A clumsy draft costs ninety seconds of editing. A fluent draft that agrees to a delivery date nobody offered costs a renegotiation, and you are less likely to catch it, because it reads like the nine you already approved.

NIST's Generative AI Profile, published July 2024, names confabulation among twelve generative AI risks and defines it as content stated confidently but not true. The confidence is the problem. Fluency is what stops you reading carefully, so the best-written tool can be the most dangerous one on this dimension.

Scattered coloured shapes on the left and the same shapes arranged in an ordered grid on the right with one square picked out in green, representing drafts judged by impression versus drafts scored against a fixed rubric
Same judgement, made in a fixed order, so two people reach the same number.

A clean ten is not a rate

Ten threads cannot show that a tool invents commitments in under ten percent of drafts. All a clean set proves is that it did not invent one in ten tries. Write that in your notes, keep the tally running after you deploy, and let the accumulated count decide whether it ever sends without you.

The protocol#

Ninety minutes of screening, then a week of ordinary work. Do not skip step four; it is the one that stops you scoring the brand.

  1. 1

    Freeze the set

    Copy ten threads into one document with IDs. Note what a good reply to each would have to contain before you see any draft. Five words per thread is enough.

  2. 2

    Build the sheet first

    One row per thread per tool, one column per dimension, plus free text for what you changed. Writing the sheet first stops you inventing criteria that fit the winner.

  3. 3

    Generate one draft each, first attempt only

    No regenerating, no coaching, no second prompt. You are scoring default behaviour, because the default is what you live with on a busy morning.

  4. 4

    Score blind, in rotated order

    Strip the tool names, shuffle the drafts, and score dimension by dimension across all tools rather than tool by tool. Brand recognition moves scores more than people expect.

  5. 5

    Do a separate commitments pass

    Reread every draft looking only for promises: dates, prices, scope, availability. Checking for invented commitments while also judging tone is how you miss them.

  6. 6

    Record the edit, not the impression

    Edit each draft until you would genuinely send it, then note roughly what fraction of the words changed and where the changes landed.

  7. 7

    Run the winner for one working week

    Keep the single tally that matters: sent untouched, edited, discarded. Five days of real mail overrules any afternoon of scoring.

Turning the edit into a number#

Machine translation solved this measurement in 2006 and email never borrowed it. Human-targeted Translation Edit Rate, from Snover and colleagues, counts the word-level insertions, deletions, substitutions and phrase moves needed to turn machine output into what a person would actually ship.

You do not need their tooling. Edit until you would send, then estimate the fraction of words you touched. A band beats a decimal you did not really measure.

Share of words you changedWhat it meansScore
Under 10%Adjusting a sentence or a sign-off. This is what a usable drafting tool feels like.3
10 to 30%Real editing, still faster than writing. Acceptable on hard threads, a warning sign on routine ones.2
30 to 60%You are rewriting inside a structure someone else chose, which is often slower than a blank page.1
Over 60%, or deletedThe draft cost you time. Count these separately: three in ten is a verdict on its own.0

Track where the edits land as well as how many. Edits in the opening line are a tone problem, and tone is configurable in most tools. Edits in the middle are a comprehension problem, and no setting fixes that.

The outcome measure: send-without-edit rate#

Everything above screens candidates in an afternoon. The number that tells you whether the purchase was right arrives a week later, and it is one fraction.

Run the finalist on real mail for five working days and keep three counts: drafts sent with no edit, drafts edited, drafts discarded. Resist adding a fourth.

The week-one tally, filled in to show the shape
Drafts generated62
Sent without editing21, or 34 percent
Edited before sending33, most edits in the opening line
Discarded8, six of them on threads with an attachment
Invented commitments1, a delivery date nobody offered. Unattended sending stays off.

Read it as a direction, not a grade. A third sent untouched on a mixed week is a tool doing real work; under one in ten and you are proofreading a machine you pay for.

The line that decides your autonomy setting is the last one, not the first.

Where this protocol scores us down#

We build AI Emaily, and the method above was not shaped to flatter it. Two places it costs us points, stated plainly, because a protocol that cannot fail its author is marketing.

Cold start. Our tone control is a Personal Context you write plus per-client profiles you set. We do not infer a voice from your past messages, by design, and we do not train on your mail. Run the ten threads five minutes after signing up and you are scoring an empty Context. A tool that mines your sent folder can beat us on that first pass.

Attachments and long threads. We summarise long threads and read attachments, but we publish no thread-length or file-type limits, so the four hardest threads in your set are ones you have to run rather than take our word on.

What we claim here is narrow: sends are approval-gated by default, every action lands in an audit log with undo, and the invented-commitment count is a number you can hold us to as easily as anyone else. OWASP's 2025 list treats Misinformation (LLM09) and Excessive Agency (LLM06) as separate risks, and a tool that drafts and can also send meets both.

When two tools finish close#

If the totals land within three points, the rubric has done its job and stopped being useful. Put the two awkward threads side by side and read the drafts.

Then ask which one you would rather correct at six in the evening. That question has a real answer, and it predicts the next six months better than a total does.

Whatever you pick, keep the ten threads. Rerun them in six months on the tool you bought. Models get swapped underneath you without a release note, and a frozen set is the only way you notice the day the drafts get worse.

Frequently asked

Nafiul Hasan

Written by

Nafiul Hasan

Nafiul Hasan is an entrepreneur and AI automation system builder with 10+ years of experience turning messy, manual workflows into reliable automated systems. He designs and ships AI enterprise solutions end-to-end — the agent logic, the data plumbing, and the product people actually use — and founded AI Emaily to give busy professionals their attention back. He writes here from the builder's seat: what works, what breaks, and how to put AI to work without giving up control.

EntrepreneurAI Automation System BuilderAI EnthusiastBuilds AI Enterprise Solutions10+ years experience
More from Nafiul
Ready when you are

Run the ten-thread test on us.

AI Emaily drafts from a Personal Context you write and client profiles you set, keeps sends approval-gated by default, and logs every action with undo. Score us on all six dimensions with your own threads, free at app.aiemaily.com/signup.

  • 7-day free trial
  • Cancel anytime
  • Every provider