Blog/ AI email prompts & use-cases

How to Test Whether an AI Really Writes in Your Voice

Nafiul HasanNafiul Hasan· 12 min read
AI Emaily blog cover: a protocol to test whether an AI really writes email in your voice, using blind colleague tests and edit-distance measurement

The short answer

Run a blind test: pick five real emails you sent, give three colleagues five AI drafts of the same briefs, and ask which are yours. Then count how many edits you make to your last twenty AI drafts. Together, both measures reveal whether your voice specification is capturing you accurately or still missing something specific.

A practical method to test if AI writes in your voice: blind colleague guesses, edit-distance counts, and a written personal style spec.

On this page
  1. 01The short answer: two measures, run together
  2. 02Before you start
  3. 03The five-step blind test protocol
  4. 04How different AI approaches perform on this test
  5. 05What to do when the test fails
  6. 06A faster way to keep voice consistent across every draft

The question of how to test if AI writes in your voice tends to come up a few weeks after you start using an AI email tool — not at the beginning, when everything feels new, but once you have settled into a routine and can no longer tell whether the drafts are actually sounding like you or whether you are just editing them quickly enough that you stopped noticing. The tool says it matches your voice. The drafts arrive fast. But are they you?

This is a measurement problem, not a prompting problem. You can spend time reading guides about how to configure voice instructions, adjust tone settings, or write a better system prompt — and still not know whether it worked. Testing is the only way to get a real answer, and most people skip it entirely because there is no obvious protocol to follow.

This guide gives you one. It has five steps: select representative emails from your sent folder, write the briefs for them, generate AI drafts, run a blind test with three colleagues, and count your edit rate across a larger sample. Run it once after initial setup. Run it again after any major update to your voice configuration. The result is a number you can act on rather than a gut feeling.

The short answer: two measures, run together#

You test AI voice matching by running two measures in parallel: a blind colleague test and an edit-distance count. The blind test catches the obvious misses — the wrong opener, the wrong level of directness, the wrong sign-off. The edit count catches the smaller, accumulating divergences you stopped noticing because you correct them automatically.

Neither measure alone is sufficient. A colleague test on a small sample can pass by chance. An edit count does not tell you which habits are off, only how many words are changing. Together, they give a clear picture: pass rate above chance on the blind test, plus an edit rate below your own threshold, means the AI is writing in your voice. Anything else means the specification needs work.

The specification is the key word. This test measures whether your written voice spec — the instructions, examples, and profile settings you have given the tool — captures your real habits accurately enough. A failing result is not a verdict on the AI's capability. It is a signal that the spec is missing something specific, and the test results point to what.

Before you start#

You need three things. First, access to your last fifty sent emails. You will only use five as test cases, but you need enough to pick ones that are genuinely representative: an internal note, a client status update, a request to a vendor, a reply to a warm lead, a short follow-up. If your sent folder is thin, run the test with what you have — the sample will be smaller and the conclusions softer, but it is still worth doing.

Second, three colleagues who read your emails regularly enough to have an opinion about how you write. They do not need to be technical. They need to know your voice — whether you open directly or with context, whether you hedge or push, how warm or clipped you run. An account manager who reads your client emails, a manager you communicate with daily, a close collaborator. Anyone who would notice if your tone shifted.

Third, a way to record your edits on twenty recent AI drafts. Paste the draft into a document before editing, paste the sent version after, and count the words you changed. A rough word-change count is sufficient — you are looking for patterns, not precision. If you did not save pre-edit versions, run the edit-count measure on the next twenty drafts going forward and revisit in three weeks.

What you are measuring

You are measuring whether your voice specification — the instructions, examples, and profile settings you gave the tool — captures your real writing habits. A failing result does not mean the AI is broken. It means the spec is incomplete, and the test tells you exactly where.

The five-step blind test protocol#

Run these steps in order. Each builds on the previous. If you skip ahead to the blind test without writing clean briefs first, you will not be able to tell whether failures are from the AI or from an ambiguous brief.

  1. 1

    Select five sent emails

    Pick five emails covering different registers: one internal, one external client, one to a vendor, one warm reply, one short follow-up. They should be emails you consider representative of how you normally write — not a message you had to revise three times or send under pressure. These five are your benchmark.

  2. 2

    Write the brief for each

    For each email, write a one- to three-sentence brief describing the situation, the recipient, and the goal. Do not include any language from the original. The brief should be the kind of input you would actually give the AI in a normal workflow — context only, not a detailed script. You are testing what the tool does with a realistic prompt, not an optimized one.

  3. 3

    Generate the AI drafts

    Run each brief through your AI tool with your current voice configuration active. Save the first draft produced for each brief and do not adjust it. You are capturing baseline output — what the tool does before any coaxing — not the best output it can produce after iteration.

  4. 4

    Strip identifiers and send to colleagues blind

    Remove recipient names and any context-specific details that would make the email identifiable from the brief alone. Present the five real emails and five AI drafts as ten numbered samples — shuffled, not grouped — and ask your three colleagues to mark each as yours or not yours, independently. Do not tell them how many are real.

  5. 5

    Score results and identify patterns

    For each AI draft: all three colleagues correctly identifying it as AI-written is a clear fail. Two out of three is a marginal fail. One out of three is a marginal pass. Zero is a clear pass. More useful than the overall score: note which specific habits the colleagues flagged — opener phrasing, sign-off, directness, sentence length. Those patterns are your fix list.

How different AI approaches perform on this test#

The blind test tends to expose a consistent gap between tool types. Understanding why each approach performs the way it does helps you interpret your own results and decide whether the fix is in your specification or in your choice of tool.

A general chatbot — ChatGPT, Claude, or Gemini used in a separate browser tab — resets every session. Whatever voice coaching you did Monday is gone by Tuesday. Drafts from a session-only prompt typically land in the marginal-fail zone: colleagues who know you well notice something is slightly off, even if they cannot name it precisely. The edit count tends to be highest on openers and sign-offs, because those are the habits most resistant to generic instruction.

An inline or extension-based assistant that operates inside Gmail or Outlook tends to improve on the blank-chatbot baseline, because it has access to the current thread and often carries a tone preference across drafts. But a tone preference is a category, not a profile. It narrows the range and still leaves you in the generic-leaning zone for anything outside that preset.

ApproachVoice sourcePersists across sessions?What the blind test typically surfaces
General chatbot (separate tab)Session prompt onlyNo — resets each timeOpeners and sign-offs caught by colleagues who know you well
Inline or extension assistantTone setting + current threadTone sticks; voice habits do notMarginal fails on directness and register shifts by recipient type
AI client with persistent profileWritten spec + per-recipient settingsYes — applied to every draftSmallest gap; fails most often on registers not covered by the spec examples
Any tool with no voice spec configuredDefault model behaviorNot applicableColleagues identify AI drafts reliably — all the generic tells are present

What to do when the test fails#

A failing result narrows the problem to one of three things: a gap in the written spec, a recipient register not covered by the examples, or a habit so automatic you did not think to describe it. The fix always moves in the same direction — make the spec more specific — but the test results tell you where to start.

If openers are consistently flagged, go back to your five benchmark emails and write down exactly how you opened each one. You will likely find a pattern you did not consciously know you had: a half-sentence acknowledgment before the ask, skipping the greeting entirely in internal threads, or opening with a question when the recipient is new. Add that as a concrete example in your spec, not as a rule description. Examples outperform rules every time.

If the directness level is what gives drafts away — too hedged where you would push, too blunt where you would soften — the fix is a per-recipient example for each register where the shift happens. Note why the register changes, not just what it looks like. That context is what the AI uses to generalize correctly to recipients not explicitly named in the spec.

If your edit rate is high on content rather than voice — the facts are right, the structure is wrong — the issue is usually in the brief, not the voice spec. A brief that omits the goal leaves the AI guessing at structure. The middle path: brief that gives recipient, situation, and goal, and lets the voice spec handle everything else.

A magnifier examining a document, representing the process of auditing AI email drafts against a personal voice specification to identify the specific habits the spec is missing
The audit step — comparing AI drafts against your benchmark emails — pinpoints which specific habits are missing from your spec rather than leaving you with a vague sense that something is off.

Failing on rare registers is expected

A well-configured spec will still fail on registers it has never seen an example of. If a colleague identifies an AI draft as not yours, check whether that email type — a difficult feedback message, a commercial negotiation, a complaint to a vendor — appears anywhere in your spec. If not, add one example for that type and run a targeted one-email retest.

A faster way to keep voice consistent across every draft#

Running this test once tells you where you stand. Running it every few months tells you whether your spec is drifting as your own writing evolves. The manual loop — update the spec, retest, adjust — is the right approach for any tool, but it is a periodic exercise rather than a continuous one.

AI Emaily is designed around the idea that you should not have to re-test and re-spec on a fixed schedule. Your voice specification lives in your Personal Context brain — a set of facts, examples, and preferences you write and control — and per-client profiles let you hold different registers for different recipients without flattening everything into one setting. Every draft applies the spec as you have written it, and you approve before anything sends. When a draft comes back off on something specific, you update that part of the spec and the next draft reflects it immediately.

We build AI Emaily. If the protocol above showed a gap between what the AI produces and what you actually write, that is exactly the problem AI Emaily is designed to close — not by learning from your past mail, but by holding the spec you write and applying it to every draft, per recipient, with you in control. A 7-day free trial is available at aiemaily.com (card required; cancel before day 7 to pay nothing). Pricing is at aiemaily.com/pricing.

Frequently asked

Nafiul Hasan

Written by

Nafiul Hasan

Nafiul Hasan is an entrepreneur and AI automation system builder with 10+ years of experience turning messy, manual workflows into reliable automated systems. He designs and ships AI enterprise solutions end-to-end — the agent logic, the data plumbing, and the product people actually use — and founded AI Emaily to give busy professionals their attention back. He writes here from the builder's seat: what works, what breaks, and how to put AI to work without giving up control.

EntrepreneurAI Automation System BuilderAI EnthusiastBuilds AI Enterprise Solutions10+ years experience
More from Nafiul
Ready when you are

Stop guessing whether the AI sounds like you.

AI Emaily holds your voice specification in a Personal Context brain and applies it per recipient on every draft. You approve before anything sends. 7-day free trial at aiemaily.com (card required; cancel before day 7 to pay nothing). Pricing at aiemaily.com/pricing.

  • 7-day free trial
  • Cancel anytime
  • Every provider