Blog/ Buyer guides

What to Ask on an AI Email Tool Demo Call: 20 Questions

Nafiul HasanNafiul Hasan· 13 min read
AI email software demo call agenda showing four question buckets — messy-thread test, unsure-state test, audit-log test, and contract-and-data test — with 20 buyer questions

The short answer

Ask 20 questions that pull the demo off its script and onto your inbox: bring a messy thread of your own and ask for a live run, ask what the tool does when it is unsure, ask to see the audit log for a real send, and pin down who owns the data and who trained on it.

20 questions for an AI email software demo that break the scripted deck: your messy thread, unsure states, audit log, and destructive-action policy.

On this page
  1. 01The short answer
  2. 02Criteria that actually matter
  3. 031. The messy-thread test
  4. 042. The unsure-state test
  5. 053. The audit-log test
  6. 064. The contract-and-data test
  7. 07Scoring table: how to score answers live
  8. 08Worked example: fifteen minutes of an actual demo
  9. 09Red flags: vendor answers that should end the demo
  10. 10What we would pick, and why (honest)

An AI email tool demo is a sales deck driving a browser. The vendor picks the account, the thread, the drafts they show you, the moments they zoom in. What to ask on an AI email software demo is really one question in twenty pieces: how do I pull the call off its rails and onto my actual inbox for the twenty minutes I have?

The twenty questions below are designed for a shared screen — not for a spec sheet you email back and forth. Each one either happens live or does not, and the answer "we can send you a video of that" is itself an answer. Print them, keep them in your notes app, ask a colleague to time you: forty-five seconds per question, on the vendor's clock.

The short answer#

A good AI email demo is one you drive. Bring a messy thread from your own inbox — one that has forwards, quoted history, an attachment, a name spelled two ways, and a decision buried three replies deep — and ask the vendor to run their triage, draft, and follow-up flow on it live. Then ask about behaviour under uncertainty, the audit trail after a mistake, and the contract shape around your data.

Everything else on this page is a version of those four moves: the messy-thread test, the unsure-state test, the audit-log test, and the contract-and-data test. Five questions per bucket, twenty questions total, twenty minutes on the vendor's clock.

Criteria that actually matter#

Most demo scripts optimise for the twenty-minute wow. That is the wrong optimisation. You are hiring a system to touch your customers' inboxes for the next two years — so the criteria are the ones that hold up on a Tuesday afternoon, not the ones that photograph well. The four buckets below are the criteria; the questions inside them are the tests.

1. The messy-thread test#

Ask the vendor to open your live inbox — or an inbox you connected on the call, from a throwaway Gmail with a few forwarded threads — and run one triage plus one draft on a thread you pick. Not on their demo account. Not on a sandbox they prepared. Yours, or one that looks like yours, in front of you.

  • Can you connect my account, or a test account I share now, and triage my last 100 messages on this call?
  • On this thread I am sharing, draft a reply and talk me through what you pulled from the quote history versus what you invented.
  • Show me the same feature on my Outlook account, if you support it — or say plainly that you do not.
  • If I hover over a drafted sentence, can you tell me which prior message, contact note, or context source it came from?
  • Rerun the same triage now. Do the categories change? If yes, why?

2. The unsure-state test#

When the model does not have enough information to reply confidently, what happens? A good product tells you so, out loud, and pauses. A bad one guesses in the same tone as its confident answers. The confidence-floor concept is real; the specific percentage a vendor publishes is often not. Ask about the mechanism, not the number.

  • Show me a thread where the tool is unsure. What does the UI look like — a warning, a lower autonomy tier, a paused draft, or nothing at all?
  • Does the tool ever refuse to draft? On what triggers?
  • Draft a reply to a partially garbled message with a mixed intent — half question, half meeting-reschedule. Which intent did it pick, and why?
  • What does the tool do with an inbound message that quotes prior discussion of PHI, a wire-transfer instruction, or a bounced payment? Show me on a real example.
  • Can I set a per-recipient rule that says never draft, only summarise? Show me the toggle.

3. The audit-log test#

Every action the AI took should be visible after the fact — every draft it wrote, every message it filed, every rule it applied, every send it made. If a mistake reaches your customer, the audit log is what tells you what the model saw, what tier of autonomy was on, and who to talk to. This is the question set the vendor's slide deck skips.

  • Show me the audit log for a real send made in your demo account today. Can I filter by user, tool, model, and outcome?
  • If a draft went out that should not have, how do I recall or undo it — and where does that undo show up in the log?
  • What is the retention period, and can I export the log to my SIEM?
  • Does the log capture the input context — the prompt, the source thread, the rule that fired — or only the outcome?
  • If I disable logging for a specific mailbox, does the tool still function? The right answer is no.

4. The contract-and-data test#

Ask these before the demo ends, not on a follow-up call. A vendor that will not answer with a screen shared answers slower once you are off the call. Frame each question atomically — a compound question gets a compound answer.

  • Do you train any model on our email content? On our prompts? On our drafts? Answer each of the three separately.
  • Where is the data processed, and where is it stored? Regions and providers, by name.
  • What are your minimum OAuth scopes for Gmail, and do you hold a current Letter of Assessment from a Google-designated third party (CASA) for restricted scopes like gmail.modify?
  • If we cancel, how long until our data is deleted, and does that include the vector store and the audit log?
  • Show me the sub-processor list and the date of your last SOC 2 Type II or ISO 27001 report.

Google's rules on Gmail scopes are not optional

Google's API Services User Data Policy requires an annual independent security assessment (CASA) for apps using restricted scopes such as gmail.modify. A vendor that cannot say when their last Letter of Assessment was dated has an answer you can verify against Google's own policy page.

Scoring table: how to score answers live#

Use the table below as your live scoresheet. Print it, or keep it open in a second window. The point is not to catch the vendor out — it is to make sure the answers you write down after the call are the ones the vendor actually gave, not the ones you remember them giving. A strong answer looks specific; a weak one sounds general.

QuestionWeak answer sounds likeStrong answer looks like
Connect my account and triage 100 messages now."We usually don't demo on customer inboxes."OAuth flow live, real triage in under two minutes, categories explained on hover.
Draft on my thread and show me your sources.A draft appears with no visible sources.Hovering a sentence surfaces the prior message, contact note, or context entry it came from.
Show me a thread where the tool is unsure."It's usually confident enough."A visible confidence signal in the end-user UI — banner, tier drop, or a paused draft with a reason.
Show me the audit log for a real send."Logging is on the enterprise plan."Filterable log with actor, tool, model, input context, and outcome, exportable to a SIEM.
Do you train on our content, prompts, or drafts?"No, we don't train on your data."Three separate written answers, plus DPA, sub-processor list, and processing region.
Undo a send made 30 seconds ago.A grey "Undo" toast that has already vanished.A per-action undo bound to the audit log, working after the delay window on autopilot sends.
Change autonomy per recipient in front of me."That's on our roadmap."Per-mailbox, per-recipient toggles that persist and appear in the audit log.

Worked example: fifteen minutes of an actual demo#

On a call last quarter with a vendor we will not name, minute one goes like this. The vendor shares a demo Gmail account seeded with clean, short threads and three named senders. You interrupt at minute two and ask: can you connect a Gmail account I created this morning?

They pause. They connect it. Your test inbox has forty-two forwarded threads from a colleague, one attachment of a mislabelled invoice, and a thread from your father with three replies and a photo. You ask the vendor to run their triage on the forty-two threads.

The triage runs. Twelve threads go to a category called "Meeting." You open two: neither is a meeting. The vendor says the categories improve as the model learns. You ask: without any training on my data, how do they improve? The answer takes forty-five seconds and does not name a mechanism.

You move to the unsure-state test. Show me a thread where you are unsure. The vendor picks a thread. It is not unsure — the model produced a confident draft. You ask if the confidence signal is visible in the end-user UI on any thread. The vendor points at a percentage in a debug panel. That panel is not visible to end users.

You end the call at minute fifteen with three clear answers written in your notes: they cannot draft with source citations, their unsure-state signal is a hidden debug field, and their audit log records outcomes but not the input context. You still have five minutes for the contract questions, which is more than most buyers save.

Red flags: vendor answers that should end the demo#

None of the phrases below are automatic disqualifiers on their own. Two or three of them in the same call are. Each one is a sign the vendor is answering the question they wish you asked instead of the one you asked.

  • "We usually don't demo on customer inboxes for security reasons." Yours is a fresh test inbox with no real data; they don't want to.
  • "That feature works, but we need to switch accounts to show it." Their demo account has a state your account does not have.
  • "The audit log is on the enterprise plan." Fine — show it anyway, on a tier we can trial.
  • "We can send you a video afterwards." The video is a rehearsed version of what you asked to see live, which was the point of asking.
  • "We're the only vendor without training on user data." Absolute claims about competitors are easy to check on their privacy pages and rarely hold up.
  • "Compliance-wise, we're covered." A vague answer to a specific question. Ask for the standard by name — SOC 2 Type II, ISO 27001, CASA — and the date of the last report.
  • "Undo works, but only if you catch it in the first five seconds." That is an unsend toast, not undo. Real undo lives in the audit log and works after the send has left.
  • "Our confidence threshold is 87%." A specific percentage is a red flag because two different vendor docs will disagree with each other. Ask about the mechanism — what the UI shows the user when the model is unsure — not the number.

One rehearsed answer is fine, three is a pattern

Vendors are allowed to be polished. A sales team that has never met your inbox and still shows you a working triage on it in ninety seconds is a good sign, not a suspicious one. The pattern to watch for is a call that keeps rerouting back to the demo account whenever a question gets specific. That is the tell.

What we would pick, and why (honest)#

By the twenty-question rubric above, AI Emaily is the tool we built to answer those questions in the affirmative. Approve-before-send is the default in Copilot mode; drafts show the sources they pulled from; the audit log is per-action and captures what the model saw, not only what it did; undo is bound to the log rather than a five-second toast; autonomy is per-recipient rather than one global setting; and we do not train models on user mail, prompts, or drafts. We build AI Emaily, and this recommendation reflects that — the disclosure is the reason to trust the rest of the paragraph.

Who we are right for: a team that has to answer a governance or security question about the AI in their inbox, and wants a product where the honest answer on a shared screen is "yes" across all four buckets — messy thread, unsure state, audit log, data contract. On Gmail, Outlook, iCloud, Fastmail, Proton, and any IMAP account, with a downloadable desktop app on macOS (Apple Silicon) and Windows, a native iOS app, and Android as a PWA.

Who we are not right for. If you live entirely inside Gmail and what you want above all else is the fastest keyboard-driven power-user workflow with the tightest hooks into Gmail's own archive search, Shortwave has built harder on that surface than we have. If your governance program is anchored in a Microsoft-tenant surface — Purview, Sentinel, ServiceNow IRM — and you expect every AI action to feed those platforms natively, Microsoft 365 Copilot has built harder on tenant-native reporting than we have, and for a Microsoft-standardised enterprise that is the correct pick.

Both of those are legitimate answers to a different question than the one this rubric asks. If the question you are asking a demo is "can you prove this on my inbox in the next twenty minutes," the twenty questions above are the ones to ask — of us, and of everyone else you are considering. Verify our current capabilities and packaging on our live pricing and features pages before you sign anything; we ship most weeks and the specifics move.

Book a demo you can drive

We prefer being asked the twenty questions above to answering a spec sheet. Bring a throwaway Gmail with a few forwarded threads, or connect your own account and disconnect it at the end of the call.

Frequently asked

Nafiul Hasan

Written by

Nafiul Hasan

Nafiul Hasan is an entrepreneur and AI automation system builder with 10+ years of experience turning messy, manual workflows into reliable automated systems. He designs and ships AI enterprise solutions end-to-end — the agent logic, the data plumbing, and the product people actually use — and founded AI Emaily to give busy professionals their attention back. He writes here from the builder's seat: what works, what breaks, and how to put AI to work without giving up control.

EntrepreneurAI Automation System BuilderAI EnthusiastBuilds AI Enterprise Solutions10+ years experience
More from Nafiul
Ready when you are

Bring the twenty questions to a demo you drive.

AI Emaily runs Manual, Copilot, and Autopilot modes with approve-before-send, a per-action audit log, one-tap undo bound to the log, per-recipient autonomy, and no training on user mail — on Gmail, Outlook, iCloud, Fastmail, Proton, and any IMAP account.

  • 7-day free trial
  • Cancel anytime
  • Every provider