ManageTime

Article

Does letting AI draft your email replies actually save time?

Letting AI draft your work emails reliably cuts the effort a reply costs and unreliably cuts the time. In the one workplace study that measured both, clinicians using AI-drafted replies reported a sharp drop in mental workload while their actual reply times did not move at all. Use it against inbox fatigue, not against the clock. The gap between those two results decides which messages are worth handing over. This article covers where the widely quoted 40 percent time saving came from, why it does not survive contact with a real inbox, and what a draft still costs you to check.

Does AI drafting actually save time on email?

The number you will meet first comes from a 2023 experiment in Science by Noy and Zhang, and it is a large one: 453 college-educated professionals were randomly given access to ChatGPT for mid-level writing tasks, the average time taken fell by 40 percent, and rated output quality rose by 18 percent. Search results about AI and email repeat that pair of figures constantly. What they leave out is the setup. Participants were handed discrete writing assignments and asked to complete them. Nobody was working an inbox, deciding which messages deserved a reply, or living afterwards with what they sent.

The closest thing to a field test points the other way. In 2024 Garcia and colleagues published a five-week study in JAMA Network Open covering 162 clinicians who were given AI-generated draft replies inside their real patient inbox. Reply times did not move. The study measured reply action time, write time and read time across the pilot, and found no change in any of them against the weeks before it. Clinicians used the drafts for about a fifth of the messages available to them.

The two results are not in conflict. They answer different questions: one asks how fast you can produce a piece of writing, the other asks what happens to your working day.

So the honest answer to whether AI saves time on email is that nobody has yet shown it does, in an inbox, on the clock. Enthusiasm thins as well. Set against workers who never got access, the group Noy and Zhang exposed to ChatGPT was twice as likely to still be using it at work two weeks later, and 1.6 times as likely at two months. None of that is a reason to skip it. Something else in the Garcia study moved a long way, and it moved in the direction most people mean when they say email is grinding them down.

Why does the effort drop when the clock does not?

What fell was the load, not the minutes. The same clinicians rated their task load at 61.3 before the pilot and 47.3 during it, a drop of nearly fourteen points on that scale, and their work exhaustion scores fell from 1.95 to 1.62. Same clock, lighter day. If you have ever sat in front of a message you knew how to answer and still not started, that gap will be familiar.

The mechanism stops being mysterious once you count what a reply actually costs. Opening the message, reading it, working out what it is asking, deciding whether you are even the right person to answer, checking one fact, picking a tone for a colleague who was short with you last week: composition is one slice of that, and rarely the widest. A draft removes the blank page. It leaves everything before the blank page untouched, and it adds a step after it, because now you have to read what the model wrote and judge whether it is true.

Effort and duration are different currencies. Email spends both, and the evidence so far says AI drafting refunds one of them.

One caveat sits under all of this. Both field studies ran in clinician inboxes: patient portal messages, a narrow genre with repeating question types, a defined relationship, and a permanent medical record attached to every word. Your inbox is probably messier and lower-stakes, which cuts both ways, because the checking burden that ate the clinicians' time savings is lighter for you, and so is the payoff on any one message. Treat the direction of these findings as informative and the exact figures as borrowed. The shape of your email day still comes from the email pillar, since how often you open the inbox decides more than what writes the replies, and a session with a hard edge, which is Timeboxing at half-hour scale, stops a drafting tool from turning a clearing session into an afternoon.

What happens when nobody checks the draft?

The checking step is where AI drafting goes wrong, and there is now a study about exactly that. In 2025 Biro and colleagues published work in npj Digital Medicine that gave primary care physicians a set of 18 patient messages with AI-drafted replies attached. Four of those drafts carried errors, either a plain inaccuracy or an omission that could cause harm. For each of the four, between 13 and 15 of the participating physicians failed to address the error adequately, and between 35 and 45 percent of the flawed drafts were sent on entirely unedited.

The same physicians liked the tool. Eighty percent agreed the drafts cut their cognitive workload, and 75 percent judged them safe. Both things held at once, and the combination is the warning: the drafts did reduce effort, and the people using them could not reliably tell which ones were wrong.

Fluent writing reads as correct writing. That is a property of the prose, not evidence about the facts.

In practice that means reading a draft for what it claims rather than how it sounds, and reading hardest for what is missing, since omissions were part of what those physicians waved through and an absence is harder to catch than a mistake. Two habits carry most of the work. Never let a draft assert a fact, a date, a number or a commitment you have not confirmed yourself, and rewrite any sentence that promises something on your behalf. The Two-Minute Rule (for tasks) supplies a useful line here: if checking the draft will take longer than writing the reply yourself, the draft was not worth asking for.

How does your working style change the answer?

The research names conditions about the message; your working style adds conditions about you. In practice, people closest to this site's Architect style gain least on routine mail. A settled template and a decided tone already do what the draft would do, and the model mostly proposes what they were about to type. People nearer the Improviser or Sprinter styles gain most. They stall at the opening line and then finish quickly once moving, so a draft is a way past the start rather than a way to the end. Visionary-leaning readers reach for it on the long explanatory message they keep postponing, then rewrite most of it. That is a fair trade if the postponing was the real cost.

The honest limits, so this page does not oversell it. Keep anything where being wrong is expensive: a negotiation, a correction, bad news, anything a lawyer or a regulator might read back to you later. Keep the two-line reply too, for the opposite reason. Checking a draft of it costs more attention than typing it. And drafting will not fix an inbox problem, because a faster reply to a message that should never have reached you is still time spent on the wrong thing.

Where it fits is volume and fatigue. Think of the long tail of routine messages that are not hard, only numerous, at four in the afternoon when the hard ones are done.

Try it for one week on one category: the routine messages you can name in advance. Hold one rule while you do, that you verify every fact and every promise before sending. If the week leaves you less worn out at five, keep it, whatever the clock says. And if you are unsure which of the patterns above is yours, the working style self-check is a faster read than another week of guessing.

Frequently asked questions

Should you tell people an email was written by AI?

Say so when the substance of the answer came from the model, and skip it when you decided the content and the model only shaped the sentences. No study settles this, so treat it as a question about accuracy rather than etiquette. Your reader is owed a correct message and an honest sender, not an inventory of your tools. In practice, teams that agreed a norm out loud argue about it far less than teams where everyone guesses.

What kind of emails should you not let AI write?

Keep anything with a real cost of being wrong: negotiations, corrections, bad news, performance conversations, and any message that could be read back to you later. Short replies are the other exception, for the opposite reason. When a message takes two lines, checking a draft costs more attention than writing it, and the studies behind any drafting benefit used substantial replies rather than one-liners.

How much should you edit an AI-drafted reply?

Enough to have verified every fact, number, date and commitment in it, which usually means two passes: one for what the draft says, one for what it left out. The 2025 npj Digital Medicine study found physicians sending 35 to 45 percent of erroneous drafts entirely unedited. The failure mode is trusting a draft that reads well, not polishing it too lightly. Rewrite anything that promises something on your behalf, since the model has no idea what you can deliver.

Which working styles get the most out of AI email drafting?

People who stall at the opening line and then move quickly get the most, which on this site sits closest to the Improviser and Sprinter styles. Readers nearer the Architect style tend to gain least on routine mail, because a template already covers it. That is a pattern from practice rather than a research finding: the published studies measured jobs and message types, not working styles.

Sources