Normal view

AI written cold emails vs human written. 3 months of real testing

we burned about $14k in tooling and 3 months of pipeline time running this test so hopefully some of you can skip the expensive part and just read this instead

ok the raw numbers first because thats what matters. this is across our full SDR team of 6, selling cybersecurity solutions into mid-market and enterprise, $98k ACV, 30 months into this role so i have decent baseline data to compare against

TOTAL SENDS (3 month test period, jan through march 2025) human written campaigns: 11,400 emails across 3 reps AI written campaigns: 10,800 emails across 3 reps

REPLY RATES human: 3.1% overall (positive reply rate 1.4%) AI: 4.7% overall (positive reply rate 2.2%)

BOUNCE RATES human group: 1.8% AI group: 1.6% (slightly lower because the AI reps happened to get cleaner lists that month, not really meaningful)

MEETINGS BOOKED human: 38 AI: 54

COST PER MEETING human: ~$312 (just tooling, not counting rep salary) AI: ~$287 (tooling plus AI costs)

ok so context. our CRO came to me in december and basically said she wanted to see real data on whether AI written sequences could outperform what our reps were writing manually. weve been using Smartlead for sending for about a year, running 22k emails a month across the full team, Maildoso inboxes, the whole setup. the question wasnt about infrastructure it was purely about copy.

the test design was pretty simple. i split the team into two groups of 3. group A kept writing their own sequences the way they always had. group B got access to Claude and a custom GPT we built that was trained on our best performing sequences from the last 18 months. group B was told to use AI for first drafts of every sequence and they could edit from there but had to start with AI output. both groups targeted the same ICPs, same titles (mostly CISOs, VP Security, IT Directors at companies 200-2000 employees), same verticals.

we kept the enrichment and verification pipeline identical for both. lists built in ZoomInfo, enriched through Prospeo for email finding, verified with NeverBounce, loaded into Smartlead. same warmup protocols same sending schedules same everything except the copy.

what surprised me was where the AI group won and where they didnt.

the AI sequences were noticeably better at first touches. like significantly. reply rates on email 1 were 5.3% for AI vs 2.8% for human. thats almost double. and i think the reason is that the AI was better at writing concise openers that felt research-driven without being fake. our reps have a tendency to write these long intros where theyre clearly just restating the prospects linkedin headline back to them and it reads as exactly what it is. the AI drafts were tighter, usually 3-4 sentences for the opener, and they got to the pain point faster.

but on follow ups the gap narrowed a lot. by email 3 and 4 in the sequence the human written stuff was actually performing about the same. reply rates on email 4 were basically identical, 1.1% human vs 1.2% AI. my theory is that follow ups are more about timing and persistence than copy quality and the humans were fine at writing short bumps.

the other thing, and this is where it gets interesting for anyone managing a team, the AI group was significantly faster at producing sequences. group A spent on average about 4 hours per new sequence (research plus writing plus internal review). group B was averaging about 90 minutes. thats a massive time savings and it meant the AI group could test more variations. they ran 14 distinct sequences in the 3 months vs 8 for the human group. more sequences means more data on what resonates which compounds over time.

now the negatives because there are real ones.

the AI copy had this tendency to sound... samey after a while. by month 2 i could read a cold email and tell you immediately whether it was AI drafted even after editing. there was this pattern of "noticed [company] is [doing thing], curious if [pain point]" that kept showing up in different variations. we had to actively fight against that by feeding in competitor examples and telling the model to avoid certain structures. our VP sales actually flagged this independently, he said the emails were starting to feel like they came from a template factory which... yeah.

the other issue was personalization quality. the AI could pull in surface level personalization really well but it would sometimes make connections that didnt actually make sense. one email referenced a prospects company "expanding into european markets" based on a job posting for a UK role, but the company was already operating in 12 countries across EMEA. the prospect actually replied to tell us we clearly hadnt done our homework. stuff like that happened maybe 4-5 times over the 3 months which isnt a lot percentage wise but each one felt bad.

also, and this drove me crazy, getting the meeting data synced properly into Salesforce was its own nightmare. we tag campaigns by sequence type and the custom fields we set up to differentiate AI vs human kept getting overwritten by our Smartlead integration. spent probably 6 hours across the quarter just fixing attribution data. i swear half my job is fighting Salesforce integrations and the other half is pretending i enjoy it.

wait i should mention cost breakdown on the AI side specifically. we were spending about $120/mo on Claude team plan, roughly $80/mo on ChatGPT plus for 2 seats, and then maybe $40-50/mo on various API calls for the custom GPT. so call it $250/mo in AI tooling. thats pretty cheap relative to the lift we got. the human group obviously had zero AI costs but their time cost was higher and they produced fewer meetings.

the conclusion im landing on, and we've since rolled this out to the full team, is that AI should be writing first drafts of everything. no exceptions. but you need a human doing a real edit pass, not just skimming it and hitting send. the reps who performed best in group B were the ones who used AI as a starting point and then rewrote 30-40% of each email. the ones who basically sent the AI output with minor tweaks had good numbers early but they degraded by month 3 as the copy got stale.

one more thing. we tested this exclusively on cold outbound to net new prospects. i have no idea if these results hold for re-engagement sequences or inbound follow up or anything else. our ACV is $98k so the math on spending time on copy quality makes sense. if youre selling a $2k product the calculus might be totally different.

since rolling it out to all 6 reps in april our reply rates have been sitting around 4.2-4.4% which is below the test group's 4.7% but well above our historical 3.1%. good enough. the time savings alone would have justified it even if reply rates stayed flat honestly.

anyway thats the data. take it for what its worth, n=22k emails across one company in one vertical selling one product at one price point. not exactly a huge sample but its real money and real pipeline so it meant something to us

submitted by /u/Jhingalalaa
[link] [comments]
❌