Nearly all A/B testing advice for cold email was written for a volume you do not have. It assumes tens of thousands of sends, which means it assumes a list you bought, which means it assumes results you would not want.

At the volume a warmed mailbox can sustain, which is somewhere around fifty a day, the statistics work very differently. That does not make testing pointless. It makes it a different activity, with different rules, and the rules are simple enough to hold in your head.

What you can and cannot detect

The core constraint, without the mathematics: small samples can only see large effects.

At a few hundred sends per variant, a difference of one or two percentage points in reply rate is indistinguishable from luck. You will see such differences constantly, because randomness is lumpy, and if you act on them you will spend six months confidently walking in circles.

Sends per variantWhat you can actually detect
100A doubling. 4% versus 8% is visible. 4% versus 5% is not.
300A large difference. Roughly 50% relative change.
1,000A meaningful difference. Roughly 25% relative change.
5,000+The small stuff everyone else writes articles about

The practical rule that follows: if you need a calculator to see which variant won, you do not have a winner. At this volume the real results announce themselves.

Test the reason, not the wording

Because you can only see large effects, you must test large things. That rules out most of what gets called A/B testing.

Here is the order I would test in, from largest effect to smallest.

1. The segment. Not an email test at all, and the highest leverage thing on the list. The same message sent to HVAC contractors and to dental practices is two different experiments, and the gap between them is usually enormous. If your reply rate is under 2%, this is almost always the variable, and no amount of rewriting will rescue the wrong audience.

2. The reason you are writing. Not the phrasing of the reason. The reason. "I noticed you have no online booking" and "I work with three other firms in your area" are different emails in the sense that matters. This is the test that most often produces a doubling.

3. The ask. A fifteen minute call, a yes or no question, a link to a two minute video, or nothing at all beyond a question. The size of the ask changes reply rate more than anything about the prose, and the smallest ask usually wins by more than people expect.

4. Length. Sixty words against a hundred and eighty. This is a genuine structural variable and worth one test, once.

5. The follow-up count and spacing. Technically a sequence test rather than an email test, and frequently the one that moves total replies most, because a large fraction of replies arrive after the third touch. The follow-up sequence has the numbers.

6. Subject line. Last, and expect little. More on why below.

Why subject line testing broke

Subject lines used to be the default first test because open rate was easy to measure. Open rate is no longer a measurable quantity.

Mail clients now pre-fetch remote images on the recipient's behalf, including tracking pixels, for messages nobody opened. Your reported open rate is inflated by an amount that varies by provider, by device and by settings you cannot see. Comparing two variants on a metric corrupted by an unknown, non-uniform amount is not a measurement.

So subject lines have to be judged on replies, and their effect on replies is small next to whether the email deserves one. Write a plain, honest, specific subject, and spend the attention on the first line instead. Cold email subject lines: what works now that open rates lie is the longer version.

A protocol that actually works at low volume

  1. Pick one variable from the list above, as high up as you can bear.
  2. Write two genuinely different versions. If a stranger reading both would call them the same email, you have not written a test.
  3. Split randomly, not by segment. Alternating rows in a list sorted by city is not random, it is a geography test wearing a costume.
  4. Send both across the same days. Not variant A this week and B next week. Weeks differ, and you will have tested the week.
  5. Stop looking. Check at 100 sends per variant, not before. Daily checking is how you get excited on day two about a difference that has vanished by day five.
  6. Judge on replies, and split them into positive replies and total replies. A variant that doubles total replies while halving positive ones has made things worse, and you will only see that if you count both.
  7. Write down what you learned, including the failures. Two years of tests where you remember only the winners is not a body of knowledge, it is a highlight reel.

The result nobody wants and everybody needs

Sometimes both variants perform identically and both perform badly. This is the most useful outcome available and almost always the one that gets discarded.

It means the variable you tested is not the problem. Stop refining it. The answer is one level up: a different segment, a different offer, or a genuinely different reason for the email to exist. I have watched teams run eleven subject line tests on a campaign whose real problem was that they were writing to people who did not have the problem they were solving.

What to do with a winner

Two things, and most people only do the first.

Roll it out, obviously. And then write down why you think it won, in one sentence, before you forget. "Naming the specific gap beat naming the industry" is a portable lesson. "Version B won" is a fact with no future in it.

The second thing is what turns a year of tests into judgment. It is also what makes the next campaign in a new niche start from a hypothesis rather than from scratch.

Where the automation fits

The mechanical parts of this (splitting randomly, holding the split across follow-ups, stopping a sequence the moment somebody replies, keeping the daily volume inside what your mailbox can sustain) are exactly the parts humans do badly by hand. That is what the A/B testing in smart sequences exists for: not to have an opinion about your copy, but to make sure the experiment you designed is the experiment that actually ran.

The judgment stays yours. So does the part where you test the segment before you test the subject line, which is the only advice here that reliably doubles anything.

Related: cold email benchmarks for what a normal reply rate looks like before you decide yours is broken, and 8 cold email templates that get replies for structures worth putting into a test.