Est. reading time: 5 minutes
The default way to A/B test subject lines is to write three variations of the same idea, split them across the list, and crown whichever one edged out the others by half a point of opens. That process feels rigorous and teaches almost nothing, because the variants didn’t isolate a variable, the sample was sized by vibes, and the winning metric is one that privacy features have been quietly inflating for years. Subject line testing in Mailchimp can be a genuine compounding lever, but only run as actual experiments, one question per test, samples sized on purpose, and a winner rule that measures something real.
Every test answers one question
A subject line is a hypothesis, not a slogan, and the test design starts with what you’re actually asking. “Increase opens” is not a question. “Does a direct benefit beat curiosity for this audience” is, and so are urgency versus authority, personalization versus none, emoji versus without. Draft two variants that isolate exactly one of those variables, meaning genuinely different concepts rather than synonym swaps, and let the send decide what your loudest internal opinion couldn’t.
Sequence the questions by size of lever. If your list has never seen benefit-led lines tested against curiosity hooks, that’s the first duel, because the answer will shape every subject line after it. Punctuation tweaks and word swaps are real tests, but they’re refinements of a dominant pattern you haven’t identified yet, so they wait. We keep a running view of which patterns are earning opens across accounts in our breakdown of the subject lines that work and why, which is a reasonable menu of first hypotheses.
And hold everything else constant, especially the preheader. It’s tempting to give each variant its own matching preheader, but then you’re testing subject-and-preheader bundles, and whatever wins, you won’t know which half did it. One variable per test is the whole method. Everything else is decoration on guessing.
Two variants, right-sized samples, honest metrics
Keep it to two variants even though Mailchimp allows three. Every additional variant splits the same sample thinner, which means slower significance and murkier reads, and two sharply different concepts produce clearer signal than three shades of one idea. This is also where the send-waste hides, since an inconclusive three-way test spent your list’s attention and bought nothing.
Size the sample deliberately. Our working allocations are 20 to 40 percent of the list combined for lists under 10,000, 10 to 20 percent from 10,000 to 100,000, and 5 to 10 percent above that, with the underlying goal being several hundred opens per variant reasonably fast. Under that, you’re crowning winners on noise. Far over it, and the test consumed the audience the winner was supposed to be sent to.
Then pick the winner rule like it matters, because it’s the most consequential setting in the campaign. Open rate is the intuitive choice for a subject line test and the wrong default, since Apple’s Mail Privacy Protection auto-registers opens for a large share of most consumer lists, meaning an open-judged test partly measures how your Apple users were distributed across variants. Click rate is the safer standard rule even for subject lines, because a subject that drives more clicks drove more real opens by definition. And with a store connected, total revenue is the gold standard for purchase-intent campaigns, for all the reasons we covered in tracking sales from your Mailchimp campaigns.
Duration is the last setting, and it should match audience behavior rather than impatience. Three to six hours generally balances speed and representation for a domestic list, six to twelve for a global one, launched when your audience is awake and scheduled so the winning send lands at peak engagement rather than whenever the timer happens to expire.
Let the automation do the handoff, then keep what you learned
Mailchimp’s A/B testing campaign type handles the mechanics. Choose subject line as the variable, enter the two variants, set the test percentage and duration, pick the winner rule, and the winning subject goes to the remainder of the list automatically, no manual scramble at hour four. The tooling is the easy part, which is exactly why the tooling was never the problem.
The compounding comes from what happens after. Log every test somewhere permanent, the winner, the losing angle, the segment, the metric, because an untracked test is a lesson your team will pay to relearn in six months. Patterns that win repeatedly get promoted to control status, and each new campaign challenges the control with a single new angle rather than starting from blank creative, a ladder that improves the baseline steadily without ever gambling the full list on an unproven idea.
Two boundaries keep the whole practice healthy. Don’t test constantly on a small list, since a few thousand contacts can’t feed weekly experiments with meaningful samples, and a monthly well-designed test beats four noisy ones. And remember that testing spends attention, which is a finite resource sitting on top of your sender reputation, the foundation we covered in why your Mailchimp emails end up in spam. Run fewer, sharper duels, judge them on clicks and revenue, bank every result, and subject lines stop being the part of email you argue about and become the part you know.









