Run paid ads and scale lead flow 22 of 23 in this group
SOP 220
Run two-round concept testing and promote to control
What this page is for. Use it to put split testing on a weekly routine. It carries how often one team tests and where the results come up in its weekly review, how the step to work on is picked, the two rounds of versions (three very different concepts, then three variations of the one that won), the stages a new version passes through before it replaces the one you run now, and the organic form of the same idea on Instagram. Under SOP 209 — Pick a lead bucket, starting with the offer and lead magnet, split tests on every component close the list of fixes for outreach (its section 1.3); this page is the routine for running them.
SOP-220-Run-two-round-concept-testing-and-promote-to-control.md
1. Where it applies
The topic is split testing in general. The claim is that it truly works in all three methods: content, ads and outreach.
The view here would be that it suits outreach and ads more specifically, because the inflow there is more consistent. Content can spike, and what kind of traffic arrives can shift depending on which piece went viral.
What follows is one team's own testing cadence: the way it does it in-house. The stages of outreach itself are set out on SOP 217 — Run outreach as an assembly process.
2. The weekly cadence
That team runs split tests each week, usually one or two of them.
In the view here, split tests need to come second on your marketing agenda. The first item is the lead count: how many leads came in, how many of them qualify, and where the main KPI stands. The next item, week after week, is which split tests ran and which ones won.
For that team this happens on Mondays, at a weekly marketing meeting.
How many tests to run at once, set against SOP 41, is in section 7.
3. Find the point of greatest leverage
The question they ask is where the point of greatest leverage sits. Typically, that is the step in the flow whose rate is lowest relative to the benchmark for that step. Looking at the percentage for each step, whichever is lowest against its benchmark is the one they go after.
This page gives no figure for the benchmark of any step.
The question that follows is what a good alternative would be to test there.
4. Two rounds: three concepts, then three variations
This testing method is judged very good. On this account it works for paid ads, and it works for thumbnails and packaging on YouTube just as well. Whether organic titles and thumbnails should be picked by click rate is flagged in section 7.
Round one. Make three versions, each wildly unlike the others: three completely different concepts for the thing under test, whether that is a home page, an opt-in, a headline or anything else. Run those three as the first test. One way to make it stick is three different fruits: an orange, an apple and a banana.
Round two. Once you know which of the three won, make three smaller variations inside that one. In the fruit picture, if the apple won, the next test is a green apple, a red apple and a yellow apple.
Across the two rounds that makes a grid of nine: three by three.
5. Why a new version earns its place in stages
Telling someone to test everything imaginable is easy. The trouble is that testing uses up resources. A test costs the traffic you spend running it, plus the sales you would otherwise have made on your control.
So, in a practice that applies to that team, a new version earns its place in stages, in this order:
- First it is tested on a percentage of the traffic, not on all of it.
- If it wins, it then goes into a true 50-50 split test.
- If it holds up there, it becomes the control.
The view here is that this is simple, and that it works great.
This page gives no figure for the share of traffic the first stage uses.
SOP 41 — Run one test a week, section 3, settles whether a before-and-after difference is real with a significance calculator. This page gives no figure for when a version has won or held.
6. Typically two tests a week, at step one and step two
Typically, that team has one test running at step one and one at step two each week. This is plainly no rule of law, only a trend noticed in that team; it applies to that team.
This is for every kind of marketing (outreach, ads and organic), though most of it will be split testing specific to a funnel.
7. What this page does not decide for you
- How many tests to run in one week. Here, that team usually runs one or two, and typically has one at step one and one at step two, offered as an observed trend, not as a rule. SOP 41 sets the cadence at one test per week per platform, does not allow testing more than one thing inside the same pathway in the same week, and among its reasons gives that split tests interfere with each other. Read as two steps of one flow (the third item below), that team's two weekly tests sit inside one pathway. SOP 41, section 7, flags the same split from its side. That team's way works on two steps each week; SOP 41's way tells you cleanly what each change did, at one change a week. This page does not settle which reading is right.
- Which step to work on. Here, the step lowest relative to its benchmark. SOP 41, section 1, adds the same number of points to every step and takes the step with the largest resulting multiplier, and works from the front of the pathway to the back (its section 6). This page's way needs a benchmark for every step, and none is given here; SOP 41's way needs none, but always points at the step with the lowest rate. This page does not settle which reading is right.
- What step one and step two are. One reading: two steps of the flow, as in section 3, where the step to attack is chosen. The other: the two rounds of section 4, since the first round has just been called step one. This page does not settle which reading is right.
- What counts as a win, and what counts as holding. Not established on this page.
- Whether to pick organic titles and thumbnails by click rate. On section 4's reading, the two rounds work for thumbnails and packaging on YouTube as well as for ads. A second view, from one team: it is moving away from picking titles and thumbnails by the best click-through rate, because the most clicked version appeals to the widest crowd rather than the people a business wants, and the platform then serves the piece to whoever clicks most. That team also finds such test data not always accurate: at times it may lead you to optimize the wrong thing. It gives thumbnails less weight and often takes the automatic one. Its conditions: a media business paid on impressions should play for clicks; if only ideal customers saw a piece, its click rate would matter; and it still tests its main show's packaging at the start, and lets software keep whichever of three generated titles draws the most clicks on each clip. The same question is flagged on SOP 276, section 11. This page does not settle which reading is right.
8. The organic version: trial reels
On Instagram, the organic version of this is trial reels, which are called wildly underused. At one point the platform had just added a limit: five permutations of a trial reel a day.
- Make sure the reels obviously differ from each other.
- The way to do that: keep one meat and give it five different hooks.
- Expect results that differ wildly. They show how much the first three seconds really carry.
Why the hook carries most of the weight in an ad, and how one winner becomes many versions, is on SOP 212 — Make better ad creative with hooks, permutations and assembly, sections 1 and 2. How many hooks to record for each kind of content is SOP 122, section 5. Making hooks apart from the meat is SOP 119 — Assemble hook, meat and call to action. The opening seconds of a video get their own treatment on SOP 34 — Front-load the opening seconds.
9. The checklist
| Step | What to do |
|---|---|
| 1 | Each week, review which split tests ran and which won, second after the lead count, qualified leads and the main KPI, in the view here (one team: a Monday meeting) |
| 2 | Find the step whose rate is lowest against its benchmark: typically the point of greatest leverage |
| 3 | Round one: three wildly different concepts, run as one test |
| 4 | Round two: three smaller variations inside the winner, for nine in all |
| 5 | As one team does, test a new version on a percentage of the traffic first |
| 6 | If it wins, run a true 50-50 split |
| 7 | If it holds, make it the control |
| 8 | As one team typically does, one test at step one and one at step two a week: an observed trend, not a rule |
| 9 | Organic, on Instagram: trial reels, one meat with five hooks, inside the daily limit in force at the time |
10. What this page does not cover
Testing for statistical significance is SOP 41 — Run one test a week, section 3. Sample size is not covered on this page.
Terms defined on this page
- Control
- The version you are running now. A new version has to beat it to take its place.
- KPI
- A key performance indicator: a main measure someone is held to. At the weekly marketing meeting it is reviewed first, with lead counts and qualified leads; at 50 to 100 people, such measures are set so someone answers for each result.
- Promote to control
- A new version proves itself in steps: first on a share of traffic, then in a true 50-50 split, then it becomes the control if it holds.
- Split testing
- Running versions side by side, often small budgets across many variations, to see which does better: find the best hook first, then the best value piece behind it, then scale the winners. It needs tracking; in the view here it suits steady inflow, such as outreach and ads, more than content.
- Two-round concept testing
- First test three wildly different ideas (an orange, an apple, a banana); then test three smaller variations of the winner (a green, red and yellow apple), nine versions in all.