
A/B testing
Making two versions of something is the easy part. Deciding what the difference between them is allowed to be is the work.
The short version
Is this for you? For tests that changed six things at once and taught you nothing. Not yet if the channel has too few impressions to separate a winner from noise; on a small channel an early result is a hint to watch, not a verdict.
What leaves your plate. The design of the test, one stated question per comparison, a single variable held between the versions, and the covers and titles built with the video rather than bolted on once it is done.
What starts working. Three covers and three titles with every video, each comparison asking one clear thing, so the next cycle begins from an answer instead of a preference.
The smallest way in. A free audit: three fields, and a written reading back, usually within about three working days. It commits you to nothing.
What to read next. The five levels to see where this fits, or the casework to see it done.
What the difference between them is allowed to be
Concept first, then the single variable
Three covers in three genuinely different concepts, each carrying its own title, so it is three whole packages to choose between rather than nine combinations to untangle: one showing a face, one showing only the finished thing, one showing the ingredients on the counter. That comparison answers which idea an audience wants. Changing the typeface at the same time as the subject answers nothing at all, so the narrow test comes second, on the concept that won, with exactly one difference left in it.
Run as a native test where the channel can, by hand only where it cannot
YouTube's own tool runs up to three variants on one video at the same time, titles, thumbnails, or the two together, and settles the winner on share of watch time, not on click-through alone, which is the honest measure: a cover that wins the click and loses the viewer has won nothing. Where a channel is eligible, that is where the covers go. A sequential swap, this cover this week and that one next, is not a controlled test, because the week changes underneath it, so we use it only where the native tool cannot reach, on older uploads, or on a question the tool was never built to answer, and we say which one a result came from. The tool also needs a fair volume of impressions before its verdict separates from noise, so on a smaller channel an early result is treated as a hint and watched rather than declared a winner.
Two voiceovers, and a decision about how to compare them
Two alternative reads over the same cut can be compared two ways: two versions of one video, or the same difference carried across two different videos. Each answers a slightly different question, and each costs something the other does not. The decision is taken before the recording session, because a voiceover recorded without knowing which comparison it belongs to is usually useless to both.
Testing is what a channel does before it commits
Where a brand has not yet settled its audience or its subject, putting everything behind one topic is a bet dressed up as a strategy. The route that spends less is a run of quick, budget-aware pieces built to be compared against each other, and read together. Once the building blocks are known, bespoke production finally has something to be bespoke about.
Variants arrive with the video itself
On this service the three covers and the three titles arrive with the video, so a comparison is never a separate piece of work somebody has to approve first. Making a comparison free at the point of use is why it actually gets run, because anything that needs its own conversation gets skipped in a busy week. The variants exist. Using them is a click.
The shape of a test
| Covers per video | three, each a concept with its own title: three packages, not nine combinations |
|---|---|
| Where it runs | up to three at once in YouTube's native tool, titles, thumbnails or the two together, where the channel is eligible; by hand where it is not |
| How a winner is judged | share of watch time in the native test; click-through and retention read together where it is manual |
| What differs in a concept test | subject, title, colour and type, all at once, on purpose |
| What differs in a narrow test | exactly one thing |
| Voiceover comparison | two versions of one video, or one difference across two videos |
| Agreed before anything is made | the question, and what each possible answer would change |
The two variables worth spending a cycle on
What gets tested first
The cover and the title, because they are the only parts of a video that everybody sees. Hooks come next, chosen against retention rather than against taste, from openings that were filmed for the purpose. What follows depends entirely on what the first two answered.
Where the answers go
Into the monthly report beside the video they came from, and into the next plan as a decision somebody has to act on. A result that changes nothing was a comparison nobody needed, which is why the question and its consequences are agreed before anything gets made.
What is not settled this way
The ideas themselves. Numbers can choose between two options that were both worth making; they cannot generate the option. A channel run entirely on what tested well last quarter slowly turns into everybody else's channel, and the graph will be perfectly happy about it.
Agree what the result would change before anything is made, or the test is a survey nobody acts on.
A test nobody will act on is a preference with a chart attached.
What changes for you
Three covers and three titles with every video, one stated question behind each comparison, and a next cycle that begins with an answer instead of a preference.
Related: Thumbnails · Hooks and the first seconds · Analytics and reporting · Content systems and planning
Before strategy is discussed, we read the channel: what it publishes, what it promises and where the two part ways. The audit comes back in writing, so the first call starts from evidence rather than introductions.
Request your free audit.
Three fields, and a written reading back: a fit snapshot, two or three prioritised opportunities, and a recommended next step.
A person reads the channel and writes the audit by hand: a considered read typically takes three working days. That is the usual shape, not a promised turnaround. We use these details only to reply to you: no lists, no lurking.
What you will get
A fit snapshot: where your channel stands, and whether we are a match.
Two to three opportunities: specific, prioritised, yours to keep.
A recommended next step, even if that step is not us.
The audit is free and commits you to nothing: nobody follows up with a call you did not ask for.
The casework shows all of this done.
The case studies are the same work with the numbers, the sources and the dates attached.
See the casework →