Performance Creative
Creative is a testing discipline, not a taste contest
We run the testing system. Your creators, editors and studios keep making the work, and every asset finally comes back with an answer attached.
Performance creative testing is the system that decides what to test, isolates a single variable, sets the rule for killing a loser, and sets the rule for scaling a winner. It is not creative production. Your creators, editors and studios keep making the work. The testing system turns what they make into evidence you can act on.
You are not short of creative. You are short of feedback.
Most ecommerce brands spending real money on paid media are already producing plenty of assets. Creators are shooting, editors are cutting, and new work reaches the ad account every week. Volume is not the constraint. The constraint is being able to say, at the end of a month, which specific decision moved the number and what should be made next because of it.
Three things break that feedback loop, and almost every account we look at has at least two of them running at once.
Nothing is isolated. The new ad changes the hook, the visual, the format and the offer at the same time. When it wins you cannot say why, so you cannot repeat it. When it loses you have thrown away four ideas instead of one, and one of the four may have been the good one.
Nothing has a kill rule. Assets stay live because nobody agreed in advance what failure looks like. Budget accumulates on ads that were never going to work, and the decision to switch them off eventually gets made on a Friday afternoon by whoever happened to open the report.
Winners are never documented. A hook works in March. By August the person who briefed it has moved on, nobody remembers what made it different from the eleven that failed beside it, and the brand buys the same lesson a second time at full price.
The outcome is a brand that produces constantly and learns almost nothing. Creative stops being a testing discipline and becomes a taste contest, settled by whoever holds the strongest opinion in the room rather than by evidence from the account.
| Decision point | Testing that teaches | Testing that does not |
|---|---|---|
| Before launch | A written hypothesis names the belief and the metric it should move | A batch of new assets goes live because the last batch got old |
| Variant design | One variable changes. Everything else is held constant | Hook, visual, format and offer change together |
| Turning things off | Kill thresholds are agreed before any spend runs | Ads run until somebody notices them in a report |
| Handling a winner | The winner is iterated on purpose and the reason it won is recorded | The winner is duplicated into more ad sets and left alone until it dies |
| Six months later | The library explains what works for this brand and this buyer | The team is re-testing ideas it already tested last year |
The anatomy of a testing sprint
A sprint is the unit of work. It has four parts. Skipping any one of them turns the sprint back into activity that happens to cost money.
1. The hypothesis
Every sprint opens with a written statement of what you believe and why it would matter. Not “test more UGC”. Something closer to this. We believe an opening three seconds that leads with the problem rather than the product will lower cost per click on cold audiences, because cold conversion rate sits far below warm and suggests new buyers cannot place what category we are in.
A hypothesis does two jobs. It forces someone to commit to a belief before the data arrives, which is what makes the result informative rather than decorative. And it makes a losing result useful, because a belief you have disproved narrows the field exactly as much as one you have proved.
2. Variants that isolate one variable
The variant set is where most creative testing quietly fails. If the control and the test differ in more than one respect, the test cannot answer anything. You get a winner with no explanation, which feels like progress and compounds into nothing.
So the sprint names the variable. Hook, opening frame, format, creator archetype, proof type, offer framing, length, caption angle. One of them moves. The rest are held constant, in writing, so the producer knows what not to improve. That last part matters more than it sounds. A talented editor will instinctively make everything better at once, and a better ad that teaches nothing is still a dead end for the next sprint.
Testing budget also has to be large enough to separate signal from noise at your conversion rate and order value. Underfunded variants produce results that reverse themselves the following week, which is worse than no test, because the team acts on them.
3. Kill criteria
Kill criteria are agreed before launch and expressed as a threshold, not a feeling. A typical rule combines a spend floor and a leading indicator, so nothing is judged before it has had a fair chance and nothing runs on hope after it has had one.
The point of writing it down first is that it removes the argument. Nobody has to defend killing an asset that the founder loved, because the rule was set when everyone was calm and neutral about the outcome.
4. Scale criteria
Winners need a rule too, and this half is skipped more often than the kill half. A winner has to earn its way into more budget by clearing a threshold that is tied to your margin, not to platform ROAS. It also needs a defined next step. Iterate the winning variable further, port the concept into a second format, or hand it to the media plan as a durable asset.
Here is a sprint written out. The thresholds below are illustrative. Real ones are set against your own contribution margin, order value and payback window.
| Sprint element | Written out |
|---|---|
| Hypothesis | Cold audiences respond better to a problem-led opening than a product-led opening |
| Variable | The first three seconds only |
| Control | Current best performing asset, product-led opening, unchanged |
| Variant A | Same asset, same creator, same edit, opening reshot as a problem statement |
| Variant B | Same asset, same creator, same edit, opening reshot as a question |
| Held constant | Creator, format, length, music, captions, offer, landing page, audience, placement |
| Kill criteria | Below a set spend floor with no purchases, or cost per acquisition above the agreed multiple of target |
| Scale criteria | Clears target cost per acquisition on new customers at the spend floor, then moves into the core budget |
| Documentation | Result, decision and reasoning written into the library the same week, win or lose |
What a weekly testing calendar looks like
Sprints only compound if they are scheduled. The testing calendar is a rolling document that sits next to the media plan, not beside it in spirit and in a different file in practice.
It holds a fixed share of spend for testing, agreed in advance, so the test budget survives a bad week instead of being the first thing cut. It names which sprint is in market now, which one is being briefed, and which one is queued behind it. It sets a weekly readout where every live test is called: kill, hold or scale. And it maps launches against what the business is actually doing, so a promotional period is not the week you choose to test a new hook.
The cadence also protects your producers. Briefs arrive on a predictable day with a predictable shape, which is what allows a creator or an editor to plan capacity instead of reacting. Testing calendars that move every week produce late assets and defensive teams.
This is the same rhythm that governs the account itself, which is why the calendar is built alongside the paid media plan rather than after it. Creative supply and budget pacing are one problem, not two.
If you want a straight read on where your current testing is losing information, that is what the free 30 minute growth audit covers. You get the findings whether or not you ever work with us.
The winners and losers library, and why documentation compounds
Every sprint ends by writing a row. Date, hypothesis, variable, variants, spend, result, decision, and the reasoning behind the decision. Losers get the same treatment as winners, because a documented loser is the cheapest way to stop a future team from paying for the same idea twice.
After a quarter the library starts answering questions the account cannot. Which hook families work on cold traffic for this brand. Which creator archetypes convert and which only generate clicks. Which proof formats hold up past the first week. Which offers have been tested to exhaustion. That knowledge is the asset. Individual ads decay. The pattern behind them does not.
It also survives turnover. When a media buyer leaves, or a creator moves on, or you change agencies, the reasoning stays with the brand instead of walking out of the building. Most brands cannot answer a simple question about their own history, which is why year two of their creative program looks very much like year one.
Compounding shows up in the numbers eventually. In one athletic apparel account, sales rose 35.7% to $9.27M across the first six months of 2026 while paid media investment rose 42%, and blended ROAS held at 3.36x. Absorbing that much additional spend without efficiency collapsing is a creative supply problem as much as a bidding problem, and a tested pipeline is what makes the supply dependable. More detail sits in the case studies.
How direction briefs work with your producers
Plaid Testing does not shoot, edit, design or staff creative. Production runs through the people you already have, whether that is an in-house team, a roster of creators, a studio, or a mix of all three. We are the layer that tells them what to make and why.
A direction brief carries the hypothesis, the single variable in play, the list of things to hold constant, the format and placement specs, the control it will run against, and the decision rules it will be judged by. It is deliberately narrow on the variable and deliberately quiet everywhere else. Your producers keep the craft. They just stop guessing about the objective.
Two practical notes. If your production capacity is not sufficient to feed a weekly cadence, we will say so plainly and tell you what capacity the plan needs, since we have nothing to sell you on that side. And if your creators are already good, this usually makes them look better, because their work finally gets judged against a fair test instead of a crowded one.
Creative testing is a media decision and a margin decision
Creative is the largest lever on cost per impression, click through rate and cost per acquisition in a mature account. Once structure and bidding are reasonable, most of the remaining performance difference between two accounts of similar size is what is inside the ads. That makes creative testing a media discipline, not a brand exercise sitting off to one side.
It only works if the scoreboard is honest. A kill rule pointed at platform reported ROAS will kill assets that were working and scale assets that were not, which is why testing sits downstream of a clean measurement setup. Server side tracking, deduplicated conversions and new customer economics are the prerequisite, covered under tracking and attribution.
Judged properly, the effect lands on new customer economics rather than on blended vanity numbers. In one womens fashion account, sales rose 99% while new customer cost per acquisition fell 21%, new customer ROAS rose 58%, and net profit rose 136%. Those are the four numbers worth reading together, because two of them can move the wrong way while the top line still looks impressive.
Where creative sits inside the wider plan, what the brand should claim, and how testing supports it are leadership questions as much as media ones, which is why this service is often run alongside fractional CMO engagements rather than in isolation.
Ad fatigue and refresh cadence
Fatigue and saturation get treated as the same problem and they are not. Fatigue is a creative problem. The same people have seen the same asset too many times, and new creative fixes it. Saturation is an audience problem. You have already reached everyone available, and new creative will not help because there is nobody new to show it to. Brands misdiagnose the second as the first regularly, then spend a quarter making ads that could not have worked.
The signals worth watching separate the two cheaply.
- Click through rate decaying while cost per impression holds steady points at fatigue.
- Frequency climbing inside your consideration window points at fatigue.
- Unduplicated reach flattening or falling while spend rises points at saturation.
- Cost per person reached doubling over a scaling period points at saturation.
Refresh cadence follows from the diagnosis. A fatigued winner is usually worth iterating rather than replacing, since the concept is proven and only the execution is tired. New hooks, new openings, new creators inside the same idea. A saturated account needs new audiences or new geographies first, and creative volume alone will not rescue it. Setting a refresh rhythm before performance decays is cheaper than reacting after it does, and the testing calendar is what makes that rhythm possible.
Who this fits
This suits ecommerce brands investing $50K or more per month in paid media, most often in apparel, athletic wear, fashion and accessories, and wellness and supplements. It fits best when you already have production capacity you trust, whether in-house or contracted, and what is missing is the discipline that turns that output into knowledge.
It is a poor fit in two situations. If your core problem is the offer or the positioning, testing will document the problem faithfully and not solve it. And if what you want is a partner to make the ads for you, that is a production relationship, and a studio or a creator roster is the right answer instead of this.
See what your last 90 days of creative actually taught you
The growth audit is 30 minutes, costs nothing, and ends with three specific fixes you can implement whether or not we work together. Bring your ad account and your last quarter of creative, and we will show you where the feedback is leaking. Book your free growth audit.
Do you produce the creative?
No. Plaid Testing runs the testing system and hands direction briefs to your producers. Your creators, editors and studios make the assets, keep the craft decisions, and keep the relationship. We decide what gets tested, how the variable is isolated, when an asset is killed and when it is scaled. If your production capacity cannot support the cadence, we will tell you what capacity the plan needs.
How long should a creative test run before we call it?
Long enough to clear a spend threshold set against your order value and conversion rate, not a fixed number of days. Most brands call tests too early on too little spend, then act on results that reverse the following week. The threshold is agreed before launch alongside the kill and scale rules, so nobody is negotiating with a live ad and a nervous founder.
Can this run alongside our existing media buyer or agency?
Yes, and it often does. The testing system needs access to the ad account, agreement on the testing budget, and a weekly slot to call every live test. Media buyers usually welcome it, because it replaces the standing request for more creative with a queue of briefs and a documented reason for each one. Where we also run the media, the calendar and the plan are built together.
How many new assets do we need each week to make this work?
Fewer than most brands assume, if they are the right ones. A sprint that isolates one variable needs a control and two or three variants, and several of those variants are edits of existing footage rather than new shoots. What matters is that assets arrive on a predictable schedule, since an irregular supply is what forces teams to test four changes at once.
Get your free 30 minute growth audit
Actionable takeaways, guaranteed. No retainer required to start.