Ad Creative Pipeline: How We Test Creatives Half Automatically

For a year we built ad creatives on gut feel. Then we started counting, and the gut was wrong more often than right.

Since then we run a pipeline. It does not produce creativity on demand. It makes sure one idea reliably turns into many tests, and that someone can read the numbers at the end. We run three brands, nano, mate and MUSTAX, with two people. Without this process, paid social would be out of reach for us.

Here is the full flow, including the places where a human still sits.

Why we built a process in the first place

Before, it went like this: somebody had an idea. A creative got made. It ran for two weeks. Afterwards we argued about whether it was any good.

The idea was fine. Three other things were the problem.

One: too few tests. We managed a handful of creatives per week. At numbers that small, every result is chance.

Two: nothing was comparable. Two creatives differed in image, copy and format all at once. When one wins, you don't know why.

Three: no traceability. The files were called "final_2_new". The ad account showed different names. After four weeks nobody could say which ad belonged to which idea.

Automation solved exactly two of those: volume and traceability. The idea stays manual, and that's how it should be. Where else we deliberately leave automation out is covered in when automation makes no sense.

Step 1: collect building blocks instead of hunting for ideas

The most important change was a mental one. We stopped looking for finished creatives. We collect building blocks.

For us a creative consists of four parts:

These blocks live in a spreadsheet. Every entry has an ID, a short description and a status.

Where the hooks come from is the underrated part: support requests, reviews and customer emails. At nano we read the support mail in a structured way anyway, because around 65 percent of requests get answered automatically there. The same pot of data gives us the sentences customers actually use.

An example from practice: at nano, the same sentence kept showing up in reviews, roughly "finally a brush that doesn't upset my gums". We tested that sentence almost word for word as a hook. It beat the variants we wrote ourselves. It was a real sentence, and that is what did the work.

Step 2: generate variants

Now comes the part the machine takes over. The building blocks turn into a combination list.

We pick a handful of hooks, a handful of visuals and one or two offers. A workflow in n8n builds the combinations and creates an entry for each: ID, copy, visual used, target format.

The copy is partly written with AI, under the same rules as our product copy: with a brief, examples and a human who signs it off. How we set that up is in AI product copy that doesn't sound like AI.

The image work is only partly automatic. Crops for the different formats, text fields and simple variants come out of templates. Shooting a new visual stays a real job with a camera, a product and lighting.

The rule behind all of it matters: per test round we change exactly one layer. Same visual, different hooks. Or same hook, different visuals. The moment two layers change at once, the result becomes unreadable.

Step 3: naming, or none of it counts later

This sounds like bookkeeping and it is the point where most pipelines die. When the names are off, you cannot evaluate anything afterwards.

Here every variant gets an ID built from its blocks. It sits in the file name, in the ad name and in the spreadsheet row. Always identical, always set by machine.

Building blockExampleAppears in the name as
Brandnanonano
Hookgum sentence from a reviewh12
Visualhand holding brush, white backgroundm04
Offerrefill pack bundlea02
Formatvertical videov9x16

Those parts produce a name like nano_h12_m04_a02_v9x16. Ugly for humans, perfect for evaluation.

The effect: later we can filter for h12 and see across all visuals whether that hook carries. Without a scheme you would have to open every ad one by one and rely on memory.

Never assign the names by hand. People shorten things, mistype and invent special cases. The workflow doesn't.

Step 4: evaluate, and not daily

We pull the numbers from the ad accounts automatically and place them next to our combination list. The report runs once a week and lands where we work anyway. Nobody opens a dashboard for it.

We evaluate on two levels:

The second level is the valuable one. Individual creatives swing wildly. A block that performs better across five variants is a real signal.

Two rules keep us from lying to ourselves:

Step 5: scale the winners

For us a winner is a building block that has proven itself repeatedly, rather than a single creative.

When a hook carries, we do three things. We give it more budget in its best combination. We build new variants around it, with other visuals and formats. And we mark it as confirmed in our block table, so it gets tested at the other brands too.

A visual that worked at mate became a test at nano this way. The image stays where it is. The principle behind it travels.

What we deliberately avoid: letting a winner run forever. Creatives wear out. The pipeline keeps running even while something is performing well. The best moment for new tests is while you don't need them yet.

What stays manual

So nobody gets the wrong impression: this pipeline is half automatic.

And one limit that surprised us: more variants only help up to a point. When the budget is too small, no variant collects enough data. At that point you are distributing noise instead of testing. On a small budget, run few tests with a clear question.

What it delivered

No invented percentages: the biggest gain was calm in the process. Better numbers in the ad account came second.

We manage several times the tests per week without anyone working longer. Four weeks later we can still say why an ad performed. And the "I like this one better" discussion has almost disappeared, because the spreadsheet gets there first.

Together with the other automated areas, we save 33 to 46 hours per week across all three brands. The creative pipeline is one part of that, and it is the part that lets us run paid social seriously at all. How the rest of our marketing works is in marketing automation in ecommerce. How we work with two people overall is in how we run 3 brands with 2 people.

Conclusion

An ad creative pipeline does not make you more creative. It makes your creativity measurable.

Start small. Break your best creatives so far into hook, visual, offer and format. Set up a spreadsheet. Give every variant a machine-generated name. Evaluate once a week at block level. That's already half the way there, and you don't need to buy any software for it.

The rest is repetition. Whoever finds out faster which ad works, wins. Building the prettiest one comes second.

If you want to know what a pipeline like this would look like in your setup: talk to us at Flowhouse. We're happy to walk you through how it runs in detail at our own brands.