TL;DR: In 2026, Meta’s retrieval and ranking systems (Andromeda and GEM) pick which ads a person sees based on what the creative is, not on who you told them to target. That makes your creative testing framework the most important account structure decision you still control. There is no single correct way to do it. This guide compares the five frameworks I see working right now: Meta’s own in-campaign creative testing, the Gladiator, the CBO Stack, Charley T’s Andromeda 2, and the Testing ABO. Each fits a different budget, creative output, and team. Pick one, set your thresholds before you launch, leave tests alone for a full week, and graduate winners by post ID rather than by duplicating them.
Creative Testing Is the Account Structure Decision That Matters
When I audit a Meta account, the structure has to satisfy three requirements, and they come in a strict order.
The first is business limitations. If you sell in three countries with different margins, or you run two product lines that cannot share a budget, the account has to reflect that. These splits are not optimisation choices. They are given to you by the business, and you build around them.
The second is data maximisation. Every campaign and ad set needs enough conversions to run efficiently. Meta’s own guidance on the learning phase is roughly 50 optimisation events per ad set in a seven-day window. Most accounts I see fail this test because they have fragmented their spend across too many ad sets, each one starved and permanently “learning limited”. Consolidation is the fix, and it constrains how much testing structure you can afford.
The third is creative experimentation, and it is the one this article is about. Once the first two requirements are met, the question becomes: how do you find the messages that attract and convert your customers, in a way that is repeatable, and without breaking requirements one and two? The account that answers this well scales. The account that does not runs the same six ads until they fatigue, then panics. There is not one right answer. The five frameworks below all work, and your job is to understand the mechanics of each and choose the one that fits your budget, your creative output, and your team.
What Changed on Meta in 2026
The old creative testing playbook was built for a system where you chose the audience and Meta found the cheapest impressions inside it. That system is gone. Four changes matter for anyone deciding how to test.
Andromeda decides which ads are even considered. In December 2024, Meta’s engineering team described Andromeda as a new retrieval engine: for every impression it narrows tens of millions of possible ads down to a few thousand candidates before the auction runs, and it does so by reading the creative itself. Meta reported a 6% improvement in recall and an 8% improvement in ad quality on selected segments from the change. It rolled out globally through 2025. The practical consequence is that your creative is now the targeting input. Interests and lookalikes are, at best, a suggestion.
GEM ranks what Andromeda retrieves. In November 2025, Meta published details of its Generative Ads Recommendation Model, a foundation model trained at large language model scale that learns from long sequences of clicks, views, and interactions. Meta attributed a 5% lift in conversions on Instagram and a 3% lift on Facebook Feed to it. Engagement history on an ad is a signal this model reads, which is why post IDs matter more than they used to.
Near-duplicate ads get grouped. Practitioners call this the Entity ID: Meta has not published the term, but the behaviour is consistent across accounts. Ads that share the same footage, the same template, or the same visual with a different headline are treated as one candidate for retrieval. Thirty hook variations of one video look like thirty ads in Ads Manager and like one ad to the algorithm. They compete with each other for the same slot instead of expanding your reach. This is the single biggest reason the “test 30 variations” approach stopped working.
Meta built its own creative testing tool. From late 2025, Ads Manager has a Creative Testing option in the ad creation flow that runs a true split test of two to five ads (some accounts now see up to ten) inside an existing ad set. Jon Loomer’s write-up is the clearest reference: spend is split evenly, nobody sees more than one variant, the default duration is seven days, and Meta recommends giving the test no more than 20% of the ad set’s budget. Meta also retired the standalone flexible ad format in March 2026 and folded multi-asset ads into the flexible media option inside Advantage+ creative, which is how you now build the dynamic ads some of the frameworks below depend on.
Put those together and the rules of testing have changed. Volume of variations is no longer the lever. Distinct concepts are. And winners are rarer than most teams expect: Motion’s 2026 benchmark across 578,750 creatives and $1.29 billion in spend found that only about 5% of ads become winners, from 3.8% for accounts under $10k a month to 8.2% for accounts over $1 million. Any framework you choose has to be honest about that hit rate.
The Rules Every Framework Has to Respect
Before comparing the five, it helps to name what they have in common. Whatever structure you pick, these five rules hold.
Give every ad set enough data. If an ad set cannot realistically reach 50 conversions a week, it will spend most of its life in learning. Testing structures that split budget into many small ad sets only work for accounts with the conversion volume to feed them. Everyone else should consolidate.
Test concepts, not variations. A concept is a genuinely different idea: a different angle, persona, format, or emotional register. A variation is the same idea with a new headline or colour. Under Andromeda, variations collapse into one candidate. Meta’s own data science team found that one ad set with 25 diverse creatives beat five ad sets of five similar creatives by 17% on conversions and 16% on cost. Diversity is the input that testing needs.
Set thresholds before launch. Decide in advance what a winner looks like (spend share, CPA or ROAS against target, hook rate) and how much you will let each concept spend before you judge it. My rule is one to two times your target CPA per concept before any decision. Deciding after the fact is how teams talk themselves into keeping ads they like.
Do not touch anything for seven days. Budget changes, pauses, and new ads added mid-test restart learning and blur the read. Most practitioners running accounts at scale, including Affect Group and TheOptimizer, now treat seven days as the floor. The only exception is an ad that is a disaster from day one.
Move winners by post ID. When a winner graduates, use the existing post rather than duplicating the ad. Duplicating creates a fresh post with zero reactions, comments, or shares, and the social proof you paid for stays behind. Reusing the post ID means every ad pointing at it draws from and adds to one shared pool of engagement.
With those in place, here are the five frameworks.
The Five Creative Testing Frameworks
1. Meta Recommended: Test Inside the Sales Campaign
This is the structure Meta itself steers you towards, and it is the simplest. You run one sales campaign, ideally an Advantage+ sales campaign with broad targeting, and you test new ads inside the same ad set that holds your existing ads. Since the February 2026 Ads Manager overhaul, Advantage+ sales campaigns have proper ad sets, each capped at 50 ads, so there is room to keep winners and tests together.
The mechanism for the test is Meta’s Creative Testing tool. You open the ad set, click Set Up Test, add two to five new ads, and Meta splits a slice of the ad set’s budget evenly across them for seven days with no audience overlap. When the test finishes, everything returns to normal delivery. The winners stay live in the campaign and keep accumulating engagement. The losers get paused.
Judge the test on cost per result against your target, then on how much spend each ad attracts once normal delivery resumes. An ad that Meta chooses to keep spending on after the test is the algorithm telling you it found an audience for it.
This framework fits lean teams and accounts with limited conversion volume. Every conversion stays in one ad set, so requirement two (data maximisation) is protected, and there is no separate testing budget to defend. The trade-off is control. You cannot test existing ads with the tool, only new duplicates, and a strong incumbent can make new ads look weaker than they are once the test ends and delivery goes back to being auction-driven.
2. Gladiator: One Arena, Winners Graduate
The Gladiator is the framework most media buyers learned first. You build a dedicated testing campaign with a single ad set, ABO or CBO, and every new ad goes into that ad set to compete for the same budget. After the review period, the winners are graduated into a separate scale campaign, and everything else is paused. The testing ad set is then refilled with the next batch.
The setup matters. The testing ad set should use the same optimisation event and the same broad targeting as your scale campaign, so what wins in the arena keeps winning outside it. Six to ten distinct concepts per batch is the range that works, because fewer than that leaves the algorithm nothing to choose between and more than that leaves half the batch without meaningful delivery.
Judging is straightforward. Spend share first: the ads that Meta keeps allocating to are the ones it found demand for. Then cost per result against your threshold. An ad that took a large share of spend and hit target is a winner. An ad that took spend but missed target is worth an iteration. An ad that never got delivery did not beat its siblings, and in a Gladiator that is the test.
The Gladiator fits accounts still searching for their first bank of winners, with enough conversion volume to feed both a testing ad set and a scale campaign. The weakness is the graduation step. If you duplicate the winner into the scale campaign, you lose its social proof and it has to earn its place again. Use the post ID.
3. CBO Stack: One Ad Set Per Angle
The CBO Stack organises testing around buying personas or angles rather than individual ads. You run one CBO campaign, and each ad set holds a single angle: “a gift for a father who is hard to buy for”, “treat yourself after a hard month”, “the version for people who tried the alternative and were disappointed”. Within each ad set, ads are added in batches, and an ad is paused if it fails to beat the minimum threshold you set for that angle.
The CBO budget flows towards the angles that are working, which is the point. You are letting Meta tell you which message resonates, while the ad set structure keeps the answer readable. If the “gifting” ad set takes 60% of the budget for three weeks, that is a strategic finding, not just a media buying one, and it should shape the next round of briefs. This is where a clear messaging architecture pays off, because the angles you test are the angles you have already mapped.
Judge each ad set on its blended performance against the angle’s threshold, and each ad within it on spend share and cost per result. When a batch is added, leave it a full week before pruning.
This framework fits brands with a persona-led creative process and mid-sized budgets, where each ad set can still hit a healthy conversion count. The risk is fragmentation. Three angles is fine. Eight angles in one CBO usually means five of them are starved, and you are back to breaking requirement two. Keep the stack short and retire angles that never earn spend.
4. Andromeda 2: A Scale Ad Set Fed by Dynamic Testing
This structure comes from Professor Charley T, who has argued louder than anyone that the answer to Andromeda is fewer, better-fed ads rather than more of them. He walks through the thinking, and the account structure, in this video.
The structure is one CBO campaign with two or three ad sets. The first is the scale ad set, and it holds a collection of proven post IDs: the ads that have already demonstrated they convert and that carry their engagement with them. The other one or two are testing ad sets, and each one runs a single dynamic ad with multiple assets.
The dynamic ad follows Charley’s 3:2:2 pattern: three assets, two primary texts, two headlines, which gives Meta twelve combinations to work with. The three assets should be the same type (all video or all static) and carry the same message. You are not testing three different ideas inside one dynamic ad. You are giving Meta three executions of one idea and letting it find the combination that lands.
The review is where this framework differs from the others. You look at two signals. The first is the blended performance of the whole dynamic ad against your target. The second is the social engagement of each individual post inside it: reactions, comments, shares, saves. When a post is both part of a strong dynamic performer and has a lot of engagement, you add that post ID to the scale ad set. The testing ad sets do not get paused after review. They keep running until they underperform, and only then do you swap in a new batch of assets.
This fits accounts that already have a bank of proven ads to put in the scale ad set, and teams that want a structure they can run week to week without rebuilding. It is the framework most in tune with how GEM reads engagement. The limitation is diagnosis. Because the dynamic ad blends three assets, you rely on engagement as a proxy for which one is carrying the result, and Charley himself is candid that 3:2:2 is a discovery structure rather than an isolated test. If you need a clean read on a single concept, use the next framework.
5. Testing ABO: Fixed Budgets, Clean Reads
The Testing ABO is the most controlled option. You run a dedicated testing campaign with ad set budgets, and each ad set is either one concept or, if budget allows, a single ad. Every ad set gets the same fixed daily budget. You let them run for a predefined period, typically seven to fourteen days, and then review all of them together.
Because budget is fixed per ad set, nothing gets starved. A concept that Meta would have ignored in a CBO gets its full allocation and a fair chance to show what it can do. That is the whole reason to choose this structure. As Foxwell Digital put it, ABO gives cleaner comparisons between ads at the cost of forcing spend into what could turn out to be low performers.
Judge on cost per result against threshold, and on secondary signals like hook rate and engagement to separate “wrong concept” from “right concept, wrong execution”. Winners graduate to your scale campaign by post ID. Losers get logged, because a concept that fails cleanly is a learning you paid for.
This fits accounts with budget to spare and a reason to need isolated reads: big creative swings, new product lines, or a strategic question the CBO structures cannot answer. It also suits smaller accounts testing big differences, as I covered in how to build a scalable creative strategy, because the gap between a good concept and a bad one is wide enough for a small budget to detect. The cost is efficiency. Every ad set has to reach meaningful spend on its own, so you cannot run many at once.
How to Choose
Start with what you have. If you do not yet have proven winners, your first job is to find some, and a Gladiator or Meta’s in-campaign testing will do that fastest. Choose between them on conversion volume: if the account cannot feed a separate testing ad set with enough data, test inside the sales campaign and keep every conversion together.
If you already have a bank of winners, the question is what kind of answer you need from testing. If you need a clean, isolated read on each concept because you are making big swings, run a Testing ABO. If your creative process is organised by persona or angle and produced in batches, the CBO Stack turns your testing into a map of which messages work. If neither applies and you want the simplest thing to run every week, Andromeda 2 gives you one campaign, one review loop, and a scale ad set that gets stronger each time a post ID earns its way in.
Treat this as a starting point. Most accounts I work on move between frameworks as they mature: Gladiator to find the first winners, then Andromeda 2 or a CBO Stack to compound them, with a Testing ABO when a big strategic question comes up.
How Much Should You Be Testing?
The framework tells you where tests run. It does not tell you how many. That number follows from your budget and your target CPA, and it is usually higher than teams expect. Motion’s data makes the arithmetic plain: at a 5% hit rate, testing four ads a week surfaces roughly 0.2 winners a week, while testing eighteen surfaces roughly 0.9. Hit rate is not the lever. Volume of distinct concepts is.
If you want a number for your account, our creative volume calculator takes your monthly budget and target CPA and tells you how many ads you should be testing to keep the account fed. The reasoning behind it is in how many ad creatives do you actually need. As a rule of thumb, allocate 10% to 20% of spend to testing once you have winners, and considerably more while you are still looking for them.
Reading Results and Graduating Winners
Whichever framework you run, the review looks the same, and it happens once a week.
Read spend share first. Meta allocates budget to the ads it can find an audience for, so the distribution of spend across a batch is the algorithm’s verdict before you look at anything else. Then read cost per result or ROAS against the threshold you set before launch. Then read the diagnostic layer: hook rate, engagement, frequency. That layer does not decide the outcome, but it explains it.
Each ad ends in one of three places. It graduates if it took a large share of spend and hit target, and its post ID moves into scale. It iterates if the hook or the engagement is strong but conversion is weak, so the concept stays and the body, offer, or format changes. Or it retires, having spent past the threshold with no signal, and you log why. Most ads retire. At a 5% hit rate, that is the system working, not failing.
Keep an eye on the winners after they graduate too. Creative fatigue now arrives in weeks rather than months, and the Creative Fatigue Score we use at Toco will tell you when a scale ad set needs its next promotion before the CPA does.
The Mistakes That Break Testing in 2026
Testing variations and calling them concepts. Five headlines on one video is one test, not five. If the new ad could be mistaken for one already running, it will be grouped with it and you will learn nothing. Run every batch through the creative checklist before it launches.
Judging on day three. Pausing an ad before it has spent enough to be judged is the most expensive habit in paid social. Set the review date when you launch and do not open the ad set before it.
Fragmenting the account to test more. Twelve testing ad sets at £15 a day each is not more testing. It is twelve ad sets that never leave learning. Consolidate the structure and increase the number of distinct concepts inside it.
Duplicating winners. Every duplicate is a new post with no engagement. Use the existing post ID and let the social proof travel.
Leaving Advantage+ creative enhancements on by default. Since February 2026, new sales campaigns launch with all enhancements pre-selected. Some of them rewrite copy and swap backgrounds. That is fine for a scale ad set. In a test it means you are not sure what you tested. Review them before launch.
Frequently Asked Questions
How long should a Meta ad test run? Seven days minimum, so the test covers a full weekly cycle and the ad set has a chance to leave learning. Extend to fourteen days for high-consideration products or low conversion volume. Meta’s own Creative Testing tool defaults to seven days.
How much budget should go to creative testing? Once you have proven winners, 10% to 20% of spend is the usual range, and Meta recommends no more than 20% of an ad set’s budget for its in-campaign tests. While you are still searching for your first winners, the share should be considerably higher. Either way, each concept needs to spend one to two times your target CPA before you judge it.
Should I test with ABO or CBO? ABO when you need every concept to get a fair, fixed budget and a clean read. CBO when you want Meta to tell you where the demand is and you can accept that some concepts will be starved. Neither is wrong. They answer different questions.
Should I duplicate a winning ad or reuse its post ID? Reuse the post ID. Duplicating creates a new post with zero engagement. Reusing the existing post keeps every reaction, comment, and share attached, and lets the ad keep compounding social proof across ad sets.
Is Meta’s built-in Creative Testing tool worth using? Yes, for launching new ads into an existing ad set with an even split and no audience overlap. Its limits are that it only tests new duplicates, not existing ads, requires the Highest Volume bid strategy, and stops protecting the split once the test ends. Use it inside the Meta Recommended framework, and use a dedicated structure when you need more control.
Sources
- Engineering at Meta — Meta Andromeda: Supercharging Advantage+ automation with the next-gen personalized ads retrieval engine — How the retrieval stage works, and the reported 6% recall and 8% ad quality improvements.
- Engineering at Meta — Meta’s Generative Ads Model (GEM) — The ranking model, the engagement signals it learns from, and the 5% Instagram and 3% Facebook Feed conversion lifts.
- Meta Business Help Centre — About the learning phase — The roughly 50 optimisation events per ad set per week guidance.
- Meta Business Help Centre — About A/B testing — Meta’s guidance on split tests, audience separation, and equal budgets.
- Jon Loomer — Meta’s Creative Testing Tool: Setup, Strategy, and Results — The in-campaign creative testing tool: setup, the 20% budget guidance, seven-day default, and limitations.
- Jon Loomer — Is Meta Expanding Creative Testing to 10 Ads? — The expansion of the tool from five to ten ads in some accounts.
- Motion — Creative Benchmarks 2026: Winning ads are rare — The 5% hit rate across 578,750 creatives and $1.29 billion in spend, and the volume-to-winners arithmetic.
- Webtopia — Entity IDs, Andromeda and the New Era of Creative-Led Targeting on Meta — How near-duplicate ads are grouped and why variations are not tests.
- SuperAds — Why Creative Diversity Is the #1 Performance Lever in 2026 — Meta data science comparison of 25 diverse creatives versus five ad sets of five similar ones.
- Foxwell Digital — How to Test Creatives on Meta in 2026 — ABO versus CBO trade-offs and concept volume by spend level.
- Affect Group — Testing Meta Ads creatives in 2026 — Seven-day minimum before judging, and integrated testing within one ad set.
- TheOptimizer — How to Test Ad Creatives After Meta’s Andromeda Update — ABO testing budgets, seven-day windows, and hit rate expectations.
- Blip — How to Scale Meta Ads with Post IDs in 2026 — Finding and reusing a post ID, and when not to.
- Professor Charley T — Copy This Simple Meta Ads Strategy, It’ll Blow Up Your Business — The Andromeda account structure and the 3:2:2 dynamic ad.
- Digitopia — The 3:2:2 Method: Scientifically Testing Creatives — Breakdown of the three assets, two primary texts, two headlines pattern.
- bir.ch — Understanding Meta’s Advantage+ Sales Campaigns (2026 Guide) — Ad sets in Advantage+ sales campaigns and the 50-ad cap.
- Campaign Builder — Meta is Removing Flexible Ads — The March 2026 retirement of the flexible format and its replacement inside Advantage+ creative.
At Toco Marketing, we specialise in growth and marketing strategies that deliver measurable results. Want to drive more engagement and conversions? Book a chat today, and let’s build a strategy that works for your business!