Since Harvard Business Review ran a piece on the Ipsos and Syracuse University study into AI-generated advertising, a lot of people have asked me what I make of it. The headline claim travels fast: AI ads perform worse than human ads, even when consumers can't tell them apart.
It's a bold claim, and it deserves a proper read rather than a reflexive take. So I read the study. Below is what it found, where I think it's right, and the five limitations that make me cautious about how the finding is being repeated.
For context on where I'm coming from: I co-founded Adomate, a data creative data platform for performance marketers and creative strategists running Meta campaigns. I have a commercial interest in this question. I'd rather be honest about the limits of the technology than pretend they don't exist.
What the study actually did
Ipsos, working with Adam Peruta and Carrie Riby of Syracuse University's S.I. Newhouse School, tested 20 video ads with 3,000 US consumers. Ten ads were human-made. Ten were AI-generated counterparts.
The pipeline: Peruta used Google Gemini to reverse-engineer a creative brief from each original :30 spot, had Gemini develop a concept from that brief in the form of a shot list, and fed the shot list into OpenAI's Sora 2 to produce the finished ad. No human creative intervention at any point. Some ads needed several attempts before they cleared a "good enough" bar for testing.
Ads were then evaluated in Ipsos' Creative|Spark environment, which shows ads under distracted, high-stimulus conditions rather than in a quiet lab.
What it found
The results are worth stating precisely, because the nuance is where the value sits.
Consumers couldn't reliably spot the AI work. Only about a quarter of viewers of an AI ad were even somewhat confident it was AI-made, and roughly 40% weren't sure either way. The perceptual gap has effectively closed.
But the effectiveness gap hadn't. On Ipsos' sales-validated measures, human-made ads scored around 14% stronger on short-term sales impact and around 17% stronger on long-term brand equity.


AI won where the job was functional. The strongest AI performers were product-driven and direct: a clear problem, a clear solution, a clear reason to believe. Febreze and Herbal Essences are the examples called out.
AI lost where the job was emotional. When a brief asked for a creative leap or a genuine point of view, the AI version fell flat. AI also reached for storytelling far less often than human creative teams do.
The brief mattered more than the maker. The Cheerios pairing produced the highest combined effectiveness of the whole set. Both versions landed at the top of the database. The differentiator wasn't the tool. It was a clear, empathetic human insight in the brief that gave the AI something real to execute against.
That last finding is, for me, the most useful thing in the entire study, and it's the one getting the least airtime.
Five reasons I'd be careful with the headline
None of this is a criticism of the researchers. The study is well constructed for the question it set out to answer. My issue is with how the finding is being generalised into "AI-Generated Ads Perform Worse Than Human-Made Ones."
1. Ten pairs of video ads is a narrow slice of advertising
Ten pairs. Video only. All :30 brand spots.
Think about the actual surface area of advertising: statics, carousels, UGC-style, product-on-white, comparison formats, testimonial cuts, and each of those across top, middle, and bottom of funnel. Ten ads is a thin sample of that space, and video is arguably the format where AI is weakest right now.
For most DTC and e-commerce brands running Meta, the workhorse asset is a static or a short UGC-style cut, not a thirty-second TV-style spot. The study didn't test the format a large part of the market is actually buying.
2. The entire AI arm rests on a model that OpenAI has switched off
Sora 2 launched on 30 September 2025. OpenAI shut down the Sora app on 26 April 2026, and the API sunsets on 24 September 2026.
A model getting retired inside twelve months isn't a neutral detail. It tells you the commercial viability wasn't there for exactly the kind of production work this study asked it to do. So the study's real finding is closer to: this particular video model, driven this particular way, underperformed human creative teams.
Meanwhile, on the image side, the ground has moved again. OpenAI shipped ChatGPT Images 2.5 on 8 September 2026, with better subject preservation from reference photos and far more reliable multi-turn editing. Image models are simply further along than video models for advertising work, and static ads are typically the better-performing format anyway for lower-funnel e-commerce.
3. "AI" is doing too much work as a term
The study's AI is Gemini plus Sora 2. That's it. Two models, one pipeline.
I don't know which Gemini version was used, and the prompts and shot lists haven't been published (as far as I know). That's a real gap, because prompt quality is not a rounding error in this kind of work. The difference between a lazy prompt and a well-constructed one is the difference between generic output and something usable. Without the prompts, the AI arm isn't reproducible, and we can't tell how much of the gap is model limitation versus instruction quality.
Every time you read "AI ads underperform," ask: which models, driven by whom, with what instructions?
4. Strategy was controlled, which means AI was graded on execution only
This is the crux for me, and it's worth being precise, because Ipsos is explicit about it.
They deliberately held strategy constant. The brief was reverse-engineered from an ad that had already been made by a human team. Gemini then developed a new concept from that inherited brief.
So the AI didn't get to ask who the customer is, what frustration to lead with, which angle beats the category convention, or what offer to put forward. It received a strategy that a human had already created and was scored on how well it re-executed it (original vs copy).
That's a legitimate scientific control. It isolates one variable cleanly. But it also means the study doesn't tell us anything about AI-led creative strategy, which is where I think most of the actual value is. And when the result comes back "the AI version was less original," that's partly a scoreboard for a task originality was never asked for.
5. No human in the loop is not how anyone actually works
The study's AI ads had zero human creative intervention by design. Again, defensible as a control. But nobody serious operates that way.
The real comparison for a working marketer isn't "human team versus autonomous machine." It's "human team versus human team with AI." That comparison is untested here, and it's the one that decides budgets.
There's also an effort asymmetry worth naming. The human ads were fully produced commercial spots with agency time and budget behind them. The AI ads were iterated until they hit "good enough." Those aren't matched investments, and the study doesn't claim they are.
Where I think the study is completely right
I don't want the critique to swamp the agreement, because on several points I think this research is a needed corrective to AI hype.
Credible is not the same as compelling. AI output clearing the bar for "looks like a real ad" is not the same as moving someone. A feed full of technically competent, emotionally inert creative is a real risk, and "good enough" at scale is how brands sand themselves down into sameness.
The brief is the whole game. The Cheerios result proves it. Give a model a sharp, human insight and it closes most of the gap. Give it a vague one and no amount of model quality saves you. This matches what we see every day: creative output quality tracks input quality far more tightly than it tracks model choice.
AI's strengths are specific, and they're worth knowing. Product-driven, direct, functional storytelling is where AI performed. That's not a consolation prize. For a DTC brand, that describes an enormous share of the work: the lower-funnel offer ads, the feature explainers, the variant testing, the volume.
The study I'd actually like to see next
If Professor Peruta reads this, I'd genuinely welcome a version two. Here's what I'd change.
Compare goals, not shot lists. Give a human team and an AI system the same brief inputs, not the same brief. Same context (internal and external), category data, same customer feedback corpus, same competitor creative set. Then let each side originate its own strategy and its own concept. Right now we've measured re-execution. Let's measure origination.
Test statics, and test them at volume. Ten video spots doesn't reflect how performance creative works. Five hundred statics across funnel stages would tell us something about creative diversity at scale, which is the actual promise of AI in this space.
Test the hybrid. Human-only, AI-only, and human-with-AI. The third arm is the one every marketing team is deciding on right now, and it's the one nobody has measured properly.
Use current models. Any study using a video model that's been decommissioned before publication is measuring history. That's not a flaw in the researchers' work; it's the pace of the field. It just means findings need a shelf-life label.
Validate in-market. Lab-predicted sales impact is valuable. Actual spend on actual Meta campaigns with actual conversion data is better.
What this means if you're buying media next quarter
Three practical takeaways.
Stop treating AI as one decision. It's two. First, the strategy layer: what should this ad say, to whom, with what angle and offer, based on what performance data, customer feedback, and competitor evidence. Second, the execution layer: turn that into a visual. Most tools sell you the second and quietly skip the first. That's precisely the setup this study was built to test, and it's the setup that underperformed.
Match the tool to the job. Lower-funnel, product-led, offer-driven work is where AI is already strong. Your brand film is not that. Spend your human creative hours where the leap is required and let AI carry the volume where the job is clarity.
Invest upstream. If the Cheerios finding is the real headline, and I think it is, then the highest-leverage thing you can do isn't switching models. It's getting sharper about the insight you're feeding them.
This is exactly why we built Adomate around raw data rather than abstraction: granular performance data, customer reviews, and competitor ad intelligence feeding the strategy step first, then execution second. The study tested a pipeline with the strategy step removed. It's not surprising that it came back "good enough."
I welcome every serious effort to study this field. It's a genuinely interesting problem, and we're all better off with data than with vibes.
FAQ
What did the study actually test?
Ipsos, with Adam Peruta and Carrie Riby of Syracuse University's S.I. Newhouse School, tested 20 video ads with 3,000 US consumers: ten human-made :30 spots and ten AI-generated counterparts. Google Gemini reverse-engineered a brief from each original, developed a concept as a shot list, and OpenAI's Sora 2 produced the finished ad. The spots were then evaluated in Ipsos' Creative|Spark environment, which tests ads under distracted, high-stimulus conditions rather than in a quiet lab.
Could consumers tell which ads were AI-made?
Broadly, no. Only about a quarter of viewers of an AI ad were even somewhat confident it was AI-made, and roughly 40% weren't sure either way. The perceptual gap has effectively closed.
So how much worse did the AI ads perform?
On Ipsos' sales-validated measures, the human-made ads scored around 14% stronger on short-term sales impact and around 17% stronger on long-term brand equity. The perceptual gap closed; the effectiveness gap didn't.
Where did the AI ads hold up, and where did they fall apart?
AI performed when the job was functional: a clear problem, a clear solution, a clear reason to believe. Febreze and Herbal Essences are the examples called out. It fell flat when the brief asked for a creative leap or a genuine point of view, and it reached for storytelling far less often than human creative teams did.
What's the most important finding in the study?
The Cheerios pairing. Both the human and AI versions landed at the top of Ipsos' database, producing the highest combined effectiveness of the whole set. The differentiator wasn't the tool — it was a clear, empathetic human insight in the brief that gave the AI something real to execute against. That's the finding getting the least airtime and deserving the most.
Why are you cautious about the headline?
The research is well constructed for the question it set out to answer. My problem is with generalising it into "AI-generated ads perform worse than human-made ones." Five things narrow the finding: the sample is ten pairs of video ads only; the AI arm runs on a model OpenAI has since retired; "AI" here means one specific two-model pipeline; strategy was deliberately held constant; and there was no human in the loop by design.
Is ten pairs of video ads enough to draw conclusions from?
For the study's own question, it's a clean design. As a picture of advertising, it's thin. Real surface area includes statics, carousels, UGC-style, product-on-white, comparison formats and testimonial cuts, across top, middle and bottom of funnel. For most DTC and e-commerce brands running Meta, the workhorse asset is a static or a short UGC-style cut — not a thirty-second TV-style spot. The study didn't test the format a large part of the market is actually buying.
Why does it matter that the study used Sora 2?
Sora 2 launched on 30 September 2025. OpenAI shut down the Sora app on 26 April 2026, and the API sunsets on 24 September 2026. A model retired inside twelve months isn't a neutral detail — it suggests the commercial viability wasn't there for exactly this kind of production work. The honest reading is narrower than the headline: this particular video model, driven this particular way, underperformed human creative teams.
Have the models moved on since?
On the image side, yes. OpenAI shipped ChatGPT Images 2.5 on 8 September 2026, with better subject preservation from reference photos and far more reliable multi-turn editing. Image models are further along than video models for advertising work, and statics are typically the better-performing format for lower-funnel e-commerce anyway.
Why do you say "AI" is doing too much work as a term?
Because the study's AI is Gemini plus Sora 2, two models, one pipeline. The Gemini version isn't specified and the prompts and shot lists haven't been published, as far as I know. Prompt quality isn't a rounding error in this kind of work, so without them the AI arm isn't reproducible and we can't separate model limitation from instruction quality. Every time you read "AI ads underperform," ask: which models, driven by whom, with what instructions?
What does "strategy was held constant" mean, and why does it matter?
The brief was reverse-engineered from an ad a human team had already made, and Gemini developed a new concept from that inherited brief. The AI never got to ask who the customer is, what frustration to lead with, which angle beats category convention, or what offer to put forward. It was scored on how well it re-executed someone else's strategy. That's a legitimate control, but it means the study says nothing about AI-led creative strategy — and when the verdict comes back "less original," that's partly a scoreboard for a task originality was never asked for.
Does the "no human in the loop" design reflect how teams actually work?
No, and it wasn't meant to. The real comparison for a working marketer isn't human team versus autonomous machine; it's human team versus human team with AI. That arm is untested here, and it's the one that decides budgets. There's also an effort asymmetry: the human ads were fully produced commercial spots with agency time and budget behind them, while the AI ads were iterated until they hit "good enough."
What does the study get right?
Three things. Credible is not the same as compelling, clearing the bar for "looks like a real ad" isn't the same as moving someone, and "good enough" at scale is how brands sand themselves down into sameness. The brief is the whole game, as Cheerios demonstrates. And AI's strengths are specific and worth knowing: product-driven, direct, functional work, which describes an enormous share of a DTC brand's output.
What would a better version of this study look like?
Compare goals rather than shot lists, give a human team and an AI system the same inputs and let each originate its own strategy, so we measure origination instead of re-execution. Test statics at volume, something like 500 across funnel stages. Add a third arm for human-with-AI. Use current models, since a decommissioned one measures history. And validate in-market with real spend and real conversion data, not only lab-predicted impact.
What should I do differently if I'm buying media next quarter?
Stop treating AI as one decision — it's two. The strategy layer decides what the ad should say, to whom, with what angle and offer, based on performance data, customer feedback and competitor evidence. The execution layer turns that into a visual. Most tools sell the second and quietly skip the first, which is precisely the setup this study tested and precisely the setup that underperformed. Match the tool to the job: let AI carry lower-funnel, product-led, offer-driven volume, and spend your human creative hours where the leap is required. Then invest upstream — the highest-leverage move isn't switching models, it's sharpening the insight you feed them.
You run an AI ad platform. Aren't you biased?
I co-founded Adomate, a data creative platform for performance marketers and creative strategists running Meta campaigns, so yes, I have a commercial interest in this question. I'd rather be honest about the limits of the technology than pretend they don't exist. It's also why we built Adomate around raw data rather than abstraction, granular performance data, customer reviews and competitor ad intelligence feeding the strategy step first, execution second. The study tested a pipeline with the strategy step removed, and it isn't surprising it came back "good enough."
Sources: Ipsos, "AI Ads Are Good Enough — And That's the Problem" (May 2026); Harvard Business Review (September 2026); Syracuse University Newhouse School; OpenAI.