Omneky Introduces TASTE BENCH for Evaluating AI Creative Quality
Omneky’s new evaluation suite assesses brand fit, creative execution, and whether an ad is ready to run.
Topics
Omneky, the autonomous AI growth platform, has introduced TASTE BENCH, an evaluation suite that tests how well image and video models turn creative briefs and brand assets into ready-to-run ads.
For image ads, GPT Image 2.5 Sunburst has the highest ready-to-run rate at 54.2% (32 of 59), with an average quality score of 7.38 out of 10. Rates across the eight models range from 22.0% to 54.2%. These are first-attempt results with Omneky’s own review step disabled.
Omneky’s AI Growth Agent chooses among image and video models from several labs for each ad.
“Taste is in the details: whether the typography feels right, whether the product looks authentic, and whether the idea comes through immediately,” said Hikari Senju, Founder and CEO of Omneky. “We built TASTE BENCH to make those judgments systematic, so we know which models to trust with our customers’ ads.”
ALSO READ: Marquee Awards Recognises Marketing Excellence at Gala Celebration
The current edition covers eight image models, 59 ad briefs, four brands, and five languages (English, Japanese, Arabic, Hindi, Spanish), with 469 generated ads and 1,371 valid blind judge reviews.
Models receive identical prompts and assets, and a blind panel of AI judges from OpenAI, Anthropic, and Google scores outputs on eight quality dimensions, including brand fit, typography, and ad effectiveness.
Eleven pass/fail checks cover exact copy, correct language, faithful logos and products, safe-zone placement, and fabricated claims or visual artefacts.
An output counts as “ready to run” when it passes the applicable hard checks under the panel’s voting rules, and a majority of judges would run it as-is. Results reflect automated creative judgments, not campaign conversions or return on ad spend.