
September 03, 2026 • 13 min read

September 03, 2026 • 13 min read
The five agencies below are ranked on one criterion: whether each can show evidence that it measures advertising causally, using holdout experiments, geo tests, or experiment-calibrated modeling, rather than reporting whatever the ad platforms attribute to themselves.
That criterion is not arbitrary. Two peer-reviewed studies established that platform-attributed numbers are not evidence of causation. Gordon, Zettelmeyer, Bhargava, and Chapsky ran 15 randomized controlled trials at Facebook covering 1.6 billion impressions and found in Marketing Science that "in half of our studies, the estimated percentage increase in purchase outcomes is off by a factor of three across all methods." In their worst case, observational methods reported lifts of 2,314 percent and 1,710 percent against a true experimental lift of 2.4 percent. Separately, Blake, Nosko, and Tadelis found in Econometrica that eBay's branded search ads had "no measurable short-term benefits," and that non-brand keywords returned negative 63 percent against naive estimates of over 4,000 percent.
An agency that cannot run an experiment is not measuring performance. It is reporting it.
Disclosure, stated once: Vibemyad publishes this blog and is not in the five below, because on this specific criterion we do not yet clear the bar the top entries set. There is a section at the end scoring us against the same test, including what we do not publish.
Three evidence classes, weighted in this order.
Published methodology counts most. Not the word incrementality, but the actual design: what the assignment unit was, how test and control were defined, what statistical method was used, and what confidence the result carried.
Independent vendor certification counts next, because it is the only evidence class that does not come from the agency's own marketing. Measured runs a formal certification requiring, in its words, "immersive training and enablement" for client-facing teams. Haus badges partners on its experiment platform. Those mean something.
One warning on badges before you use them. Triple Whale's agency tiers are banded by monthly recurring revenue, from Bronze under $1,000 to Platinum above $10,000. An agency advertising Platinum partner status is telling you how much subscription revenue it refers, not how rigorously it tests. Northbeam's tiers work similarly. Treat commercial tiers and competence certifications as different things.
Named, dated case studies count last, and almost nobody has them.
One pattern is worth naming before the entries, because it inverts what you would expect. The agencies with the loudest incrementality headlines tend to anonymize their clients and omit their methods. The agencies that publish actual methods tend to report smaller, less impressive numbers. Specificity and modesty travel together here.
Common Thread Collective, founded in 2012 and based in Costa Mesa, California, publishes more real methodology than anyone else in this category, and it is not close.
Its team describes the statistical approach by name: "the method of incrementality that we use is trying, is called synthetic controls." It documents two test designs, a straight holdout where media runs only in selected states and an inverse holdout where media is excluded from selected states. Its measurement methodology piece frames the problem exactly as the research does, describing a gap "between Reality (the true incremental impact of your spend) and Fiction (what your platform dashboards report)," and gives a worked example of Google branded search reporting 12.5x return on ad spend on-platform against 3.1x incremental. That is roughly a fourfold overstatement on branded search, which is the same failure mode the eBay experiment found.
The detail that settles the ranking is that they publish significance levels. Their example reports p equal to 0.01 alongside an incremental return of about 1.33 against a platform-reported 1.21. No other agency in this research published a p-value. They also cite a proprietary database of 146 tests across 80 stores, and their stated principle is worth quoting: "There is no perfect measurement. There is only less wrong measurement."
They also publish pricing, which almost nobody does. Their entry tier is $1,500 onboarding plus $500 a month for integration with Statlas, their analytics platform, month to month with no contract.
Two things to weigh. The flagship incrementality walkthrough anonymizes the client, and the headline figures on their incrementality page carry no dates. And the ownership changed: The Acacia Group announced a strategic investment on July 15, 2025, with majority or minority position undisclosed. Founder Taylor Holiday says he is staying. Ask what changed.
Wrong for: brands under roughly $1 million in revenue, enterprises, and anyone outside ecommerce.
Tinuiti, founded in 2004 and running roughly $4 billion in media under management with more than 1,000 staff, has the strongest third-party validation of any agency here.
It is the only agency badged by Haus, the experiment platform. And in the Forrester Wave for Media Management Services in the fourth quarter of 2024, which evaluated 12 providers against 22 criteria including Dentsu and Omnicom Media Group, Tinuiti scored 5 out of 5 on measurement and attribution. That is an independent analyst assessment rather than a self-description, which puts it in a different evidence class from most of this category.
Its Incrementality Lab is genuinely experiment-based, described as using "real-time experimentation to isolate causal lift" and comparing "users who saw your ads to those who didn't." The distinctive claim is a method that uses agency-wide impressions as a control group, avoiding the cost of buying public service announcement inventory to build controls.
Three caveats, and the first is the one that matters for a list like this. Tinuiti's May 2026 launch headline claims a "47% Average Incremental Conversion Impact" for YouTube, and discloses no sample size, no test design, no time period, and no client names. Both supporting case studies are anonymized. By the standard this article is applying, that specific claim fails. Second, the agency describes its control-group method as patented, and a search of public patent records did not surface an incrementality patent assigned to Tinuiti or Bliss Point Media, so ask for the patent number. Third, "independent" here means not owned by a holding company. Tinuiti has been backed by New Mountain Capital since December 2020, which is a different thing.
Wrong for: small DTC brands. No minimum is published, but the scale implies mid-market and above.
Brainlabs discloses more about its model architecture than any other agency in this set, and what it discloses is textbook correct.
InsightMix, published in January 2026, is built on Google's open-source Meridian library, is Bayesian, and, critically, incorporates "previous incrementality test results into our models as priors." That last part is the whole game in modern measurement. Experiments establish causal ground truth, the model consumes those results as priors, and the model then allocates budget. Naming the open-source basis is a transparency signal almost nobody offers, because it lets a technical buyer inspect the method rather than trust the brand.
The gap is proof. InsightMix has no named clients and no published case studies. The only supporting claim is an aggregate: studies delivering results in four to six weeks that "validated budget shifts of 10-30% to historically under-invested in mid-funnel channels," undated and unattributed. Brainlabs took a growth investment from Falfurrias Capital Partners in September 2023.
Wrong for: DTC brands specifically. Named clients are WeTransfer, Estée Lauder Companies, and adidas. Retail and ecommerce is one vertical among many. This is an enterprise media agency with excellent modeling, not a DTC shop.
New Engen uses the cleanest causal language of any agency site here, promising measurement of "true incremental growth (geo, matched-market, causal)" and an always-on marketing mix modeling framework, with named internal tooling called LIFT and RippleAI.
More importantly, it was named a Measured Certified Partner in the program's first cohort in March 2025. That certification requires training and enablement rather than a commercial arrangement, which is why it carries more weight than a partner logo. Its analytics lead, Andrew Richardson, is a senior hire specifically for advanced analytics and measurement.
The gap is the same as Brainlabs but wider. The measurement page carries no named clients, no case studies, and no numbers. Strong claims, a strong badge, and no published proof. DTC is also not a stated specialism.
Wrong for: buyers who want to inspect the work before signing, since there is currently nothing published to inspect.
Jellyfish ranks fifth on DTC fit and first on a single artifact, which is why it is here at all.
Its Yves Rocher case study, run with Meta in April 2023, is the only fully documented named-client experiment this research found anywhere. It describes a genuine three-cell design: group A exposed to the Meta campaign, group B not exposed, group C exposed at 80 percent through Meta's Conversion Lift. It used Meta's open-source GeoLift, which applies synthetic control methods, and cross-validated the result against Meta's own Conversion Lift study, with the two methods agreeing. The reported outcome was a 3 percent increase in new customers with positive lift across online and offline sales.
Named client, disclosed design, stated statistical method, independent cross-validation. That combination appeared exactly once in this entire research process, which says more about the category than about Jellyfish.
One thing not to be fooled by. Jellyfish's headline proprietary product, Share of Model, measures how large language models perceive a brand. It is interesting, and it is not causal media measurement.
Wrong for: DTC brands. This is a global enterprise agency with a largely European client base.
Three near-misses and one instructive absence, because the reasons are more useful than the ranking.
Power Digital is a Rockerbox partner, so it does run a measurement tool. But its published incrementality content is generic education built on a hypothetical example, and its nova platform mentions incrementality modeling in a case study headline while disclosing no methodology at all. The named client results cited, for Casper and others, are not incrementality results. This is the clearest example of measurement used as vocabulary rather than practice.
Croud publishes a marketing mix modeling explainer that is purely educational, with no tooling, no named clients, no methodology, and a closing invitation to contact the analytics team. There is no published evidence of causal measurement practice.
Monks, which dropped "Media" from its name in July 2024, publishes modeling content that leads with multi-touch attribution used "in conjunction with" incremental techniques. Leading with multi-touch attribution is specifically what the Gordon and Blake results caution against.
WITHIN is the interesting one and the inverse of everyone else. It holds three independent vendor credentials, as a Measured Certified Partner, a Rockerbox partner, and Northbeam Premier, yet its own site never mentions incrementality, modeling, or measurement anywhere. If you want an agency that quietly does the work and does not market it, this is the candidate. We left it out because ranking on badges alone would violate the criterion.
Wpromote deserves a note as the opposite case. It holds two independent vendor credentials, Measured Certified and Northbeam Elite, which is genuinely strong. But its published methodology is thin, and its claim of "market-leading 95% accuracy" for its planning model is unsubstantiated and difficult to interpret without knowing what is being predicted against which holdout. Good badges, thin proof.
One request separates agencies that run experiments from agencies that describe them.

An agency that runs experiments already has the test design document, because it cannot run one without it.
Ask for a test design document before you ask for results. It should specify the assignment unit, whether geographic or user-level, how test and control groups are defined, the length of the pre-period, the duration of the test, the minimum detectable effect, and the significance threshold. An agency that runs experiments has this document already, because it cannot run one without it. An agency that talks about incrementality will send you a case study instead.
Three follow-ups worth the time. Ask how they treat branded search, because if the answer does not acknowledge the substitution problem they have not read the eBay result. Ask what happens when the platform, the analytics, and the order count disagree, because a good answer names one as ground truth and explains the gap. And ask which measurement vendor they are certified on and what the certification required, so you can tell a training program from a reseller agreement.
For the pricing side of the decision, our comparison of retainer versus outcome pricing works through the trade-off, and our explainer on what a performance marketing agency is covers the three-layer measurement stack these agencies are operating inside. If your constraint is Meta specifically rather than performance across channels, our ranking of Meta ads agencies scores five on a different criterion, which is what each publishes before a sales call. Common Thread Collective appears in both lists and places differently in each, because the question is different.
Applying our own criterion to ourselves produces an uncomfortable answer, which is the point of publishing the criterion.
What we do not have. We do not publish a test design document. We do not publish p-values or confidence intervals on client results. We hold no measurement vendor certification, so nothing in this section is independently verified the way Tinuiti's Haus badge or New Engen's Measured certification is. On the three evidence classes this article ranks by, we clear the third and not the first two. A brand for whom causal measurement rigor is the deciding factor should hire Common Thread Collective or Tinuiti, and we would tell you that on a call.
What we do. Vibemyad is an AI-native marketing agency working with United States direct-to-consumer brands, mostly on Shopify, in beauty, apparel, food and beverage, and wellness. We price against the outcome rather than a retainer, with no contract beyond the current month, which puts our fee at risk against the result rather than against a scope of work. That structure creates a specific obligation: an agency paid on outcomes cannot afford broken tracking, so event integrity and reconciliation against the store's own order count come before any spend increase. Our platform reads live category advertising and classifies what it finds by hook, format, offer, and funnel position, which feeds the creative pipeline rather than guesswork.
What we are changing. The honest reading of this research is that publishing methodology is the cost of being credible in this category, and that most agencies, us included, have not paid it. Publishing our test design is on the roadmap. Until it exists, treat this section as a statement of intent rather than evidence, and weigh it accordingly.
Wrong for: brands that need a fixed retainer line for budgeting, pre-product-market-fit stores, and anyone below the spend level where professional management pays for itself.
Get notified when new insights, case studies, and trends go live — no clutter, just creativity.
Table of Contents

Arpita Mahato
Content Writer, Vibemyad

Arpita Mahato
Content Writer, Vibemyad

Arpita Mahato
Content Writer, Vibemyad