Google Ads

Google Ads Experiments: How to Run an A/B Test You Can Actually Trust

Google Ads Experiments let you split a portion of a live campaign's traffic between your original setup and a draft, then compare performance before applying the change account wide. This guide translates Google's own significance math, split ratio guidance, and the new Experiment Power score into one place.

K
By Kampaio TeamSenior PPC strategy at KampaioAugust 11, 2026 · 11 min read

Google Ads Experiments let you split a portion of a live campaign's traffic between your original setup and a draft, then compare performance before applying the change account wide. This guide translates Google's own significance math, split ratio guidance, and the new Experiment Power score into one place.

What is a Google Ads experiment (and what it isn't)?

A Google Ads experiment splits a portion of your campaign's traffic and budget between your original setup and a draft, then compares performance for a duration you control (Google Ads Help, 2026). Google renamed the feature "Experiments," though its own support page title still can't quite let go of "formerly drafts and experiments," so plenty of searchers are still typing in the old name.

Confirmed experiment types, as of August 2026:

  • Ad variations
  • AI Max experiments (one-click, 50/50 split, Search only; see AI Max experiments)
  • App asset experiments
  • Custom experiments (Search, Display, App; typically used to test Smart Bidding strategies, match types, landing pages, audiences)
  • Demand Gen experiments
  • Performance Max experiments
  • Video experiments (2 to 4 arms, measured on brand lift or conversions)

One constraint applies across all seven types: only one experiment runs per base campaign at a time (Google Ads Help, 2026).

How to set up a Google Ads experiment

  1. Create a draft from the base campaign

    The draft inherits the campaign's structure, so the only real differences are the ones you introduce.
  2. Make the change you want to test

    Keep it to one variable: swap the landing page URL for a product category, or shift the daily budget on a single ad group. That's a cleaner example than a bid-strategy test, already covered in a dedicated walkthrough elsewhere on this site.
  3. Choose up to two goals to measure

    Pick the metrics that decide whether the draft wins.
  4. Define the traffic split

    How much live traffic the draft sees relative to the original (more on the ratio decision below).
  5. Set the timeframe

    A start date and, ideally, a fixed end date rather than an open-ended run.
  6. Decide whether to sync

    Sync keeps everything you're NOT testing aligned with live edits to the base campaign while the experiment runs. Turn it off to isolate the test from unrelated mid-run changes.
  7. Launch the experiment

    From there the Experiments tab does the traffic splitting and the significance reporting for you.

Some practitioners still duplicate campaigns manually instead of using the native tool, per a January 2024 r/PPC debate. The native tool's advantage: it splits traffic and reports significance automatically; a duplicate leaves both jobs to you.

What split ratio and duration actually mean

Get either number wrong and the test still runs, it just won't tell you anything true. The split decides how much live traffic actually reaches your draft; duration decides how long you have to wait before you can trust what it says. Neither number is arbitrary.

50/50 split

Best for typical accounts, and the most reliable read on a small or medium sample

  • Yes: Safe on a small conversion sample
  • No: Needs 100+ conversions in the campaign
  • No: Fastest to launch on a large baseline
  • No: Minimizes live traffic exposed to the draft

80/20 split

Best for accounts with 100+ conversions in the tested campaign

  • No: Safe on a small conversion sample
  • Yes: Needs 100+ conversions in the campaign
  • Yes: Fastest to launch on a large baseline
  • No: Minimizes live traffic exposed to the draft

90/10 split

Best for the same 100+ conversion accounts, with the least exposure to the draft

  • No: Safe on a small conversion sample
  • Yes: Needs 100+ conversions in the campaign
  • Yes: Fastest to launch on a large baseline
  • Yes: Minimizes live traffic exposed to the draft

Split ratio decision table. 50/50 is the default and carries the lowest statistical risk, though it fills slowest on a low-traffic account. The 100+ conversion threshold for 80/20 and 90/10 comes from DataFeedWatch's Google Ads Experiments guide; below that volume, the treatment arm stays noisy and shrinks further with every point you shift back to the original.

Google's own guidance is direct: run an experiment for "at least 4 to 6 weeks to collect enough data to evaluate the results" (Google Ads Help, 2026). Practitioner minimums run lower and shouldn't be blended with that figure: per DataFeedWatch's guide, its own minimum is 2 to 4 weeks, and it flags ending a test after 2 or 3 days as too short; Karooya flags anything under 2 weeks as a pitfall, paraphrasing Google as "4-6 weeks for some experiments" (Karooya, 2025).

4-6
weeks before you evaluate

is Google's own recommended minimum run time for an experiment, not a target you beat by ending early.

Source: Google Ads Help, About the Experiments page, 2026 (support.google.com/google-ads/answer/10682377)

Cookie-based splits assign a user once, so the same visitor always sees either the base or the draft. Search-based splits can show either version across sessions, since each search is assigned independently, per Karooya's guide; Google itself flags search-based as the "(recommended option)" in the UI, per DataFeedWatch's guide.

How to read the significance readout (p-value, confidence, and the new Experiment Power score)

"Statistically significant" means the difference between draft and original is unlikely to be chance, and the change should keep performing similarly once applied (Google Ads Help, 2026).

TermWhat it measuresAct on it when
p-valueProbability the result happened by chanceLower is more significant; Google's own example uses p ≤ 0.05 (95% confidence) as illustrative, not a fixed rule
Point estimateEstimated percent lift of draft vs. originalPositive means the draft is outperforming
Margin of errorRadius around the point estimateSubtract it from the point estimate before trusting the number
Confidence intervalRange the true result likely falls inDefault confidence level 80%, adjustable
Experiment Power score0-100% odds of reaching significance, from Campaign GuidanceBelow 50% (Low band): don't trust a result even if you get one

Significance readout cheat sheet. Field definitions from the Google Ads API docs on experiment reporting and Google Ads Help, "Monitor your experiments".

The rule that always applies: subtract the margin of error from the point estimate. If still greater than zero, the result counts as real and positive (Google Ads API Docs, 2026). Google layers a p-value threshold on top in its worked example, p ≤ 0.05 for 95% confidence, but frames it as "for example," not a fixed requirement.

Google's own example: "There's a 95% chance that your experiment notices a +10% to +20% difference for this metric when compared to the original campaign" (Google Ads Help, 2026). Default confidence level is 80%, adjustable.

Campaign Guidance is a newer layer on top of this readout: the Experiment Power score (Google Ads Help, 2026). It estimates, before your test finishes, how likely you are to reach significance, banded Low, Medium, and High. As of August 2026 it covers Performance Max and Broad Match Search only; check your account, since coverage is expanding.

Five factors drive the score: campaign selection (enough volume), campaign variability (consistent history), traffic split (enough traffic in the experiment arm), duration (a minimum tied to volume), and experiment type (bidding-strategy tests need longer minimums) (same source). Google calls it an estimate, not a promise: it "could vary in the actual experiment." Fair enough. A score that admits its own uncertainty up front still beats finding out at week six that the traffic was never going to get you there.

Barry Schwartz described the score as showing "the likelihood of achieving statistically significant results, along with actionable recommendations" (Schwartz, Search Engine Roundtable, 2026), the same family as Google's recommendation-adoption metric, optimization score.

BandRangeWhat it means
Low0-49%Don't trust a significant result even if you get one; fix the setup first
Medium50-79%Workable, but budget extra weeks past the 4 to 6 week minimum
High80-99%Best odds of a trustworthy read inside the standard window

Experiment Power bands, per Google Ads Help, "About Campaign Guidance" (2026).

🛡️Aegis· Risk review
Say your Experiment Power score comes back at 30%. That's a Low-band read, under the 50% floor, and I flag it before the clock starts: don't spend four to six weeks on a test that was never built to reach significance. Fix the traffic split or wait for more conversion volume first.

What invalidates a Google Ads experiment

Four things reliably invalidate an experiment: stopping it before 2 weeks, changing the tested element or budget mid-run, testing something outside the account's real traffic mix, and narrowing the scope so far the sample stays thin. None of these are exotic mistakes. They're the ones anyone managing five accounts at once makes on a busy Friday.

Google's own reasons a result comes back not significant: the experiment hasn't run long enough, the campaign doesn't receive enough traffic, the traffic split was too small, or the change didn't produce a real difference (Google Ads Help, 2026).

Karooya notes pitfalls to check before launch: avoid trivial changes, avoid overlapping experiments (only one runs per base campaign anyway), watch for cannibalization between arms, and remember an experiment can't be reactivated once ended. DataFeedWatch flags a similar failure mode: testing irrelevant elements, and stopping too soon before the sample stabilizes.

When a Google Ads experiment is the wrong tool

An account whose campaign doesn't generate enough weekly conversions for a meaningful Experiment Power score, or enough traffic to fill even a 50/50 split, won't produce a trustworthy result no matter the duration.

Treat a Low-band reading (0-49%) as the go/no-go gate before launch, not something you discover after 4 to 6 wasted weeks (Google Ads Help, 2026).

A Low Experiment Power reading shows up most on SMB and e-commerce accounts spending $3,000 to $50,000 a month, where conversion volume is thin, and the instinct is usually to launch the test anyway. The honest alternative: wait for more volume, consolidate a bigger change into one test, or use geo holdout / lift testing instead.

kampaio's agents watch a running experiment's scorecard: Aegis flags a stalled Experiment Power reading early instead of at week 6, and Buzz surfaces the likely bid-side impact. Neither runs or bypasses Google's native tool; the Experiments tab is still the only place an experiment runs.

Google Ads Experiments FAQ

What are Experiments in Google Ads?

Experiments let you test how a change performs against your current setup by splitting live traffic between the two (Google Ads Help, 2026). If the draft wins your chosen goal, you apply the change; if not, nothing changes.

Does Google Ads still have experimental ads?

Yes. The Google Ads API supports experimental campaigns for A/B testing structure or bidding changes, built by drafting changes on a special campaign and comparing them to the base (Google Ads Help, 2026).

How long should a Google Ads experiment run?

Google's recommendation is at least 4 to 6 weeks before evaluating results (Google Ads Help, 2026). Practitioners run lower: DataFeedWatch's floor is 2 to 4 weeks; both it and Karooya flag under 2 weeks as too short.

What split ratio should I use for a Google Ads experiment?

50/50 is the default and most reliable for a typical account. 80/20 or 90/10 only fit accounts already generating over 100 conversions in the tested campaign, per DataFeedWatch (DataFeedWatch, 2026).

What does "statistically significant" mean here?

Statistically significant means the observed difference is unlikely to be chance (Google Ads Help, 2026). Google confirms this with a point estimate and margin of error: a positive result needs the point estimate minus the margin of error to stay above zero (Google Ads API Docs, 2026).

What is the Experiment Power score?

A 0-100% score from Campaign Guidance, launched mid-2026, estimating your odds of reaching significance before the test finishes: Low (0-49%), Medium (50-79%), High (80-99%) (Google Ads Help, 2026). As of August 2026 it covers Performance Max and Broad Match Search only.

Can I run more than one experiment on the same campaign?

No. Only one runs per base campaign at a time; a second test waits until the first ends (Google Ads Help, 2026).

Reading a significance readout correctly is the harder half of running a Google Ads experiment; watching it stay on track for four to six weeks is the tedious half. kampaio's agents watch a running experiment's scorecard: Aegis flags a stalled Experiment Power reading or a drifted traffic split, and Buzz surfaces the bid-side implications once a result lands. Neither runs, edits, or bypasses Google's native Experiments tool; the Experiments tab is still the only place a test executes.

Catch a stalled experiment in week one, not week six

Aegis watches the Experiment Power reading and the traffic split while the test runs, and Buzz reads the bid-side impact once a result lands.

See how Kampaio watches a live experiment

Results may vary. This article is informational and does not constitute professional advice. Feature availability (including Campaign Guidance and the Experiment Power score) was verified against Google's own documentation on August 11, 2026, and can change; check your own account before acting.

Sources

  1. Google Ads Help, "About the Experiments page" (2026). Experiment types taxonomy, the 4 to 6 week recommendation.
  2. Google Ads Help, "Monitor your experiments" (2026). Statistical significance definition, confidence interval, reasons a result is not significant.
  3. Google Ads Help, "Campaign experiment: Definition" (2026). Definition of the split, one experiment per base campaign.
  4. Google Ads Help, "About Campaign Guidance" (2026). Experiment Power score, bands, and the five factors behind it.
  5. Google Ads API Docs, "Report on experiments" (2026). p-value, point estimate, margin of error definitions.
  6. DataFeedWatch, "Google Ads Experiments Guide". Split ratio thresholds, 2 to 4 week minimum, search-based split as the recommended option.
  7. Karooya, "Google Ads Experiments Guide" (2025). Cookie vs search based splits, pre-launch pitfalls.
  8. Barry Schwartz, Search Engine Roundtable (June 9, 2026). Coverage of the Campaign Guidance Experiment Power launch.

Keep reading