Google Ads Experiments let you split a portion of a live campaign's traffic between your original setup and a draft, then compare performance before applying the change account wide. This guide translates Google's own significance math, split ratio guidance, and the new Experiment Power score into one place.
Google Ads Experiments let you split a portion of a live campaign's traffic between your original setup and a draft, then compare performance before applying the change account wide. This guide translates Google's own significance math, split ratio guidance, and the new Experiment Power score into one place.
A Google Ads experiment splits a portion of your campaign's traffic and budget between your original setup and a draft, then compares performance for a duration you control (Google Ads Help, 2026). Google renamed the feature "Experiments," though its own support page title still can't quite let go of "formerly drafts and experiments," so plenty of searchers are still typing in the old name.
Confirmed experiment types, as of August 2026:
One constraint applies across all seven types: only one experiment runs per base campaign at a time (Google Ads Help, 2026).
Some practitioners still duplicate campaigns manually instead of using the native tool, per a January 2024 r/PPC debate. The native tool's advantage: it splits traffic and reports significance automatically; a duplicate leaves both jobs to you.
Get either number wrong and the test still runs, it just won't tell you anything true. The split decides how much live traffic actually reaches your draft; duration decides how long you have to wait before you can trust what it says. Neither number is arbitrary.
Best for typical accounts, and the most reliable read on a small or medium sample
Best for accounts with 100+ conversions in the tested campaign
Best for the same 100+ conversion accounts, with the least exposure to the draft
Split ratio decision table. 50/50 is the default and carries the lowest statistical risk, though it fills slowest on a low-traffic account. The 100+ conversion threshold for 80/20 and 90/10 comes from DataFeedWatch's Google Ads Experiments guide; below that volume, the treatment arm stays noisy and shrinks further with every point you shift back to the original.
Google's own guidance is direct: run an experiment for "at least 4 to 6 weeks to collect enough data to evaluate the results" (Google Ads Help, 2026). Practitioner minimums run lower and shouldn't be blended with that figure: per DataFeedWatch's guide, its own minimum is 2 to 4 weeks, and it flags ending a test after 2 or 3 days as too short; Karooya flags anything under 2 weeks as a pitfall, paraphrasing Google as "4-6 weeks for some experiments" (Karooya, 2025).
is Google's own recommended minimum run time for an experiment, not a target you beat by ending early.
Cookie-based splits assign a user once, so the same visitor always sees either the base or the draft. Search-based splits can show either version across sessions, since each search is assigned independently, per Karooya's guide; Google itself flags search-based as the "(recommended option)" in the UI, per DataFeedWatch's guide.
"Statistically significant" means the difference between draft and original is unlikely to be chance, and the change should keep performing similarly once applied (Google Ads Help, 2026).
| Term | What it measures | Act on it when |
|---|---|---|
| p-value | Probability the result happened by chance | Lower is more significant; Google's own example uses p ≤ 0.05 (95% confidence) as illustrative, not a fixed rule |
| Point estimate | Estimated percent lift of draft vs. original | Positive means the draft is outperforming |
| Margin of error | Radius around the point estimate | Subtract it from the point estimate before trusting the number |
| Confidence interval | Range the true result likely falls in | Default confidence level 80%, adjustable |
| Experiment Power score | 0-100% odds of reaching significance, from Campaign Guidance | Below 50% (Low band): don't trust a result even if you get one |
Significance readout cheat sheet. Field definitions from the Google Ads API docs on experiment reporting and Google Ads Help, "Monitor your experiments".
The rule that always applies: subtract the margin of error from the point estimate. If still greater than zero, the result counts as real and positive (Google Ads API Docs, 2026). Google layers a p-value threshold on top in its worked example, p ≤ 0.05 for 95% confidence, but frames it as "for example," not a fixed requirement.
Google's own example: "There's a 95% chance that your experiment notices a +10% to +20% difference for this metric when compared to the original campaign" (Google Ads Help, 2026). Default confidence level is 80%, adjustable.
Campaign Guidance is a newer layer on top of this readout: the Experiment Power score (Google Ads Help, 2026). It estimates, before your test finishes, how likely you are to reach significance, banded Low, Medium, and High. As of August 2026 it covers Performance Max and Broad Match Search only; check your account, since coverage is expanding.
Five factors drive the score: campaign selection (enough volume), campaign variability (consistent history), traffic split (enough traffic in the experiment arm), duration (a minimum tied to volume), and experiment type (bidding-strategy tests need longer minimums) (same source). Google calls it an estimate, not a promise: it "could vary in the actual experiment." Fair enough. A score that admits its own uncertainty up front still beats finding out at week six that the traffic was never going to get you there.
Barry Schwartz described the score as showing "the likelihood of achieving statistically significant results, along with actionable recommendations" (Schwartz, Search Engine Roundtable, 2026), the same family as Google's recommendation-adoption metric, optimization score.
| Band | Range | What it means |
|---|---|---|
| Low | 0-49% | Don't trust a significant result even if you get one; fix the setup first |
| Medium | 50-79% | Workable, but budget extra weeks past the 4 to 6 week minimum |
| High | 80-99% | Best odds of a trustworthy read inside the standard window |
Experiment Power bands, per Google Ads Help, "About Campaign Guidance" (2026).
Four things reliably invalidate an experiment: stopping it before 2 weeks, changing the tested element or budget mid-run, testing something outside the account's real traffic mix, and narrowing the scope so far the sample stays thin. None of these are exotic mistakes. They're the ones anyone managing five accounts at once makes on a busy Friday.
Google's own reasons a result comes back not significant: the experiment hasn't run long enough, the campaign doesn't receive enough traffic, the traffic split was too small, or the change didn't produce a real difference (Google Ads Help, 2026).
Karooya notes pitfalls to check before launch: avoid trivial changes, avoid overlapping experiments (only one runs per base campaign anyway), watch for cannibalization between arms, and remember an experiment can't be reactivated once ended. DataFeedWatch flags a similar failure mode: testing irrelevant elements, and stopping too soon before the sample stabilizes.
An account whose campaign doesn't generate enough weekly conversions for a meaningful Experiment Power score, or enough traffic to fill even a 50/50 split, won't produce a trustworthy result no matter the duration.
Treat a Low-band reading (0-49%) as the go/no-go gate before launch, not something you discover after 4 to 6 wasted weeks (Google Ads Help, 2026).
A Low Experiment Power reading shows up most on SMB and e-commerce accounts spending $3,000 to $50,000 a month, where conversion volume is thin, and the instinct is usually to launch the test anyway. The honest alternative: wait for more volume, consolidate a bigger change into one test, or use geo holdout / lift testing instead.
kampaio's agents watch a running experiment's scorecard: Aegis flags a stalled Experiment Power reading early instead of at week 6, and Buzz surfaces the likely bid-side impact. Neither runs or bypasses Google's native tool; the Experiments tab is still the only place an experiment runs.
Experiments let you test how a change performs against your current setup by splitting live traffic between the two (Google Ads Help, 2026). If the draft wins your chosen goal, you apply the change; if not, nothing changes.
Yes. The Google Ads API supports experimental campaigns for A/B testing structure or bidding changes, built by drafting changes on a special campaign and comparing them to the base (Google Ads Help, 2026).
Google's recommendation is at least 4 to 6 weeks before evaluating results (Google Ads Help, 2026). Practitioners run lower: DataFeedWatch's floor is 2 to 4 weeks; both it and Karooya flag under 2 weeks as too short.
50/50 is the default and most reliable for a typical account. 80/20 or 90/10 only fit accounts already generating over 100 conversions in the tested campaign, per DataFeedWatch (DataFeedWatch, 2026).
Statistically significant means the observed difference is unlikely to be chance (Google Ads Help, 2026). Google confirms this with a point estimate and margin of error: a positive result needs the point estimate minus the margin of error to stay above zero (Google Ads API Docs, 2026).
A 0-100% score from Campaign Guidance, launched mid-2026, estimating your odds of reaching significance before the test finishes: Low (0-49%), Medium (50-79%), High (80-99%) (Google Ads Help, 2026). As of August 2026 it covers Performance Max and Broad Match Search only.
No. Only one runs per base campaign at a time; a second test waits until the first ends (Google Ads Help, 2026).
Reading a significance readout correctly is the harder half of running a Google Ads experiment; watching it stay on track for four to six weeks is the tedious half. kampaio's agents watch a running experiment's scorecard: Aegis flags a stalled Experiment Power reading or a drifted traffic split, and Buzz surfaces the bid-side implications once a result lands. Neither runs, edits, or bypasses Google's native Experiments tool; the Experiments tab is still the only place a test executes.
Aegis watches the Experiment Power reading and the traffic split while the test runs, and Buzz reads the bid-side impact once a result lands.
See how Kampaio watches a live experimentResults may vary. This article is informational and does not constitute professional advice. Feature availability (including Campaign Guidance and the Experiment Power score) was verified against Google's own documentation on August 11, 2026, and can change; check your own account before acting.