Ecommerce Marketing Management
How to Run Performance Max Experiments
By Anata Inc. ·

The short answer.
Start a Performance Max experiment only after writing one falsifiable hypothesis, one primary metric, a fixed comparison window, and a decision rule. Google Ads provides experiment workflows that can compare Performance Max with Standard Shopping and other eligible campaign types. Keep conversion goals, economics, locations, feed quality, attribution settings, and campaign changes controlled so the comparison answers the stated question. Record the traffic split, comparable campaigns, start and end dates, and every intervention. Read the experiment scorecard only after the planned window and normal conversion delay. Apply the treatment, keep the control, or end the test according to the prewritten rule, then preserve the evidence so the same question is not tested again without a reason.
Section 01
Write the decision before opening Experiments
Begin with a business decision, not a feature to try. State the current campaign, the proposed change, the expected direction, the primary metric, the acceptable guardrails, and what the team will do for each possible result. Google's experiment guidance recommends a clear hypothesis tied to a business goal and testing one variable at a time. That discipline prevents a test of campaign type, creative, bidding, and landing pages from producing a result that nobody can interpret.
Freeze the measurement definitions before launch. Record the conversion actions included in the campaign goal, primary versus secondary status, conversion values, attribution settings, lookback windows, currency, margin source, and expected reporting delay. Choose one or two metrics before the test begins, as Google advises. Do not select a winner because a secondary chart looks favorable after the primary measure misses the declared threshold. A test is useful when its decision rule survives contact with an inconvenient result.
Section 02
Choose a supported comparison and eligible campaigns
Use the experiment type that matches the question. A Standard Shopping versus Performance Max experiment compares an eligible Shopping campaign with an existing or new Performance Max campaign using a treatment and control split. It is not a generic before-and-after comparison. Confirm that the base campaign is active, not already participating in another experiment, and suitable for the selected workflow. If Google Ads does not list the campaign, investigate eligibility rather than recreating it immediately.
Align the treatment with the control on the settings that are not under test. Keep locations, languages, conversion goals, feed, merchant account, audience definitions, and economic value inputs documented. Google advises using the same target ROAS when conversion value is the chosen focus for this comparison. If the treatment uses a different goal set and a different target, the experiment cannot isolate campaign type. Record every unavoidable mismatch as a limitation before scheduling.
Section 03
Control the traffic split and live campaign changes
Save the experiment name, control and treatment identities, traffic split, budget treatment, scheduled dates, and expected learning period in a change record. Confirm that the split matches the declared design before launch. Do not alter the control to compensate for an early chart. Google's general guidance warns that base-campaign changes can make it difficult to understand what affected the result when experiment sync is not handling them.
Create a short list of allowed operational changes, such as correcting a rejected product or a broken destination, and require a dated note for each one. Pause unrelated creative refreshes, landing-page tests, value-rule changes, and feed restructuring until the experiment ends. If a material outage, policy disapproval, inventory shock, or tracking defect affects only one arm, do not quietly extend the test. Mark the period invalid or restart with a new record after the defect is fixed.
Section 04
Monitor health without repeatedly choosing a winner
During the run, monitor serving, policy status, spend allocation, feed coverage, conversion collection, and material business interruptions. Keep a dated observation log, but avoid ending the test because one short interval favors a preferred arm. The declared window and primary metric should govern unless the test becomes unsafe, unmeasurable, or operationally invalid. Separate a platform warning from a confirmed defect and preserve screenshots or exports that show the exact state.
Review comparable campaigns in the experiment report. Google states that comparable Performance Max campaigns are excluded from reporting by default unless comparable-campaign reporting is enabled, and that their membership can be edited until the experiment concludes. Record the final comparison set because an unseen change in comparable campaigns can alter interpretation. Also wait for the ordinary conversion delay before treating the scorecard as complete.
Section 05
Make the apply-or-end decision from the plan
At the scheduled review, export the scorecard, date range, primary metric, guardrails, campaign set, and change log. Compare the result with the prewritten threshold. A favorable point estimate without sufficient evidence is not permission to claim a causal lift. If the experiment interface does not support a confident conclusion, describe the result as inconclusive and decide whether another test is worth the cost. Do not manufacture certainty from spend or impression differences alone.
Google's workflow allows an operator to update the original campaign or create a new campaign after learning from the experiment. Select the path that matches the written decision and preserve campaign history when that matters. If the treatment violates a guardrail, end it even when the primary metric rises. If applying the result changes budgets, targets, creative, or customer definitions, ship those as a separate governed change with its own rollback point.
Section 06
Archive the experiment as reusable evidence
Close the record with the decision, exact configuration, dates, exclusions, observed result, uncertainty, applied changes, owner, and next review. Google's guidance recommends keeping experiment records so future tests can build on earlier learning. Save links to the experiment and campaign, but also preserve a portable summary because account interfaces and campaign names change. The archive should let another operator understand what was tested without reconstructing the account months later.
Do not rerun the same hypothesis because performance later moves. First check whether seasonality, product availability, pricing, tracking, auctions, value rules, or customer mix changed enough to create a new question. A disciplined experiment program is a decision system, not a collection of favorable screenshots. Its output is a controlled change or an honest inconclusive result, both connected to the evidence that produced them.


