Research

The False Diagnosis

Why performance changes are easier to explain than to diagnose.

Lei · Sep 23, 2026

On a single day, a Google Ads team running a trading-platform account made three changes at once: they removed the primary conversion goal from one campaign, removed its target CPA, and cut its daily budget by roughly a third.

The next day, conversion rate roughly halved. CPA more than doubled.

Every explanation was immediately available. The budget was too low. The bidder had lost its target. The creative was tired. The traffic had changed. Each one fits.

A story that fits is cheap. Diagnosis starts when you look for evidence that separates the ones that do.

01

Is there anything to diagnose?

The first day's numbers came from a small base — a handful of conversions either side. A swing that size, on its own, can be noise. What made it worth investigating was that the decline held over the following several days, not just the first.

There was a second problem with the comparison. The conversion definition changed the same day the numbers did. Before and after, CPA was counting different things. An anomaly exists only against a baseline, and only if the metric means the same thing on both sides of it.

02

Three changes, one timestamp

Bundling changes destroys its own evidence. Whatever happens next, nobody can say cleanly which change did it.

But the data wasn't silent, and "we can't tell" would have been lazy. The useful question is what each change could plausibly explain — and what the available evidence makes more or less convincing.

The budget was cut by roughly a third. That could reduce available traffic and, depending on how delivery changed, alter its composition as well. But a budget change does not by itself explain why conversion rate fell by half. Without verified spend, click and traffic-mix data around the change, we could not push that explanation further.

Removing the goal and the target changed the optimization setup, so it could also have changed traffic selection. We had no basis to predict where. The data showed an uneven pattern afterwards: conversion rate fell substantially in one broad query segment, while brand held roughly flat. That fits a selection effect. It doesn't confirm one — a split found after the fact earns no credit as a test, only as a lead.

The split also gave us a rough internal comparison. Because both segments belonged to the same account and measurement environment, a system-wide failure became somewhat less convincing. It did not rule out a segment-specific funnel or tracking problem, and the comparison was imperfect because the users differed in intent — it shifted weight without killing anything.

Impressions were flat and top-of-page rate held or rose slightly. "We lost the auction" got weaker too.

03

The obvious explanation, checked

The tempting story: removing the conversion goal from the count dropped the numbers by accounting alone. We tested it. The removed event was rare to begin with — a few a day, often zero — so taking it out of the count couldn't account for a fall of that size in total conversions, and the remaining goal's own volume dropped as well.

That kills one mechanism, the count. It doesn't touch the other: removing the goal and the target changed what the bidder was pursuing, not just how it was scored.

04

Mix explained part of it

Traffic mix explained part of the decline and not all of it. Given the shift in query mix alone, the expected conversion count was meaningfully higher than what the account actually delivered.

We weren't choosing between "bidding" and "mix." Both were in the answer, in unknown proportions.

05

The change didn't just resize the target. It redefined it.

The account had been live for months. Across most campaigns, the primary goal was FTD — a first deposit, close to the actual business outcome — with several hundred conversions over the prior quarter. One campaign was the exception: registration, an earlier funnel event with substantially more volume, was Primary.

That exception was the campaign edited on day zero — which changes how we read the intervention.

The team had removed the target CPA and cut the budget, both obvious bidding levers. But the same intervention also removed FTD from the campaign's goals. It wasn't only changing how the campaign pursued its objective. It was changing the objective itself — a different kind of intervention, and one that had largely disappeared from view by the time the account was being diagnosed.

06

What we did not conclude

None of this proves that the goal change caused the decline.

The budget changed. The target disappeared. Traffic mix shifted. The conversion definition changed. Traffic segments behaved differently. Several mechanisms remained compatible with parts of the evidence.

But one finding changed what we could justify doing next: the intervention had been misclassified. What looked like a bidding adjustment had also changed the optimization objective. That mattered because another bidding change would have added another intervention before the existing one was properly understood.

So the recommendation wasn't another bidding strategy. First, restore a consistent optimization objective across campaigns pursuing the same business outcome. Keep earlier funnel events Secondary unless there is a deliberate reason for a campaign to optimize toward them. Then evaluate bidding against a stable definition of success.

This wasn't a treatment for the original performance decline. It was a way to make the next result interpretable — so that if performance remained poor afterward, bidding, traffic mix, creative, product and the other surviving explanations would still be there, now with better evidence to separate them.

07 · Method

The method in five questions

The questions are general. Next to each one: where it landed in this case.

  1. Is the change real, against which baseline, and what is the metric counting?

    In this caseSmall base on day one, but the decline held for several days. The conversion definition changed the same day, so before-and-after CPA counted different things.

  2. What materially contributes to it?

    In this caseThe shift in query mix explained part of the decline, not all of it. Bidding and mix were both in the answer, in unknown proportions.

  3. What else could produce the same pattern?

    In this caseThe budget cut, the lost target, lost auction position, creative, a funnel or tracking fault — and the removed goal falling out of the count.

  4. What would look different across those explanations?

    In this caseFlat impressions and top-of-page rate weakened the auction story. The removed event's low volume ruled out the count. The broad-versus-brand split weakened a system-wide failure — as a lead, not a test.

  5. What can we justify doing now, and what is the cheapest reversible way to learn more?

    In this caseNot another bidding change. Restore a consistent optimization objective first, then evaluate bidding against a stable definition of success.

Prioritize checks by information value, cost and risk. Cheap configuration and measurement checks are often good early candidates, because a failure there can invalidate everything downstream. There is no universal order.

08

Where AI helps, and where it hurts

A language model will turn five observations into one clean story in seconds. That is the failure this piece is about.

Its useful job is the first pass: pull the change history, line up the edit with the metric break, do the arithmetic, list which hypotheses each observation weakens, and say what it cannot see — platform-side delivery changes, auto-applied recommendations, events outside the tracker. "No changes found" is only as good as the change history it read.

09

The point

The hard part of performance diagnosis isn't finding something wrong — there's usually plenty of that. It's establishing which of those things explains the change in front of you.

Back This

Research stays free. Support keeps it that way.

If a piece saved you time, you can back the Lab directly.