Most accounts get audited by feel. You open the search terms report, scroll a bit, see a few obviously stupid queries, add some negatives, and close the tab with a vague sense that it is “better now.” The problem is that feel does not tell you whether you just trimmed the edge of a healthy account or dabbed at a structural haemorrhage. Industry estimates consistently put wasted ad spend at 20 to 40 percent for accounts without active hygiene — a band wide enough that eyeballing cannot place you inside it.
A scorecard fixes that. Instead of one impression, you compute five query-level numbers, each with a threshold that maps to an action. Together they answer the only question that matters before you spend an afternoon cleaning up: does this account have a structural problem — the wrong queries being matched in the first place — or a tuning problem you can trim your way out of? As Conner Crowe puts it, the headline metric “takes fifteen minutes to calculate and tells you whether the account has a structural problem or a tuning problem”. The other four keep that verdict honest. None of this replaces the full hygiene audit — it is the triage you run first to decide whether the audit is even worth starting.
Metric 1: Search-term irrelevance rate
This is the headline number, and the one you compute first. Pull the search terms report over a 30-day window, sort by cost, and mark each visible row as on-theme or off-theme against the intent the campaign is meant to serve. Sum the cost of the off-theme rows, divide by the cost of all visible rows, and you have the irrelevance rate: the share of the money you can see that went to queries you did not want. Do it per campaign, never account-wide, because a tidy brand campaign will average out a leaking broad-match prospecting campaign and hide the exact thing you are looking for.
The threshold is what makes it actionable. Above roughly 25 percent off-theme, the campaign has a structural problem — the match types and structure are pulling in the wrong intent faster than negatives can block it, and the fix is upstream, not another list of exclusions. Below about 10 percent, you have a tuning problem at most, and ordinary term-by-term cleanup will keep it healthy. Between the two, mine the report but pair it with a structural review. The number itself does the deciding, which is the entire point: it turns “this looks bad” into “this campaign at 34 percent needs its match types rebuilt this week.”
Metric 2: Visible-spend share
The irrelevance rate is only as honest as the share of spend you measured it on, so this metric sets the confidence level for the whole scorecard. Compute it the same way: sum the cost of every visible row in the search terms report and divide by the campaign’s total cost for the window. That percentage is how much of the budget you can actually itemise and judge. The rest is spending on queries that never surface as rows — the low-volume terms Google withholds and the anonymous “Other search terms” aggregate.
Read the two metrics together or you will fool yourself. A clean 8 percent irrelevance rate calculated on a campaign where visible-spend share is only 30 percent is not a clean campaign — it is a clean slice, describing a third of the money while the expensive majority hides unmeasured. When visible share drops below about 30 percent, the irrelevance rate stops being trustworthy and the real problem is visibility itself: you cannot negate what you cannot see, so the fix moves to match types and structure that pull spend back into rows you can read. The full method for this one metric, and what to do at each visibility band, is in the visibility ratio audit.
Metric 3: Zero-conversion query spend
This metric isolates the money spent on queries that produced no conversions at all. Filter the search terms report to rows with zero conversions over a window long enough to be fair to your conversion lag — usually 30 to 60 days — then sum their cost and divide by total visible cost. Unlike the irrelevance rate, which is a judgement call about relevance, this one is arithmetic: the query either converted or it did not. It catches a different failure mode — queries that look on-theme, and so survive the relevance read, but quietly never pay back.
Interpret it with a little care, because a healthy account still has some zero-conversion spend: not every legitimate query converts inside the window, and cutting all of it would starve the account. The signal is concentration, not the raw total. If a handful of individually expensive queries make up most of the zero-conversion spend, those are precise negation or bid-down candidates. If the zero-conversion cost is a thin smear across hundreds of one-off queries, that is a pattern problem, and the fix is a pattern negative or a match-type change rather than hundreds of exact negatives — the reasoning is in where the rest of the waste hides.
Metric 4: Conversion-rate spread by match type
Break conversion rate out by match type — exact, phrase, and broad — over the same window, and read the gap between them. Exact and phrase should convert at a materially higher rate than broad, because they match tighter intent; that is the whole reason the tighter types exist. When broad-match conversion rate collapses to a small fraction of exact, the broad spend is buying volume that does not convert — wasted spend wearing a different label. The size of the gap is the diagnostic. A narrow gap means broad is behaving and earning its budget; a wide one means broad is dragging in queries the tighter types would never have triggered.
This metric points to the lever, not just the leak. A large broad-versus-exact gap says the fix is match type: move budget out of the broad terms that are underperforming and toward the phrase or exact coverage that converts, or tighten the broad keywords behind the worst queries. It is more precise than the irrelevance rate for this purpose, because it names which match type is doing the damage rather than just confirming that damage exists. Which way to move each keyword, and the trade-offs of tightening, are laid out in the match-type decision for 2026. Close variants deserve their own look here too, since they can drag exact and phrase down from the inside — quantified in the close-variant wasted-spend study.
Metric 5: Spend in the “Other search terms” bucket
The final metric is the mirror image of visible-spend share, and it is the one most audits ignore because there is nothing to click on. Find the “Other search terms” or unattributed aggregate — the difference between campaign cost and the sum of the itemised rows — and track it as a percentage of spend. This is the money going to queries Google will not show you individually: low-volume terms and the long tail of near-unique broad and AI Max expansions. You cannot audit it query by query, so the metric itself is the whole story: a big, growing number is a structural warning.
Watch its direction more than its level. A gradual rise on a campaign you have not changed almost always means broad match or AI Max query expansion has widened the net, pushing more spend into terms below the reporting threshold. That is a cue to act structurally — tighten match types, segment the campaign so each ad group chases a narrower intent and lifts per-query volume back above the threshold, or apply channel and brand exclusions on Performance Max. Every one of those moves shrinks the bucket and returns spend to rows you can measure, which is why re-running the scorecard after a structural change is the only way to confirm the change actually worked rather than just moved the problem.
Reading the scorecard together
No single metric is a verdict; the pattern across the five is. The most common structural signature is a high irrelevance rate and a low visible-spend share and a fat, rising Other-terms bucket — that account is matching the wrong intent and hiding most of the evidence, so the fix is match types and structure, and adding negatives will barely move the needle. The tuning signature is the opposite: high visible share, irrelevance under 10 percent, a small Other bucket, and a healthy broad-to-exact conversion spread. That account just needs its weekly trim, and a structural rebuild would be wasted effort and risk.
Run the full five monthly, per campaign, and log the numbers so you are reading direction and not just level — a metric that is drifting the wrong way on an untouched campaign is the platform changing underneath you, which the search-engine-marketing press has framed as the ongoing work to cut waste and improve ROAS as automation widens matching. Between full runs, a weekly glance at the irrelevance rate on your top campaigns is enough to catch a new leak early. The scorecard is the diagnosis; the weekly search-terms-report routine is the habit that keeps each metric from drifting in the first place.