GTM Container Health Scorecard

Score your GTM container health across the 8 things a real audit checks.

Container health checklist
Weight 15
Weight 15
Weight 15
Weight 15
Weight 15
Weight 10
Weight 10
Weight 5
Container health score
13/100
Data is likely unreliable
Highest-priority fix
Every tag has a documented purpose and owner
Lowest-scoring dimension

Multiple core checks failed, starting with "every tag has a documented purpose and owner." At this state, treat every number your ad platforms and GA4 report with real suspicion until the audit below is run in full.

This is a self-assessment, not the audit itself, read the full checklist below for how to verify each of these eight items in Preview mode and GA4 DebugView.

About this calculator

A GTM container can look fine and still be reporting numbers 20-40% off from reality, duplicate tags, unscoped triggers, and missing consent handling rarely throw errors, they just quietly corrupt the data everyone downstream trusts. This scorecard is a weighted self-assessment across the eight checks that actually predict clean data, based on the technical audit checklist in the Google Tag Manager Audit guide.

How to use it

  1. Answer each of the eight checklist items honestly: Yes/confirmed, Partial/not sure, or No.
  2. "Not sure" is a valid answer, and should be treated as a gap, if you can't confirm it, you can't rely on it.
  3. Read the weighted container health score out of 100, and the single highest-priority fix, the lowest-scoring dimension.
  4. Use the full audit guide below to actually verify each item in Preview mode and GA4 DebugView, this scorecard tells you where to look, not how to fix it.

Methodology

Each of the eight checklist items has a weight (15, 15, 15, 15, 15, 10, 10, 5) reflecting how much it typically affects data reliability, tag inventory, duplicate detection, trigger specificity, and Consent Mode carry the highest weight; reconciliation carries the lowest since it's a symptom-detector rather than a root cause.

Each answer scores 1 point (Yes), 0.5 points (Partial), or 0 points (No), multiplied by that item's weight. The container health score is the sum of weighted points divided by total possible weight, as a percentage.

Scores band into three tiers: 85+ is a clean container, 60-84 needs a targeted cleanup, below 60 means data is likely unreliable across the board. These bands are calibrated conservatively, a container with three or four unchecked items typically compounds into meaningfully distorted reporting even if no single item looks catastrophic alone.

This is a self-assessment based on your own honest answers, it can't verify your container directly. Treat a high score as a reason to spot-check rather than a guarantee, and re-run it after any agency handoff, site rebuild, or CRM migration, the three events most likely to quietly break a previously clean container.

FAQ

What's the single most common cause of a low score?

Unscoped triggers, tags firing on "All Pages" that should be scoped to specific events or conditions, combined with duplicate tags tracking the same event through different mechanisms. Together these two issues account for most of the inflated or double-counted conversion numbers we see in container audits.

Why does Consent Mode carry so much weight?

Consent Mode v2 directly affects whether Google Ads and GA4 can model conversions for users who decline cookies. Getting it wrong, or not implementing it at all, silently suppresses or corrupts conversion data for a meaningful share of EU traffic, and increasingly matters outside the EU as more regions adopt similar consent requirements.

How often should this scorecard be re-run?

At minimum quarterly, and always after any of the three high-risk events: an agency or freelancer handoff, a site rebuild or CMS migration, or a CRM/marketing-automation platform migration. Each of these commonly introduces duplicate tags or breaks data layer variables without anyone noticing immediately.

My score is high but my GA4 and ad platform numbers still don't match. Why?

A high score on this checklist reduces the odds of the eight most common causes, but reconciliation gaps can also stem from attribution model differences between platforms, sampling, or timezone mismatches, issues outside what a container-health checklist alone can catch.