How to Build a Customer Health Score
A health score built on gut feel or one login metric will miss churn until it's too late. Here's how to build one that actually predicts it.
Most customer health scores fail the only test that matters: they don’t predict churn earlier than a CSM’s gut feeling would. A health score built on login frequency alone, or worse, on a subjective 1-10 rating a CSM assigns during a QBR, gives you a lagging indicator dressed up as a leading one. The point of a health score is to surface risk while there’s still time to intervene — 60-90 days before renewal, not during the renewal conversation itself. Building one that actually does that requires more rigor than most teams put in.
Start from churned accounts, not from intuition
The single biggest mistake in health score design is starting with “what feels like it should matter” — feature adoption, login frequency, NPS score — and assembling those into a weighted formula without ever checking whether they actually correlate with your churn. Instead, start backward: pull every account that churned in the last 12-18 months and look at their behavior in the 90 days before they left. What patterns show up repeatedly?
This is a real analysis, not a guess — export usage data (login frequency, feature usage breadth, key action counts specific to your product) alongside support ticket volume and sentiment, and NPS or CSAT responses if you collect them, for both churned and retained accounts over the same window. Compare the two groups. You’re looking for the signals that reliably diverge between accounts that churned and accounts that renewed, because those are the signals worth weighting heavily in your score — and just as importantly, you’re looking for signals that don’t diverge, because those don’t deserve a place in the formula no matter how intuitively important they seem.
A pattern I’ve seen repeatedly in this kind of analysis: login frequency alone is a weak predictor, because plenty of healthy accounts have infrequent-but-high-value usage (a monthly reporting tool used once a month by design isn’t unhealthy), while a much stronger predictor is usage breadth — the number of distinct core features or workflows an account touches — because accounts that only ever use one narrow feature are more vulnerable to a competitor or budget cut than accounts embedded across multiple workflows, even if both log in with similar frequency.
Build the score from four categories of signal
Once you’ve identified what actually correlates with churn in your own data, organize signals into four categories, because each measures a different dimension of risk and blending them into one undifferentiated score obscures which dimension is actually failing for a given account:
- Usage depth and breadth — how many core features/workflows are actively used, not just login count. This is usually the strongest predictive category for product-led or self-serve accounts.
- Engagement trend, not absolute level — a account using the product less than it did 60 days ago is a stronger risk signal than an account with moderate but stable usage, even if the moderate account’s absolute usage number is lower. Trajectory matters more than snapshot.
- Relationship signals — for accounts with a CSM or account manager, whether scheduled check-ins are happening and being attended, whether the champion who originally bought the product is still at the company (a departed champion with no identified replacement is one of the highest-risk signals in B2B SaaS and is frequently missed because it requires someone to actually notice, not just pull from a usage dashboard).
- Support and sentiment signals — unresolved critical tickets, repeated tickets about the same unresolved issue, declining NPS/CSAT, or (for accounts that don’t respond to surveys) simply the absence of any sentiment data at all, which is itself often a risk signal because disengaged accounts don’t bother responding to surveys.
Score each category separately (e.g., 0-100 or red/yellow/green) before combining into a composite score, because a composite-only view hides why an account is unhealthy. An account red on relationship signals but green on usage needs a very different intervention (re-engage the champion, find a new one) than an account green on relationship but red on usage trend (product adoption problem, possibly needs training or a workflow redesign).
Weight signals by validated predictive power, not by category count
Once you’ve built the four categories, resist the temptation to weight them equally by default (25% each) just because there are four of them. Go back to your churned-vs-retained analysis and check which category showed the clearest divergence between the two groups, and weight accordingly. If usage breadth diverged sharply (churned accounts used 1.2 features on average, retained accounts used 3.8) while support ticket volume showed almost no difference between groups, usage breadth should carry meaningfully more weight in the composite score than support signals.
This weighting isn’t static — revisit it every 6-12 months as you accumulate more churn data, because the signals that predict churn can shift as your product, customer base, and competitive landscape evolve. A signal that was highly predictive two years ago (say, a specific onboarding milestone) may lose predictive power once your onboarding process changes, and a health score that’s never recalibrated slowly drifts from accurate to decorative.
Set thresholds using percentile distribution, not arbitrary round numbers
A common mistake: setting the “at risk” threshold at a round number like “score below 50” without checking what percentage of your actual customer base that threshold captures. If 40% of your accounts fall below 50, the threshold is too loose to be actionable — your CS team can’t meaningfully triage 40% of accounts as “at risk,” the signal has lost discriminating power. If only 2% fall below the threshold, it may be too strict, missing meaningfully at-risk accounts that don’t quite cross the line.
Instead, look at the actual distribution of scores across your base and set thresholds based on what’s operationally actionable — often the bottom 10-15% as “at risk” (red), the next 15-20% as “watch” (yellow), and the remainder as healthy (green), adjusted based on your CS team’s actual capacity to intervene. A health score that flags more at-risk accounts than your team can act on isn’t more useful than one that flags fewer — it’s less useful, because it trains the team to ignore the alert.
Build in time-to-action, not just a static score
A health score that updates monthly and gets reviewed quarterly during QBR prep is nearly useless for actual churn prevention, because by the time a declining trend shows up in a quarterly review, the intervention window has often already closed. The score needs to update on a cadence that matches how fast your churn signals actually move — for high-touch enterprise accounts with slow-moving relationship dynamics, monthly may be sufficient; for self-serve or low-touch accounts where usage can drop off within days of a workflow change, weekly or even daily automated recalculation matters more.
Pair the score with automated alerting, not passive dashboard availability. A CSM checking a dashboard once a month will miss an account that dropped from green to red in week two and stayed there for six weeks before the next scheduled check. An automated Slack or email alert triggered the moment an account crosses from green to yellow, or yellow to red, converts the score from a reporting artifact into an actual operational tool, and this single change — from passive dashboard to active alert — is often the difference between a health score that gets referenced occasionally and one that actually changes CS team behavior day to day.
Validate the score against outcomes continuously
After launch, the score itself needs ongoing validation, not a one-time build-and-forget. Every quarter, pull the accounts that churned in that period and check where their health score sat 60-90 days before churn. If accounts are consistently churning from a “green” or “healthy” score with no warning, the score has a blind spot — some risk factor isn’t being captured, and this is worth investigating specifically (a common blind spot: contract-level risk like a shrinking budget or a merger/acquisition at the customer’s company, which no product usage signal will ever capture, and which requires a manual “known risk” override flag layered on top of the automated score).
Conversely, if a large share of “red” accounts are consistently renewing anyway, the score may be over-flagging based on a signal that doesn’t actually predict churn as strongly as assumed, and the weighting needs adjustment. This ongoing validation loop — checking real outcomes against predicted risk every quarter — is what separates a health score that stays useful for years from one that quietly becomes noise within the first year because nobody checked whether its predictions were actually holding up.
A Worked Example: Two Accounts, Same Composite Score, Different Reality
Imagine two accounts both land at a composite health score of 62 out of 100 in a given month. Account A scores 85 on usage breadth and engagement trend (heavy, growing product use across multiple workflows) but 20 on relationship signals, because its original champion left the company three months ago and nobody has confirmed a replacement contact. Account B scores the mirror image: 85 on relationship signals (an engaged, responsive champion, regular QBR attendance) but 20 on usage breadth, because adoption has been stuck at a single narrow use case since onboarding.
A single blended number treats these as identical, equally-at-risk accounts, but the intervention each needs is completely different. Account A needs an urgent effort to identify and re-engage a new internal champion before renewal, since a departed champion with no successor is one of the highest-risk signals in B2B SaaS regardless of how well the product itself is performing. Account B needs a hands-on adoption push — training, a use-case expansion conversation, possibly a services engagement — because the relationship is healthy but the product hasn’t become load-bearing enough to survive a budget review. This is the concrete argument for scoring categories separately before blending them: a CS leader looking only at “62” would apply a generic save-play to both accounts and likely fail both, while a CS leader looking at the category breakdown knows exactly which lever to pull for each.
The Failure Mode: Building the Perfect Model Before You Have Enough Data
Teams building their first health score often spend months trying to perfect the weighting formula, debating whether usage breadth should be worth 35% or 40% of the composite, before the model has ever been checked against a single quarter of real outcomes. This is backwards, and it’s the most common reason health score projects stall indefinitely instead of shipping. A health score with a rough, admittedly imperfect weighting that’s live and being validated against real churn and renewal outcomes teaches you more in one quarter than a theoretically elegant model that’s still in a spreadsheet six months after the project kicked off.
The related version of this failure mode is over-indexing on the churned-account analysis from a very small sample — if your company has only churned 15 accounts in the last year, drawing firm conclusions about which signals “reliably diverge” from that sample alone risks mistaking noise for pattern. In low-churn-volume environments, supplement the quantitative analysis with structured interviews of the CS team about what they observed in those churns, and treat the early model as directional rather than statistically rigorous until more data accumulates.
Sequencing the Build So It Ships Instead of Stalling
A practical build order: start with the churned-account analysis using whatever usage and support data you already have access to, even if it’s incomplete — don’t wait for a perfect data pipeline before starting the analysis, since the goal at this stage is directional signal, not precision. Next, build the simplest possible version of the four-category score using only the two or three signals that showed the clearest divergence in your analysis, resisting the urge to include every available data point just because it’s technically accessible. Ship that simple version to a pilot group of CSMs managing a subset of accounts, gather their feedback on whether the score matches their own gut sense of account risk for a month or two, then expand to the full book of business. Only after the full rollout has run for at least one full quarter should you invest in refining weights, adding more signals, or building more sophisticated alerting — sequencing it any other way means spending the most effort on refinement before you have enough real-world feedback to know what’s actually worth refining.
Make the score a conversation starter, not a verdict
Finally, resist treating the health score as the final word on an account’s status — it’s a triage tool that tells the CS team where to look first, not a replacement for actual judgment about a specific account’s context. A score can’t capture that an account just signed a new multi-year contract with a different department, or that a champion mentioned in a call that budget season is coming and they’re bullish on renewal. Build a lightweight process for CSMs to log manual overrides or context notes against the automated score, and review the gap between automated score and CSM judgment periodically — persistent, systematic gaps in one direction usually indicate the model needs recalibration, while occasional individual overrides are exactly the human judgment layer the score was never meant to replace.
