Building a Single Source of Truth for Marketing Data
How to reconcile ad platform reports, CRM data, and analytics tools into one number everyone in the company trusts, and why most attempts fail.
Every marketing team eventually hits the same meeting: someone pulls up Meta Ads Manager showing 340 conversions, someone else pulls up the CRM showing 210 leads from paid social that quarter, and a third person pulls up Google Analytics showing a completely different number for the same channel over the same period. Nobody’s lying and nobody made an error — each tool is accurately reporting what it measured, using its own definitions, its own attribution window, and its own way of counting. The problem isn’t data quality. It’s that there’s no single, agreed system reconciling these numbers into one answer, so every report becomes a negotiation instead of a fact.
Understand why the numbers disagree before trying to fix it
Before building any reconciliation system, get specific about why the discrepancy exists, because “the numbers don’t match” has several distinct causes that each need a different fix.
Different attribution windows. Meta’s default reporting window (commonly a 7-day click / 1-day view window) counts a conversion if it happened within that window of an ad interaction, regardless of what else the customer did in between. Your CRM, tracking last-touch at the moment of form submission, may attribute that same conversion to whatever channel touched the customer last — which could be organic search three days after the ad click. Both are “correct” by their own definition; they’re just measuring different things.
Different counting units. Ad platforms often count conversions per ad account, which can double-count a single customer who clicked ads from two different campaigns before converting once. Your CRM counts unique leads. Reconciling these requires deciding which unit — impressions of conversion, or actual unique humans — is the one the business cares about, and it’s almost always the latter.
Platform self-reporting bias. Every ad platform has some incentive, structural or otherwise, to report generously on its own contribution, using the most generous attribution model available in its own dashboard by default. This isn’t necessarily deceptive — it’s just that the platform’s default settings are tuned to make the platform look good, not to give you an objective comparison across channels.
Naming which of these is causing a specific discrepancy is the first real step. Without it, reconciliation attempts usually turn into picking whichever number is most convenient for the story someone wants to tell, which defeats the purpose entirely.
Pick one system of record for identity, and route everything else through it
The core architectural decision behind a real single source of truth is choosing one system where a “customer” or “lead” is defined once, with one unique ID, and having every other tool’s data map back to that ID rather than maintaining its own separate definition of who converted. For most B2B companies this is the CRM; for many ecommerce and PLG companies it’s a dedicated customer data platform or the product database itself.
Every marketing tool — ad platforms, email tool, analytics — should ideally pass data into or reconcile against this one system, rather than each tool being treated as an equally valid, independent source of truth for the same underlying event. In practice this means the ad platform’s reported “conversions” are treated as a directional signal for platform-level optimization (useful for the platform’s own bidding algorithm and quick-glance monitoring) while the actual company-wide reporting number is derived from the CRM or customer database, which reflects what genuinely happened to a real, de-duplicated human being.
This is a bigger technical lift than it sounds — it usually requires consistent UTM tagging discipline, a clean way of passing a unique identifier (an email hash, a client ID) across systems, and some engineering time to build the pipeline connecting ad platform data to CRM records. But without this step, every other part of a single-source-of-truth effort is cosmetic — you’re just choosing which conflicting number to display more prominently, not actually resolving the conflict.
Define every metric once, in writing, before building any dashboard
A huge share of “our numbers don’t match” problems trace back to the same metric name meaning different things to different teams, not to any actual data pipeline issue. “Lead” might mean a form fill to one team and a sales-qualified opportunity to another. “Customer” might include free-trial signups to one dashboard and only paying accounts to another. Two dashboards using the same label for genuinely different definitions will always disagree, and no amount of pipeline engineering fixes a definitional mismatch.
Build a short, explicit data dictionary — a living document, not a one-time exercise — defining every metric that appears in a company-wide report: what counts, what’s excluded, what attribution window applies, what date field is used (created date vs. close date matters enormously for anything measured “this quarter”). Get every team that touches these numbers — marketing, sales, finance, product — to actually sign off on the definitions, not just marketing unilaterally deciding and hoping everyone else adopts it quietly. Disagreements surface here, in a document, where they’re cheap to resolve, rather than in a live meeting where two conflicting numbers are already on the screen and someone’s credibility is on the line.
Build the reconciliation layer, don’t just pick a favorite tool
Teams sometimes shortcut this whole problem by declaring one tool the official source (e.g., “we only trust Google Analytics numbers now”) rather than actually reconciling across systems. This avoids the meeting-room conflict but doesn’t solve the underlying problem, because the declared “official” tool still has its own blind spots and biases, and the other tools’ data — which may capture something real that the chosen tool misses — gets silently discarded rather than incorporated.
The more durable approach is a genuine reconciliation layer: a data warehouse or BI layer where raw data from each source (ad platforms, CRM, analytics, email tool) lands separately, tagged clearly by source, and gets combined using the identity-matching approach and metric definitions established above. This is more infrastructure than picking a favorite dashboard, but it’s the only approach that actually produces one trustworthy number instead of one convenient one. Modern warehouse tools (a cloud data warehouse plus a transformation layer) have made this meaningfully more accessible for mid-size teams than it was a few years ago — this no longer strictly requires a dedicated data engineering team, though it does require someone with the skill and time to own the pipeline.
Assign explicit ownership, or the reconciliation will drift within a quarter
A single source of truth built once and left unmaintained degrades quickly, because ad platforms change their default attribution settings, new tools get added to the stack, and definitions quietly drift as new team members interpret existing metric names their own way without checking the data dictionary. Without an explicit owner responsible for maintaining the pipeline and defending the definitions, the carefully built reconciliation layer slowly reverts to the same multi-number chaos it was built to fix.
Name one person or small team as the owner of the marketing data model — not necessarily a dedicated data engineer at smaller companies, but someone whose job explicitly includes noticing when a new tool gets added, when a platform changes a default setting, or when a team starts reporting a metric using a definition that doesn’t match the documented one, and correcting it before it becomes an entrenched habit.
A worked example: reconciling one channel’s numbers end to end
Take paid social for a given month. Meta Ads Manager reports 340 conversions using its default 7-day click / 1-day view attribution window. The CRM shows 210 leads with “paid social” as the last-touch source. Google Analytics, using yet another model, reports 275 paid-social-assisted conversions. None of these numbers is wrong; each is answering a slightly different question, and the reconciliation work is figuring out which parts of the gap are real and which are artifacts.
Start with identity matching: de-duplicate the 340 Meta-reported conversions against actual unique CRM records using a hashed email match. This alone might drop the number to around 260, because a meaningful share of “conversions” were the same person converting through two different ad sets, counted twice by the platform. Next, apply your agreed attribution rule — if the company has decided first-touch-within-30-days is the standard for channel credit (a defensible choice for a business with a longer consideration cycle), recheck those 260 people against their actual first touchpoint rather than Meta’s self-reported claim. Some portion of them — say 55 — actually had an organic or direct first touch that happened to be followed by a paid social ad later in their journey; under a first-touch standard, those 55 belong to a different channel’s ledger entirely.
The reconciled number that should show up in the company-wide dashboard is roughly 205 paid-social-attributed leads — close to, but not identical to, the CRM’s own last-touch count of 210, and meaningfully below Meta’s self-reported 340. Documenting each step of that bridge (raw platform number → de-duplicated → re-attributed under the agreed model) is what makes the final 205 defensible in a room rather than just another number competing with the other three.
A common failure mode: reconciling once and never again
Teams that do the reconciliation work described above often treat it as a project with an end date — build the pipeline, present the newly reconciled numbers, move on. The problem is that every input to this system changes on its own schedule: ad platforms periodically change default attribution windows without much announcement, a new marketing tool gets added to the stack with its own tracking parameters, and UTM tagging discipline degrades gradually as new campaign creators skip the naming convention under launch-day time pressure.
The result, six or nine months after a successful reconciliation project, is often a slow return of the exact symptom the project was built to fix — numbers that don’t quite match, except now harder to diagnose because everyone assumed the reconciliation layer was still authoritative and stopped checking its assumptions. The guard against this is treating reconciliation accuracy itself as a metric with an owner and a recurring check — a quarterly spot-check where someone deliberately re-traces one channel’s numbers by hand, the way the paid social example above was traced, and confirms the automated pipeline still produces the same answer. If it doesn’t, that’s the signal something upstream changed silently, and it’s far cheaper to catch in a quarterly spot-check than in a live meeting where a stakeholder notices the numbers have quietly drifted apart again.
Reconcile publicly when discrepancies surface, rather than quietly picking a number
When a discrepancy does show up — and it will, even in a well-maintained system, because ad platforms and analytics tools change constantly — resist the instinct to quietly pick whichever number supports the narrative already being told and move on. Investigate and document the specific cause (attribution window mismatch, a new tracking gap, a platform change), and communicate the resolution back to whoever raised the discrepancy.
This has a compounding trust benefit beyond the specific number in question: teams that see discrepancies get genuinely investigated and explained, rather than hand-waved away, start trusting the reporting system as a whole. Teams that see numbers get quietly adjusted without explanation start privately maintaining their own shadow spreadsheets, which is exactly the fragmented-truth state the whole project was meant to eliminate — and once that shadow-spreadsheet habit takes hold across a few teams, it’s far harder to undo than it would have been to prevent in the first place.
