The Ecomm Analyst

Growing stores, one honest take at a time.

What I do when two measurement tools disagree

Two tools showing different numbers for the same channel is the normal state of affairs, not a malfunction. The mistake I see most often is treating disagreement as a problem to be eliminated rather than as information about what each tool is measuring.

That said, you eventually have to act, and acting requires picking a number. Here is the order I work through it.

First, work out if it is even a disagreement

Most of what gets reported to me as a discrepancy is a comparison error, and it takes about ten minutes to rule out.

Check the date range boundaries, because tools disagree about whether a day ends at midnight in your timezone or theirs. Check the attribution model on both sides, since comparing one tool’s last-click against another’s multi-touch is not a comparison of anything. Check the lookback window. Check whether one is showing gross and the other net of refunds and discounts. Check whether subscription renewals are included in one and not the other.

I would guess more than half of the discrepancies I have chased came down to one of those six. It is worth being systematic about it before you start theorising about tracking.

Then ask what kind of disagreement it is

There are only two kinds and they need completely different responses.

A counting disagreement means the tools do not agree on how much money came in at all. Total revenue differs, or order counts differ. This is a plumbing failure and one of the tools is simply wrong. It is fixable and it must be fixed before anything else is worth reading.

An allocation disagreement means both tools agree on total revenue but split it differently across channels. This is not a failure. It is the tools doing their jobs and reaching different conclusions from different models. There is no fix, and looking for one wastes weeks.

The test is quick. Sum revenue across all channels in both tools and compare against Shopify net sales. If both totals land within a few percent of Shopify and each other, it is allocation. If either total is off, it is counting.

For counting problems, Shopify is the referee

Shopify’s net sales figure is the one number in the stack that corresponds to money that actually arrived. Whichever tool is further from it is the one with the problem.

The causes are unglamorous and repeat endlessly. A pixel not firing on a particular template or on mobile checkout. Refunds netted in one system and not the other. Test orders nobody deleted. Currency conversion on international orders. POS or wholesale orders included in one side of the comparison. Subscription orders counted at signup rather than at each renewal.

Order counts diagnose this faster than revenue does, because an order either exists or it does not and value differences cannot muddy it. If one tool sees 940 orders where Shopify has 1,000, you are missing 6% of tracking, and that missing 6% is not randomly distributed. It clusters, usually somewhere specific enough to find in an afternoon.

For allocation problems, do not average them

The instinct when one tool says a channel drove $30,000 and another says $22,000 is to split the difference and call it $26,000. That number is not more accurate than either input. It is a number with no model behind it at all, which makes it harder to reason about than either original.

What I do instead is pick one tool as the decision-making source and use the other as a sanity check. The criteria I use, roughly in order: which one reconciles more closely against Shopify, which one I understand the methodology of well enough to explain to someone else, and which one has been stable longer.

Then the second tool earns its place a different way. I stop reading its absolute numbers and start reading the gap. If tool A consistently shows a channel at 30% above tool B, and that gap suddenly becomes 60%, something changed. The gap is often a more sensitive signal than either number, because it moves for structural reasons rather than seasonal ones.

When the disagreement is the actual finding

Sometimes the gap between two tools is telling you something no single tool would.

A channel where a first-click model shows far more revenue than a last-click model is doing discovery work. That is real, and a last-click-only view would have you cut it. A channel where the two models agree closely is probably capturing existing demand rather than creating it, which is worth knowing before you scale it. Branded search almost always shows this pattern, and it is the most reliable way I know to spot it without running a test.

So before resolving a disagreement, it is worth asking whether the disagreement itself answers a question you were already trying to answer. My post on first-touch vs last-touch goes into how to read that gap.

The one case where you settle it properly

If the disagreement is large, persistent, and attached to a channel where real money is at stake, stop arbitrating between models and run a test.

Turn the channel down or off in one geography for two weeks and watch total revenue there against a comparable region. That measures whether the spend is incremental, which is the question underneath the disagreement and the one neither tool can answer. It is more work than reading a dashboard and it is the only thing that actually resolves the argument.

I would not do this routinely. I would do it before a decision big enough that being wrong costs more than the test does.

What I have stopped doing

I no longer try to get tools to agree. I spent a fair amount of time earlier on trying to configure two platforms into alignment, and it is not achievable when the models differ by design. Agreement would actually be suspicious, because it would mean one of them is not doing what it claims.

I also stopped treating the more conservative number as automatically the honest one. Lower is not the same as truer. A tool that under-attributes a channel is wrong in exactly the same way as one that over-attributes it, and the failure mode is worse, because the correction is to cut spend on something that was working.

Pick a source, understand its model, reconcile it against Shopify regularly, and use the second tool to tell you when something moved. That is most of the job.

Leave a comment

Navigation

About

Six years in e-commerce. Three Shopify stores across different niches, one scaled past seven figures. I’ve tested hundreds of ad creatives, obsessed over email flows, and learned more from my failures than my wins.

Now I focus on conversion optimization, retention marketing, and the analytics behind it all. This blog is where I share what actually works, backed by real numbers. No fluff, no guru energy.