Your customer data isn't objective. You helped create it.

The customer list you're using to plan next year was assembled by the campaigns you ran last year.

Written down like that it sounds obvious. It stops being obvious the moment somebody in a meeting says let's build a lookalike audience off our best customers, and everyone nods, because who would argue with that.

Here's the loop underneath it.

You run a campaign. It reaches certain people and not others, for reasons that include your targeting, your budget, your channel mix, and a lot of things nobody in the room chose. Some of the people it reached buy something. Those buyers become your customer data. You build your next audience out of that data. It reaches people who resemble the first group. Some of them buy. Now they're in the file too.

Run that cycle a few times and your customer file stops being a picture of your market. It becomes a record of your own past decisions, wearing the costume of objective data.

There's a well-documented version of this problem in an adjacent field: recommendation systems, and researchers there have a name for it.

Chaney, Stewart and Engelhardt called it algorithmic confounding in a 2018 paper, and the definition is almost exactly our problem: the model is trained on data that was already shaped by the model that came before it. So you can no longer tell what people actually prefer from what the system taught them to do.

They simulated it. Communities of users, a thousand cycles, six different recommendation methods. Every method except recommending things at random made users more similar to each other over time, and the effect compounded with each pass through the loop. The part I find most damning is that the homogenizing came without a matching gain in usefulness. The system wasn't narrowing people down because narrow was better for them. It narrowed because narrowing is what the loop does.

Marketing has it worse than a recommender does, and it's worth being honest about why.

A recommendation engine is at least trying to serve the person in front of it. A lookalike model starts with the people you already reached. So whatever else it learns, your own history is baked into the starting point.

I have run straight into this, and not only in the obvious place.

Years ago I ran audience research for a set of consumer brands by studying the people already following them. It was good work and I still believe in the method. But every person in that dataset was someone previous marketing had already found. Everything I concluded was downstream of decisions I hadn't made and couldn't see.

The subtler version showed up in destination marketing.

At Visit Central Oregon we built a measurement system connecting campaign exposure to confirmed visitation. We could see far beyond clicks: what someone had been served, what they engaged with or passed over, and ultimately whether they arrived.

But even that couldn't tell us about everyone.

Before we could measure someone's response to our marketing, they had to enter the universe our marketing could reach. The targeting decisions, media plan, flight markets and campaign strategy all helped determine who entered that universe in the first place.

So even the richest behavioral dataset I'd ever had still contained the fingerprints of the decisions that created it.

Which is why the best audience work I've done was for a bank entering markets where we had no customers at all.

First Interstate had just acquired Great Western, and we were launching the brand into eight new states and more than twenty markets. We couldn't simply assume First Interstate's historical customer profile described these new markets. Each had its own population.

That constraint turned out to be a gift.

With no customer list to mine, we went and built a demographic picture of who actually lived in each market, then matched the creative to the dominant audiences we found there. Some of those markets included tribal reservation communities. Some had large Hispanic populations. None of that would have surfaced from a lookalike model built on the bank's existing book, because the existing book was assembled somewhere else, by somebody else's marketing, among people who don't live there.

We weren't studying our customers. We were studying the population.

The same instinct showed up in Central Oregon, in the one place we got it right. We ran focus groups in key markets with direct flights in. Not markets that were already sending us visitors. Markets we wanted.

So here's the practical version, for anyone who has to make a plan out of a CRM export.

Treat your customer file as evidence about your past marketing, not as a description of your market. It is genuinely both, but only one of those is what people usually use it for.

Ask who isn't in it, out loud, in the meeting. Not as a philosophical exercise. Pull the census or the category data for your footprint, lay it next to your customer profile, and look at the gap. The gap is the part your data cannot tell you about, and it is often where the growth is.

Budget something for reaching people your model would never select. Call it prospecting, call it a test, call it whatever gets it approved. If every dollar goes to audiences derived from your existing customers, you have built a machine that finds you more of what you already have, forever, with increasing efficiency and decreasing reach.

And be suspicious when your audience gets easier to describe over time. That usually gets reported as clarity. It's often just the loop tightening.

None of this means the data is wrong. It means it has a memory.

And what it remembers is you.

Sources

Allison J.B. Chaney, Brandon M. Stewart and Barbara E. Engelhardt, "How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility," Proceedings of the 12th ACM Conference on Recommender Systems, 2018.

Next
Next

AI can simulate your customer. It can't be one.