Imagine two eCommerce brands. Same category, similar traffic, comparable ad budgets. One of them can answer, in a sentence, questions like: Which customers bought twice last year but haven't visited in ninety days? Which products get browsed heavily but returned constantly? Which of yesterday's signups came from actual humans?
The other one can't. Not because the team is less capable — but because, three years ago, nobody made sure the answers were being collected.
That's the difference this course is about. First-party data isn't a report you can pull when you need it. It's an asset you either built or you didn't — and the moment you want to use it is exactly the moment it's too late to start collecting it.
This first lecture is a map of the territory: what first-party data is, how it differs from its neighbors, what it's made of, and why its value — which was always there — is suddenly becoming visible to everyone. The rest of the course goes deep on how to actually capture it right and keep it clean. This lecture is free and open; the rest is free too, but we'll ask who you are at the door — more on that at the end.
The three kinds of data we work with
When we build data foundations for eCommerce brands, we're typically working with three distinct categories of customer data. They differ in one fundamental way: how the information comes into existence.
Three categories, three verbs: what they do, what they say, and whether you can reach them. Hover a card to explore — click to pin.
First-party data — what customers do. Observed behavior. Nobody fills in a form; the data is generated by actions. A product page viewed, an item added to cart, an order placed, a refund requested, an email opened. The customer acts, your infrastructure records. This is the largest and richest of the three categories, and it's what this entire course is about.
Zero-party data — what customers tell you. Disclosed information. You ask a direct question — a microsurvey about size preferences, a quiz about skin type, a "who are you shopping for?" prompt — and the customer answers. It's smaller in volume but uniquely valuable, because it captures intent and preference that behavior alone can't reveal. It's also a big enough discipline that we've given it its own course in the Academy; we'll only gesture at it here.
Email & SMS contact information — permission to reach them. An email address or phone number, captured with consent. We treat it as its own category because it plays a different role from the other two: beyond being data, it's a channel. There's plenty to analyze here — where contact information gets collected most effectively, how different subscriber offers move collection rates — but its defining property is what it unlocks. Without it, everything you know about a customer can only be activated on your own website; with it, you can reach out. This category, too, gets its own dedicated Academy course.
Three categories, three verbs: what they do, what they say, and whether you can reach them. A strong data foundation collects all three deliberately — but the first one is where most of the volume, most of the complexity, and most of the quiet failure lives.
Where this sits in the system
Readers who've spent time on our site have met the Technical Marketing Pyramid — our map of what a strong technical marketing department is built from. Three tiers: Data Collection at the base, Analytics in the middle, Activation — the layer where money is made — at the top.
The logic of the pyramid is simple: each tier is built on the one below it, and each tier sets a hard limit on what can be built above it. Analytics can only see what was collected. Activation can only act on what analytics can see. The wider the base, the taller the pyramid can become.
This course lives in the base — specifically in its largest block. Everything a brand eventually wants to do at the top of the pyramid — the abandoned-cart flow that actually converts, the churn prediction that actually predicts, the personalized homepage that actually feels personal — is capped, from day one, by decisions made down here.
First-party data, unpacked
"First-party data" sounds like one thing. In practice it's a family of data sources, each generated by a different kind of customer action, each with its own tracking challenges, and each unlocking different things downstream. Here's the tour — we go deep on tracking best practices for each of these later in the course.
Website behavior. The stream of events generated as customers move through your store: page views, product views, searches, filters applied, add-to-carts, checkout steps. Enormous in volume, and the source most exposed to everything that can go wrong in tracking — which is why so much of this course orbits around it.
Campaign engagement. System-generated interaction data from your own marketing: email opens and clicks, SMS clicks, weblayer interactions, push notification responses. Often overlooked as "channel metrics," but at the individual level it's behavioral gold — it tells you which customers are leaning in right now. And it's frequently the largest first-party source by volume, often outweighing website behavior itself.
Transactional data — online and offline. Orders, order values, items, payment methods — from the webshop and from retail, where applicable. And, critically, the part of the transaction story most brands under-collect: refunds and exchanges. A customer's real value isn't what they ordered; it's what they kept. We've seen "top customer" segments quietly harbor serial refunders simply because returns never made it into the data model.
Order lifecycle updates. Shipping confirmations, delivery events, pickup notifications — where they're useful. Not every brand needs these in the customer data layer, but for some use-cases (post-purchase experience, delivery-triggered campaigns) they're the difference between a message that lands perfectly and one that arrives absurdly early or late.
Customer service interactions. Tickets, complaints, chat transcripts, call outcomes. A customer who contacted support twice about a late delivery is in a very different state than their purchase history suggests — and a brand that doesn't see this data will happily send them a cheerful upsell in the middle of a complaint.
Other channels — physical catalog sendouts or a loyalty platform, for example, wherever they play a role in the brand's mix. Who received which catalog, who earned or redeemed which reward, and when — it all belongs in the same customer record as everything digital, or those channels stay forever unmeasurable.
Notice what these have in common: none of them require asking the customer anything. The data already exists the moment the interaction happens. The only question is whether your infrastructure catches it — completely, correctly, and attached to the right person. Each of those three words is harder than it sounds, and each gets its own part of this course.
Why this asset is suddenly visible
Here's the thing about first-party data: it was always this valuable. A complete, clean record of what every customer browsed, bought, returned, and responded to has been the foundation of great retail since long before anyone called it data. What's changing isn't the value — it's how visible that value has become.
Two things are making it visible.
Activation is getting dramatically easier. For years, the honest reason many brands under-invested in their data foundation was that using it was expensive. Every personalization use-case meant briefs, designers, developers, weeks of coordination — so most of the data's potential just sat there, theoretical. AI is collapsing that cost. In our own work, AI-created weblayers and email assets have multiplied the number of use-cases and A/B tests we can ship every month; AI analysts connected to a data warehouse have turned week-long reporting requests into conversations. When activation was the bottleneck, mediocre data was survivable — there was only so much you could do with it anyway. Now the machinery on top runs fast, and it will happily run fast on whatever you feed it.
Which is why data quality is becoming the differentiator. This is the realization we see landing across the industry right now: the constraint has moved down the pyramid. The brands pulling ahead aren't the ones with the cleverest AI on top — everyone increasingly has access to the same models. They're the ones whose base actually holds: events that fire completely and correctly, transactions that include the refunds, identities that connect the anonymous browser to the known customer, analytics that aren't quietly inflated by bot traffic. AI doesn't fix a broken foundation; it amplifies whatever the foundation contains, garbage included.
And unlike the models, this asset is exclusively yours. Your competitors can rent the same AI, run the same ad platforms, copy the same playbooks. They cannot buy, scrape, or replicate the record of how your customers actually behave. First-party data is the one input to the whole system that compounds privately, in your favor, every single day it's collected well — and the one that's silently lost every day it isn't.
Capturing all the data that matters, correctly, so you can activate on it later — that is, quite literally, the topic of this course.
What's ahead
The course runs in two arcs.
Capturing it right:
- Consent & compliance. Before anything else: what you're allowed to collect, consent for tracking vs. consent for communication, and what this means for data completeness.
- Earning first-party status — and its limits. "First-party" is a status browsers grant, not a label you choose. What determines the classification, how to structure tracking so your data genuinely qualifies, and the limitations that persist even then.
- Client-side vs. server-side tracking. What browsers block and why, what server-side actually solves, and what it doesn't.
- Digital identity. The concept the whole system hinges on: connecting an anonymous cookie to a known customer, across sessions and devices.
- The categories, in depth. Best-practice tracking recommendations for each first-party data source — and what each one unlocks downstream in analytics and activation.
Keeping it clean:
- Common tracking errors. The hygiene story: broken events, duplicate firing, schema drift, and the quiet ways tracking degrades.
- Bots in the AI era. The invasion story: how AI-driven traffic inflates and distorts your analytics, and how to detect it.
- Monitoring systems. How we keep tracking honest continuously, instead of auditing it once and hoping.
Everything in it is free. To request access, we'll ask for your email address and a little about you — including the company you work for. And we want to be transparent about why: we grant access to people at eCommerce companies. We don't grant it to agencies. This is the knowledge agencies typically gate, and we're happy to hand it directly to brands — handing it to other agencies is a different matter.
(You'll also notice that collecting your email address is us practicing category three of our own framework.)











