How Does Real-Time Prediction Work in Retail Media?

How real-time prediction works in retail media: scoring customer behavior at the first mile, before it ever reaches the warehouse.

How Does Real-Time Prediction Work in Retail Media?

Share with others

I asked our Head of AI & Product to explain it like I'm five.

The whole industry is arguing about agentic commerce right now. Is it existential? Do off-site retail media businesses survive in a world where agents do the shopping?

Andrew Lipsman has been the loudest voice on the "everyone calm down" side. His argument, roughly: we keep over-crediting what the technology can do and under-crediting what real humans will do. And the variable agents are worst at is the one that decides most purchases in the first place — context. Not what you bought last time. What's going on with you right now.

I agree with him. My own read is that the curve is bending, not exploding. ChatGPT more than doubled its users this year, sure. But the growth is clearly decelerating, and the market's fragmenting across four or five assistants instead of sprinting toward one agentic future. So I'm not losing sleep over the timeline.

But there was one word I kept nodding along to in meetings without fully getting it. Inference.

Small confession. When I tested our website messaging with people outside MetaRouter, more than one told me they'd never heard the word before. Which is either a problem or an opportunity. I filed it under opportunity, and then realized I owed it to myself to actually understand the thing I was about to go educate people on.

So I did what I do when I'm out of my depth. I booked time with someone smarter than me. Patrick Harrington runs AI and Product here at MetaRouter. I asked him to explain how real-time prediction works like I'm five. Then I made him do it again.

Here's what I finally got.

What is real-time inference?

Scoring every event the moment it happens, before it lands in a database, in order to predict what someone will do next. That is the whole definition, and everything below is why it turns out to be harder and worth more than it sounds.

Where does the scoring happen?

Every event on a retailer's site — searches, hovers, how long you lingered, your device, your time zone, where you came in from, hundreds of inputs — moves through pipes before it lands anywhere. Before the CDP, before Snowflake or Databricks, before it is a row in anyone's database. Patrick's version is the plumbing under your sink, which catches the water before it ever reaches the city system. Most sites score behaviour after it has arrived in the warehouse, which is minutes or hours later. First-mile inference scores it while it is still moving.

What does the model actually see?

A stranger lands on the site, not logged in and with no email attached. She searches white dress, looks at one, then another, then a third. Patrick described it as watching someone add dresses to a closet one hanger at a time, where each hanger tells you a little more about where the session is heading and with what intent.

I told him I am a visual person and needed an actual robot in my head to picture the mechanism, so he gave me R2-D2. Raw behaviour goes in and a score comes out, then new behaviour arrives and a new score replaces it. A hundred anonymous people on the site means a hundred small scores all recalculating at once, based on what those hundred humans — or bots, because it clocks bots too — are doing right now.

Why do two people searching white couch see different things?

Today they don't, which is the part that clicked for me. Same query, same sponsored listings. But I came in from a Google search, on my phone, in Lisbon, at 9am, and you typed the URL directly on a laptop in Boulder after already looking at nine things. Same words, different person. Inference folds in the route each of us took to that search rather than only the search itself, which means the context is doing the deciding — and we happen to see that context before anything else in the stack does.

Does the same product mean the same thing at a different hour?

No, and this is my favourite thing Patrick said. At a storage company there are people searching for units at three in the morning, and one of them may have just had a fight and needs somewhere to put their life by sunrise. That is a completely different customer from the one browsing calmly at ten on a Saturday night, even though the product is identical. The context is the entire story, and it is the piece that an agent can pick up on in a conversation but cannot act on inside your store, which is a gap that belongs to you for as long as you bother to hold onto it.

Why does first-mile behaviour stay scarce?

The AI labs already bought the internet, so every catalogue, every SKU and every image is baked into the training data. They are not blind to intent either — every question you type into a chat tells them more about yours than the one before it did. What they don't have is your catalogue sitting next to your conversion history, and that pairing is what turns a signal about intent into an actual prediction about a purchase. It only exists on your own site. Which leads to the least subtle sentence in this post: don't put an AI lab's pixel on your website.

What does this buy a retailer?

Three things, and the order matters more than I expected. Speed rather than smarts, because a prediction that arrives while the person is still on the page is worth something and the same prediction tomorrow morning is worth almost nothing. Cost, because scoring in the pipes means you are not shipping every event to Snowflake or Databricks in order to ask the question there, and you are not standing up a team of AI engineers to go and ask it. And control, because the prediction gets made with identity and consent already attached, inside your own cloud, so less sensitive data has to leave the building in order to be useful — the same rules you already apply to collection, now applied to the model.

Telling a human from a bot before you have paid to serve the bot is only one of the predictions you can run, because it is the same model answering whatever you have trained it on. Is this person here to buy, to check the opening hours, or to apply for a job at the deli counter?

Does any of this depend on agentic commerce happening?

No, which is the part I find clarifying, because the whole industry is arguing about whether agentic commerce is existential and whether off-site retail media survives a world where agents do the shopping. Andrew Lipsman has been the loudest voice on the "everyone calm down" side. He called the thing a collective hallucination, took a fair amount of incoming for it, and has mostly been proven right — instant checkout is dead, and Walmart reported conversion running about three times worse through it than through a plain click-out. AI search sends top retailers somewhere between a tenth of a percent and two percent of their traffic, against roughly sixty percent that arrives direct and isn't going anywhere.

The behaviour that did move is the more instructive number. Sensor Tower's holiday analysis put forty percent of Amazon's transactions as having touched Rufus along the way, which is the intelligence coming to the site rather than the shopper leaving it, and it is what Lipsman was getting at when he asked why we assume traffic will move upstream into the LLMs "rather than the LLM technology moving onto the site?" Meanwhile more than half of all web traffic is already automated, at 53% in Imperva's 2026 report, up from 51% the year before. The industry is bracing for a two percent channel and shrugging at the fifty-three percent that has already arrived.

Human to computer, human to AI, AI to AI. Those are the three relationships, traffic is going to keep sliding across all of them in ways none of us can time, and the work underneath doesn't change either way. Capture the behaviour at the first mile, in your own cloud, before the moment is lost, and predict on it in real time — will they buy, is this a human, is this fraud. That holds whether the visitor is a person, a person's agent, or an agent talking to another agent.

So what is actually defensible?

Not software. That used to take fifty million dollars and fifty engineers, and now it takes a twenty dollar subscription and six months. Context is what is left, which means seeing it before anyone else does and owning where it lives. Patrick put it more plainly near the end, and I keep chewing on it, because once you have heard the whole thing it sounds almost obvious: "why wouldn't we be doing this?"

The honest answer is that you can't RFP for it yet. There is no line item called first-mile context, no box to tick and no grid to score, and Lipsman keeps asking what the single unifying metric for the middle of the funnel is going to be without anyone answering him. I don't have that metric either. But I would bet it starts with whoever is capturing the context the middle of the funnel actually runs on, which is why the retailers who win the next few years won't be the ones who found the best vendor on a checklist. They will be the ones who understood the shift before it had a name, and started capturing the one asset nobody can buy back later.

It took me a dress, a couch, a storage unit at three in the morning and a small robot to get there, which is a little embarrassing for a CMO.

Learn more about Metarouter Inference solutions: https://www.metarouter.io/inference