What the model gets to see

ᐧ

You open Netflix on a Sunday evening. You haven't searched for anything, but it already has a few ideas about what you might want to watch.

Maybe you spent the last two evenings on a crime series. Last week you finished a crime film. You gave one title a thumbs-up and abandoned another after twenty minutes. Yesterday you watched something on your phone. Tonight you're on the TV.

Netflix knows all of this, and much more.

But the better question isn't what Netflix knows about you. It's this:

What does Netflix choose to tell the model about you before the model decides what to recommend?

"The model" is the AI that does the choosing. A recent Netflix research paper on GenRec, an AI-driven recommendation ranker, shows how much that choice matters, and how much of it is a product decision rather than an engineering one.

From measuring you to describing you

Older recommendation systems represent you through measurements.
Crime genre: strong. Recent viewing: high. Completion: high. Device: TV.

Engineers decide what is worth measuring, turn behaviour into those signals, and build systems that know how to combine them. That is feature engineering.

Netflix's paper says its production systems rely on large numbers of hand-crafted signals and specialised designs, and that this has become hard to extend. Supporting a new type of content or a new screen often takes substantial work.

GenRec takes a different route. It writes you down. History, details about the titles, and the situation you're in are turned into text the model reads. Something closer to:
This person watches a lot of crime, tends to finish what they start, liked one film, and is on a TV on a Sunday evening.

That sentence is mine, not the paper's. The paper describes turning history, item information and context into natural language or lightly structured text.

The traditional system represents you through measurements. GenRec gives the model a richer description of what you've been doing and lets it infer which relationships matter. The authors say they rely on the model to learn those higher-level patterns instead of encoding them by hand.

Netflix calls this a shift from feature engineering to context engineering.

traditional recommender vs genrec

The old question was what to measure about the person. The new one is what to tell the model about the person.

What gets left out

Your history is far too large to hand over whole. The paper notes that Netflix members generate hundreds of billions of interaction events across viewing, playing, thumbing and adding to lists. So GenRec has to choose.

As the paper describes it, long plays and thumbs-up are kept with richer detail. Short plays and noisy clicks can be left out. Repetitive behaviour, like binge-watching, is compressed instead of listed event by event. New or important titles can get more description, because the model knows less about them. Older history is either dropped or folded into a short summary of interests.

Back on the couch, that means three episodes of the same series don't need three equally detailed entries. A three-minute play may not deserve the weight of a film you finished.

Genrec

Caption: As the paper describes it. Not what happens to any one viewer.


These sound like engineering choices. They are. They are also decisions about what counts as meaningful behaviour.

A three-minute play could mean you hated it. It could mean the doorbell rang. It could mean you were only curious. The system doesn't know the reason from that event alone. It has to decide whether that moment is useful enough to influence what comes next.

That is a judgment about you, made in advance, at scale.

There is a practical reason for all this selection. The model processes text in small units called tokens. The more it reads, the more it costs to run.

Netflix tested how much history GenRec really needed. The paper reports cutting the context from roughly 5,000 tokens to about 1,700, with negligible degradation in offline ranking quality, meaning tests on past data. Serving cost fell to roughly a third.

This doesn't mean less is always better. It means that in this system, on this task, much of what could have been included wasn't helping the decision. Extra information can also dilute the model's attention, making the rest of it less useful. The authors' own tests found a point beyond which more history added little.

More information isn't automatically more useful information. Information has a cost, and it isn't only the bill. The question isn't how much we can give the model. It's what is worth giving it.

Memory is also selection
Netflix doesn't call this memory. I do.

We usually talk about AI memory as a capability. Can it remember? How much? For how long? But look at what GenRec does: some history stays sharp, some is blurred into a pattern, some never reaches the model. That is memory too, built from selection. To remember something for a decision is to keep it available to that decision. To forget it is to leave it out. To summarise it is to decide which parts survive.

One clarification, so as not to overstate. Leaving something out of what the model reads isn't deleting it. The paper describes how the model's input is built, not erasure of the underlying records. But from the model's side, what isn't in the description doesn't exist for that decision.

If AI products can remember, who decides what they should remember?

And the question that gets asked far less: who decides what they are allowed to forget?

Take it beyond Netflix. An assistant that remembers your preferences. A support agent that reads your earlier conversations before replying. In each, before any decision happens, someone chose which past details reach the system, which get shrunk to a line, and which are treated as noise. Usually nobody chooses on purpose. Whatever the pipeline kept, or whatever fit, becomes the policy.

The model is not the product

It would be easy to conclude that Netflix added an AI and recommendations improved. That isn't what the paper describes.

The authors note that off-the-shelf language models tend to over-recommend what's globally popular, invent titles that don't exist and ignore business constraints. So the system starts from a model adapted to Netflix's catalogue and behaviour, is trained specifically for ranking, and is limited to titles Netflix actually has. It is also rewarded for more than the next click. The paper states its target as long-term member satisfaction rather than short-term engagement alone, using indirect signals such as whether people return to the service, because the real long-term outcomes are slow and noisy to measure. Other rewards help balance exposure across movies, series, games and live content.

None of that comes from the model. Cost and speed set the edges of what's possible, but inside those edges, what counts as a good recommendation depends on what the product is for. Reward the next click and a three-minute play may look different than it does when you're rewarding the viewer who is still here next month. Someone has to decide what "good" means.

The results are interesting, but worth keeping in perspective. Offline, GenRec showed about a 1.6% relative improvement in ranking quality over Netflix's existing system, using roughly 40 times fewer labeled training examples at the final training stage. In an online A/B test on about 10% of traffic for four weeks, on particular batch-compute surfaces and in a deliberately low-data configuration, Netflix reported statistically significant improvements on both short-term and long-term metrics, including a +0.006% relative improvement on its core online metric, which the authors call meaningful at their scale.

The paper presents GenRec as an initial step, to be tested rigorously against strong non-LLM systems, not as proof that this approach is better. The interesting part, to me, isn't simply that the numbers moved. It's where the work moved: from designing more and more features to deciding what context the model receives.

The model is not the product. The decisions around the model are part of the product.

Upstream

Product teams have always decided what users see. The default option, what gets prominence, what stays hidden. We argue about it openly.

The product job hasn't disappeared. Part of it has moved upstream.

Teams still own the person, the experience and the outcome. But there is another layer underneath. What does the model know about this person, and what it doesn't know? What are we assuming about someone's behaviour when we call something noise? What are we rewarding? How would we notice a bad decision? And what does a better answer cost?

Those questions don't require a product manager to become an engineer. They require knowing where the decisions are being made. That is my interpretation, not Netflix's conclusion.

Now go back to the couch. A title sits at the top of your screen. You can see the recommendation. You can't see the history that shaped it: which moments were kept, which were left out, which were folded into a summary, which the system decided mattered.

The product decision is no longer only what the user sees.

It is also what the model gets to see.

And increasingly, that decision is invisible to the person using the product.

So when the model makes the decision, who decides what the model gets to know?

/a


Source: Li et al., “GenRec: An LLM-Backed Recommendation Ranker at Netflix,” arXiv:2608.10257v2. Diagrams are my own, based on the paper; interpretation is mine.



|