Start with the capability gap

Tiancheng XuWorking draft

A useful model strategy connects valuable work, a specific failure, and a plausible way to teach the missing capability.

Data markets are often described through categories: expert demonstrations, human feedback, work traces, synthetic tasks, and reinforcement-learning environments. Those categories describe how data is produced. They do not, by themselves, explain what a model will learn.

A more useful starting point is a task the model cannot yet perform reliably.

Make the gap specific

“Improve at knowledge work” is too broad to guide a data investment. “Resolve conflicting instructions across several documents, ask for missing information, and produce a verifiable recommendation” describes a more useful target.

The target suggests what an example needs to contain. A polished final answer might teach presentation. A trace of the work might reveal how conflicts were identified. An interactive environment could test whether the model notices a change and revises its plan.

Each format supplies a different learning signal. Choosing among them requires a hypothesis about the failure.

A domain is more than a label

Healthcare, finance, law, and software each contain many kinds of work. Some tasks have clear answers and rich feedback. Others have delayed outcomes, disputed judgments, or consequences that make realistic experimentation difficult.

To compare candidate domains, ask:

This framing can reveal similarities across industries. Resolving inconsistent records, maintaining constraints over a long workflow, and knowing when to ask for clarification may be shared capability problems, even when the surrounding documents look different.

Realism and control serve different purposes

Recorded work can expose interruptions, tacit assumptions, and recovery from mistakes. It can also contain missing context and habits we would not want a model to imitate.

Constructed tasks make it easier to control difficulty and judge outcomes. They can also simplify away the very ambiguity that makes real work hard.

The useful question is which properties of the real task must survive the transformation into training data. An environment does not need to reproduce every button or every document. It does need to preserve the decisions whose consequences we want the model to learn.

Close the learning loop

A disciplined sequence is to observe failures, define a capability hypothesis, collect or construct targeted data, and test on held-out tasks that preserve the original challenge.

If the score improves only on closely related examples, the evidence for broader progress remains weak. If the capability transfers, the result can inform both the next data investment and the product experiences now worth attempting.

The most useful unit of strategy is a capability that unlocks valuable work, together with evidence that we can teach and measure it.

← All essays

Related: Context is part of the product