Each layer solves a different problem, and the sequence cannot be skipped:
Provides the raw capability for reasoning and generation. A necessary condition, not a sufficient condition—the same model with a different Harness can show a 20–30 percentage point difference in measured performance.
Turns the model's "intent" into real executable actions—calling tools, maintaining context, handling loops; this is the first conversion layer where capability lands.
Encodes the experience, judgment, and tools of an X expert (e.g., cyber security; can be replaced with any vertical industry) into the execution environment, determining whether it can "take on the job."
For the same Harness capability, who it is delivered to determines the magnitude of value realized—the most easily overlooked yet most return-determining环节.
EXAMPLE · X = CYBER SECURITY
The same vulnerability intelligence unearthed by the same Harness yields completely different realized value depending on the delivery path: delivered to a large enterprise like Microsoft, becoming a direct input for its security hardening and patching process, the magnitude is tens of billions; delivered to mid-sized and small clients or long-tail channels, the same intelligence can only realize a fraction of the value, with a magnitude of billions. This shows that the Harness itself does not equal commercial value—the choice of delivery path is the switch that determines the magnitude of value.
Business results either become a data flywheel or are wasted. The chain is: ① Business Feedback (real customer usage results—whether intelligence is verified as effective, not whether the model thinks it answered correctly) → ② Reward Chain (verifiable parts go through RLVR for automatic scoring; ambiguous parts go through human preference/KPI evaluation, beware of reward hacking) → ③ Data Flywheel (real execution trajectories沉淀 as training and evaluation datasets, the more used, the more accumulated, the more accurate) → ④ Feedback Iteration (verifiable tasks fine-tune base model weights; ambiguous tasks optimize Agent/Harness engineering)—then back to the first layer, the closed loop is complete.
The model is a necessary condition, not a sufficient condition; the Harness determines whether expert capability can be encapsulated and executed; and the delivery target and feedback closed loop are the two keys that determine whether this chain can truly realize value and continuously self-reinforce. When evaluating any "AI落地" narrative, first ask whether these four layers of questions are each answered correctly, rather than just looking at model benchmark scores.
First published 2026-07-24