Liang, Garg & Moghaddam
Measurement Research
Models memorized the benchmark, not the skill
Category
Date
Share
Most evaluation tooling assumes the agent is already deployed. That creates a chicken-and-egg problem: trust is required to deploy, but the measurement signal only appears after deployment. Hue builds the environment first. We reconstruct the workflow from real operational data and the people who perform it, then expose models and agents to the full distribution of cases before they touch live work. The result is evidence at the moment enterprises actually need it: before the bet is made.

