We build the evaluations that measure frontier AI.
Evaluations, environments, and ground truth for frontier models and agents.
Backed byCapability is moving faster than the ability to measure it.
Benchmarks saturate, leak into training corpora, or were never hard enough to begin with. The frontier needs instruments built to its own standard.
KT-22 builds those instruments. We publish little and disclose less.
Problems built past the point where current benchmarks saturate: hard enough that frontier models fail them, and reproducible enough to run as RL environments.
Evaluation and rollouts by people who work at that level themselves: olympiad medalists, PhDs, and licensed practitioners, under NDA in isolated, audited environments.
Every task ships with run records from the strongest models available: what they tried, where they broke, measured rather than estimated.
KT-22 is a peak at Palisades Tahoe. In 1948, Sandy Poulsen got down its steep north face the only way she could: traverse, kick turn, traverse again. From the valley floor her husband Wayne counted 22 kick turns, and the mountain had its name. We took it for this lab because the job is the same.
Hard terrain, every turn counted.