Finer-grained outcome parameters expose latent flaws.
To efficiently probe the frontier, we prefer highly discriminating question formats. Binary thresholds lazily collapse nuanced beliefs; by converting to continuous or multiple-choice formats, we expose underlying algorithmic disagreements that drive model improvement.
- More discriminating question formats
- Rejecting simple binary thresholds
- Exposing latent predictive variance








