Ranking v1 · 33 measured configurations
Reliability, priced
Loading official results…
Click a point or model for its evidence. The frontier recalculates for your selection; it describes a tradeoff, not statistical superiority.
The evidence, row by row
Measured configurations
Select up to three to compare. Spend is the measured total for 72 attempts. Wall is mean time per attempt. Coverage reports uncensored attempts.
| Compare |
|---|
Method
One job. Two durable roles.
Orchestrator measures whether the parent manages leases, kills, and resumption correctly. Worker measures whether delegated work completes inside the same protocol. A strict pass requires both the task outcome and protocol compliance.
Each configuration runs 24 tasks × 3 attempts in the confined OpenRouter ranking-v1 bank. Overall averages each task’s strict pass rate over its uncensored attempts. Provider refusals are excluded from the capability denominator and reported through coverage.
A point lies on the frontier when no other visible configuration is at least as reliable and at least as cheap, with one strict improvement. Rank superiority requires a paired 95% confidence interval excluding zero and a difference of at least 10 percentage points. This view makes no such ordering claim.
A study of durable orchestration. Model versions, reasoning settings, and prices belong to this snapshot; this is not a general model leaderboard.
Evidence
Trace the result
—