Jev confidence threshold (below it, Sonnet 4.6 decides)
faster per decision
cheaper
Time per decision
Mean wall-clock seconds; the median is in the tooltip.
Cost per 1M decisions
Jev + Sonnet 4.6 at list price on billed tokens.
Accuracy
Agreement with the label. Nothing is given up.
All three setups
Every number is measured, per call, on the same 200 items: Sonnet 4.6 through Evaluator("bedrock/us.anthropic.claude-sonnet-4-6"), and the cascade through CascadeEvaluator(jev="jev/jev-latest", judge=<that Evaluator>, min_confidence=t).evaluate(). Prices: Jev $0.042 per 1M input tokens; Sonnet 4.6 $3 / $15 per 1M tokens in / out. Label: whether Kimi K2-thinking called a tool next, so accuracy is agreement with K2-thinking, not a human judgement. GeneralFunctionCall-Test (evalscope, Apache-2.0), 100 tool-call and 100 reply turns, seed 7. Source: benchmarks/tool_call/bench.py.