Jev confidence threshold (below it, Sonnet 4.6 decides)
faster per decision
cheaper

Time per decision

Mean wall-clock seconds; the median is in the tooltip.

Cost per 1M decisions

Jev + Sonnet 4.6 at list price on billed tokens.

Accuracy

Agreement with the label. Nothing is given up.

All three setups

Every number is measured, per call, on the same 200 items: Sonnet 4.6 through Evaluator("bedrock/us.anthropic.claude-sonnet-4-6"), and the cascade through CascadeEvaluator(jev="jev/jev-latest", judge=<that Evaluator>, min_confidence=t).evaluate(). Prices: Jev $0.042 per 1M input tokens; Sonnet 4.6 $3 / $15 per 1M tokens in / out. Label: whether Kimi K2-thinking called a tool next, so accuracy is agreement with K2-thinking, not a human judgement. GeneralFunctionCall-Test (evalscope, Apache-2.0), 100 tool-call and 100 reply turns, seed 7. Source: benchmarks/tool_call/bench.py.