Measure it, then trust it.
I'm Josh Longenecker. I build grounded-ai, an open-source library for evaluating what LLM applications say and do. I write here about evaluation, decision models, cutting the cost of AI systems without losing accuracy, and whatever else I find worth measuring.
Latest writing
- Half the latency, 61% cheaper: cascading a decision model in front of an LLM
CascadeEvaluator lets TypeSafe’s Jev settle the decisions it is sure of and passes the rest to any LLM judge you pick. Against Claude Sonnet 4.6 on 200 real agent turns, it was 51% faster and 61% cheaper, with no loss in accuracy.