Measure it, then trust it.
I'm Josh Longenecker. I build grounded-ai, an open-source library for evaluating what LLM applications say and do. I write here about evaluation, decision models, cutting the cost of AI systems without losing accuracy, and whatever else I find worth measuring.
Latest writing
- Hardware entitlement to roofline: a 7-day curriculum
Take an op and a chip, work out from the documentation how fast the op can run, measure how fast it does run, and explain the difference. Seven days, built on AWS Trainium and Inferentia2, with a check-yourself problem for each.
- Tensor Engine clockwork: one matmul on Trainium, cycle by cycle
An interactive walk through what the Trainium/Inferentia Tensor Engine does on every clock of one nc_matmul: weights parked in a systolic array, inputs streamed in, partial sums falling into PSUM. Plus how the same picture shows up in a Neuron Explorer profile.
- Half the latency, 61% cheaper: cascading a decision model in front of an LLM
CascadeEvaluator lets TypeSafe’s Jev settle the decisions it is sure of and passes the rest to any LLM judge you pick. Against Claude Sonnet 4.6 on 200 real agent turns, it was 51% faster and 61% cheaper, with no loss in accuracy.