About

I'm Josh Longenecker. I build grounded-ai, a Python library with one interface for evaluating LLM outputs across providers: hosted judges (OpenAI, Anthropic, Bedrock), local models, and decision models such as TypeSafe's Jev.

This site is where I write up what I measure: which evaluations hold up, what they cost, how fast they run, and what I learn along the way, plus other topics I find interesting.

Code is on GitHub. New posts are in the RSS feed.