
LLM benchmarks: what each bench actually measures
Capability evals and systems benches answer different questions. A plain map of LLM benchmarks for eng leads choosing what to measure before they pick a model or agent stack.
Read the articleNotes on the technology decisions behind the systems we build.

Capability evals and systems benches answer different questions. A plain map of LLM benchmarks for eng leads choosing what to measure before they pick a model or agent stack.
Read the articleWhen should an agent use a decision model vs an LLM? A practitioner look at TypeSafe Jev use cases for tool approval, routing, and typed Choice/Score outputs.
Releases got faster and bugs got sneakier. How QA roles, tools, and techniques moved from classic automation to AI-assisted testing, eval harnesses, and the work we still keep human.
Agentic coding made generation cheap. The bill moved to CI, browser gates, review tokens, and avoidable reruns. How we cut that tax without softening quality.
Nest, Laravel, Django, ASP.NET Core, and Spring Boot compared for coding agents. Not AI SDKs. Golden paths, types, CLI, magic, and what public benches actually measure.