Skip to content

Free live workshop — Eval-driven development for LLM apps.

Save your seat

Writing

Notes from production

What actually broke, what we measured, and what we changed. Written by the people teaching the cohorts, from systems they run.

Featured

Build the eval harness before you build the feature

Teams that ship reliable LLM systems almost always wrote the measurement first. Here is what that looks like in practice, and why the ordering matters more than the tooling.

AI Engineering · 21 Aug 2026 · 1 min read

Read the article