Skip to content

Free live workshop — Eval-driven development for LLM apps.

Save your seat
Back to workshops
Free workshop90 minStarts 19 Nov

Cutting inference cost without cutting quality

A free 90-minute live session on finding the spend inside an LLM feature and removing it without anyone noticing.

Save your seat

Bring a prompt and a bill. We will find the spend in it together.

What we cover

  1. Where the money actually goes, measured rather than guessed
  2. Caching the part of the prompt that never changes
  3. Routing by difficulty instead of by default
  4. Proving quality held, with the eval you already have

Who it is for

  • Teams whose model bill grew faster than their usage did
  • Engineers asked to cut spend without a quality regression

Ninety minutes is a start. The cohort is the rest.