Back to workshops
Free workshop90 minStarts 19 Nov
Cutting inference cost without cutting quality
A free 90-minute live session on finding the spend inside an LLM feature and removing it without anyone noticing.
Save your seatBring a prompt and a bill. We will find the spend in it together.
What we cover
- Where the money actually goes, measured rather than guessed
- Caching the part of the prompt that never changes
- Routing by difficulty instead of by default
- Proving quality held, with the eval you already have
Who it is for
- Teams whose model bill grew faster than their usage did
- Engineers asked to cut spend without a quality regression