How do we keep LLM costs under control without losing quality?
Short answer
Start with evaluations that define what good enough means for each task. Then use the least expensive model that passes them, and keep the most capable models for the requests where they make a measurable difference. Track cost per request alongside quality, so savings never quietly cost you accuracy. Where tasks are too varied for formal evaluations, rely on thorough manual QA and be explicit about the risk that leaves.
Evaluations come first
Without evaluations, every model choice is a guess, and the safe-looking guess is the most expensive model. Build a set of real examples with expected results, and run it on every change to prompts, models or retrieval.
When formal evaluations do not fit
Some tasks are too varied for a fixed set of examples to mean much: every request is different, and what good looks like depends on context. Forcing an evaluation onto that work gives false confidence.
There, I develop with a great deal of manual QA: reviewing real outputs by hand, sampling production use and listening to the people who rely on it. That is weaker evidence than a formal evaluation, so I say so explicitly, record the risk and agree with the business how much of it to accept.
Choose models by task, not by default
- Route simple, high-volume requests to smaller, less expensive models
- Send classification-style decisions, such as routing, triage and scoring, to purpose-built decision models like TypeSafe’s Jev, which return typed answers with confidence scores, rather than to a general-purpose LLM
- Keep the most capable models for tasks where evaluations show they are needed
- Re-run evaluations when providers release new models; the right answer changes
- Cache and reuse results where the same question is asked repeatedly
Measure cost and quality together
Report cost per request, or per successful outcome, next to evaluation scores and customer-facing quality. A cost cut that lowers quality is not a saving.
I have spent much of my career on cost, including cutting Remo’s infrastructure and vendor spend by more than 80% and Deepnote’s server costs by 40%. The same discipline applies to models.
Related questions
Should we build our own model?
Usually not. I have developed custom machine learning with data scientists for years, but most problems are now solved with existing models, tuned and customised for the task. I remain open to research-based work, particularly on making custom models more cost-effective.
How often should we revisit model choices?
Whenever providers release new models or your usage changes, and at least every few months. Evaluations make that review quick.