Shipping LLM features for customers
Most FDE engagements today involve large language models. The difference between a demo and a deployment is not the model — it is everything around it: choosing the simplest pattern that works, proving quality on the customer's data, and rolling out in a way users trust.
Choose the simplest pattern that works
| Pattern | Good for | Cost of getting it wrong |
|---|---|---|
| Prompting with structured output | Classification, extraction, drafting from provided text | Low — easy to iterate |
| Retrieval-augmented generation (RAG) | Answering from the customer's documents and knowledge | Medium — retrieval quality dominates results |
| Tool use / function calling | Taking actions or querying systems | Higher — needs permissions and safeguards |
| Agents (multi-step tool use) | Open-ended tasks across several systems | Highest — harder to test, slower, costlier |
| Fine-tuning | Consistent format or style at scale, specialised domains | High — data, maintenance and evaluation burden |
Start at the top and move down only when evidence says you must. Most customer problems are solved by good prompting plus good retrieval. For the retrieval side, see the RAG Handbook.
Evals are your acceptance criteria
You cannot ship an LLM feature responsibly without an evaluation set — a collection of realistic inputs with known-good outputs or grading criteria, built with the customer's subject-matter experts.
- Source it from reality — real questions, real documents, real edge cases from discovery.
- Grade consistently — exact match for extraction, rubrics for open-ended answers, SME review for a sample.
- Run it on every change — prompt edits, model upgrades and retrieval tweaks can all regress quality.
- Report it in business terms — "92% of answers correct and complete, 0 fabricated policy references."
When the customer asks "how accurate is it?", the only good answer is a number from an evaluation set they helped build. Everything else is opinion.
Design for being wrong
LLMs will sometimes be wrong. Good deployments make that safe:
- Show sources so users can verify answers.
- Express uncertainty — let the system say "I couldn't find this" instead of guessing.
- Keep a human in the loop where errors are costly; make review fast.
- Make feedback one click — thumbs up/down with an optional reason feeds your eval set.
Guardrails
- Validate structured outputs against a schema; retry or fall back on failure.
- Restrict tools to least privilege, and require confirmation for consequential actions.
- Treat retrieved documents and user input as untrusted — defend against prompt injection.
- Filter or redact sensitive data according to policy.
Cost and latency budgets
Agree budgets early: target response time and cost per request at production volume. Measure tokens per request in the pilot, then use caching, smaller models for simpler steps, and tighter context to stay inside budget.
Staged rollout
- Shadow mode — the system runs on real traffic, but outputs are only logged and evaluated.
- Assist mode — users see suggestions and decide.
- Automate selectively — automate only the cases where measured quality is consistently high, and keep humans on the rest.
Model providers update models. Pin versions where you can, and re-run your evaluation set before switching — otherwise a "free upgrade" can silently change behaviour for your customer.
"The customer says the assistant gives wrong answers sometimes. How do you fix it?" Start with data: collect the failures, classify them (retrieval missed, wrong source, reasoning error, unanswerable), fix the biggest category, and prove it with the eval set.