Forward Deployed Playbook field guide for FDEs Bipin Singh
Build & deliver

Shipping LLM features for customers

3 min readChapter 11 of 20By Bipin Singh

Most FDE engagements today involve large language models. The difference between a demo and a deployment is not the model — it is everything around it: choosing the simplest pattern that works, proving quality on the customer's data, and rolling out in a way users trust.

Choose the simplest pattern that works

Pattern Good for Cost of getting it wrong
Prompting with structured output Classification, extraction, drafting from provided text Low — easy to iterate
Retrieval-augmented generation (RAG) Answering from the customer's documents and knowledge Medium — retrieval quality dominates results
Tool use / function calling Taking actions or querying systems Higher — needs permissions and safeguards
Agents (multi-step tool use) Open-ended tasks across several systems Highest — harder to test, slower, costlier
Fine-tuning Consistent format or style at scale, specialised domains High — data, maintenance and evaluation burden

Start at the top and move down only when evidence says you must. Most customer problems are solved by good prompting plus good retrieval. For the retrieval side, see the RAG Handbook.

Evals are your acceptance criteria

You cannot ship an LLM feature responsibly without an evaluation set — a collection of realistic inputs with known-good outputs or grading criteria, built with the customer's subject-matter experts.

Key idea

When the customer asks "how accurate is it?", the only good answer is a number from an evaluation set they helped build. Everything else is opinion.

Design for being wrong

LLMs will sometimes be wrong. Good deployments make that safe:

Guardrails

Cost and latency budgets

Agree budgets early: target response time and cost per request at production volume. Measure tokens per request in the pilot, then use caching, smaller models for simpler steps, and tighter context to stay inside budget.

Staged rollout

  1. Shadow mode — the system runs on real traffic, but outputs are only logged and evaluated.
  2. Assist mode — users see suggestions and decide.
  3. Automate selectively — automate only the cases where measured quality is consistently high, and keep humans on the rest.
Watch out

Model providers update models. Pin versions where you can, and re-run your evaluation set before switching — otherwise a "free upgrade" can silently change behaviour for your customer.

Interview

"The customer says the assistant gives wrong answers sometimes. How do you fix it?" Start with data: collect the failures, classify them (retrieval missed, wrong source, reasoning error, unanswerable), fix the biggest category, and prove it with the eval set.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I turn ambiguous customer problems into production systems — cloud, data and AI.

Work with me