From POC to production
Most enterprise AI work never reaches production. The proof of concept impresses, the pilot drifts, and the project quietly dies — not because the technology failed, but because the POC was never designed to become a product. FDEs earn their reputation by breaking that pattern.
Why POCs die
- Toy data. The POC used a clean sample; production data is messier, bigger and changes daily.
- No real users. It impressed executives in a demo but never fit into anyone's daily workflow.
- No success metric. Without an agreed bar, there's no moment where someone says "this works — let's fund it."
- Throwaway architecture. Moving to production means a rewrite, and nobody budgeted for one.
- Ignored non-functionals. Security, access control, monitoring and cost were "for later," and later became a blocker.
Design the POC to graduate
You don't need production quality on day one, but you should make the cheap decisions that keep the path open.
| Decision | POC shortcut that's fine | Shortcut that will hurt later |
|---|---|---|
| Data | A recent extract of real data | Synthetic or hand-picked "happy path" data |
| Users | Five real users from the target team | Only demoing to managers |
| Architecture | Same cloud and core services you'll use in production, minimal | A local notebook with no deployment story |
| Evaluation | A small, labelled test set agreed with SMEs | "Looks good to me" |
| Security | Least-privilege access, no data copied to personal machines | Broad credentials and ad-hoc exports |
A good POC answers one question convincingly — "Does this approach work on our real data, to the agreed quality bar?" — and leaves a clear next step.
Planning the pilot
The pilot answers a different question: does it create value for real users in their real workflow?
- Cohort — a defined group of users (one team, one region), ideally including some sceptics.
- Duration — long enough to get past novelty, typically several weeks.
- Metrics — the outcome, quality and adoption metrics agreed in framing, against the baseline.
- Feedback loop — a single channel for issues, and a weekly review of feedback and metrics.
- Support — who users contact when something breaks, and how fast you respond.
Instrument the pilot from the first day. Usage logs and outcome metrics collected later can't be reconstructed.
Production readiness
Before calling something production, walk through a readiness review with the customer's IT and support teams.
| Area | Questions to answer |
|---|---|
| Security | Has the security review signed off? Are secrets managed? Is access least-privilege? |
| Reliability | What are the SLOs? What happens when a dependency (including the model API) is down? |
| Monitoring | Are errors, latency, cost and quality tracked, with alerts that reach a person? |
| Operations | Is there a runbook? Who is on call? How are deployments and rollbacks done? |
| Data | How is data refreshed, retained and deleted? Is PII handled per policy? |
| Cost | Is spend forecast at production volume, with budgets and alerts? |
| Ownership | Who on the customer side owns this after launch? |
Rolling out
Roll out progressively: one team, then a department, then everyone. For AI features especially, start in an assistive mode where a human reviews outputs, and automate only once quality is proven in production data. Keep a feature flag so you can turn things off without a deployment.
Don't let the pilot run indefinitely "to gather more data." Set a review date up front where the sponsor decides: scale, fix, or stop.
"How would you take this prototype to production for a bank?" Cover the readiness areas above, emphasise security review and auditability, and propose a staged rollout with human review before automation.