The manager — Step Functions and sagas
Some work is too important to leave to a crowd: taking a payment, cooking, and refunding if the kitchen runs out of ingredients. Here a manager helps — someone who knows the whole plan, tells each human what to do, waits for results and decides what happens when something goes wrong. AWS Step Functions is that manager.
What a manager does
- Knows the steps and their order.
- Waits — seconds, hours or days (Standard workflows can run for up to a year).
- Retries failed steps with backoff.
- Branches on results ("if payment fails, notify the customer").
- Runs steps in parallel.
- Records every step, so you can see exactly where a process is — or where it failed.
The order lifecycle as a workflow
Start
→ Cashier: authorise payment ──(fails)──► Messenger: "payment failed" → End
→ Chef: cook order ──(out of ingredients)──► Cashier: refund (apology) → Messenger: "sorry" → End
→ Wait for "order ready"
→ Messenger: "your food is ready!"
End
Sagas: apologising instead of rolling back
In one database, a transaction can roll back everything. Across many humans, each with their own memory, you can't. Instead you use a saga: if a later step fails, run compensating actions for the earlier ones — like a refund after a payment. It's how people handle mistakes: you can't undo, but you can apologise and make it right.
| Step | Compensation |
|---|---|
| Authorise payment | Void or refund the payment |
| Reserve a table | Release the table |
| Send confirmation | Send a correction message |
The manager in CDK
import * as sfn from "aws-cdk-lib/aws-stepfunctions";
import * as tasks from "aws-cdk-lib/aws-stepfunctions-tasks";
const authorise = new tasks.LambdaInvoke(this, "Cashier: authorise payment", {
lambdaFunction: cashierAuthoriseFn,
outputPath: "$.Payload",
});
authorise.addRetry({ errors: ["States.TaskFailed"], maxAttempts: 2, backoffRate: 2 });
const cook = new tasks.LambdaInvoke(this, "Chef: cook order", {
lambdaFunction: chefCookFn,
outputPath: "$.Payload",
});
const refund = new tasks.LambdaInvoke(this, "Cashier: refund (apology)", {
lambdaFunction: cashierRefundFn,
outputPath: "$.Payload",
});
const notifyReady = new tasks.LambdaInvoke(this, "Messenger: food is ready", { lambdaFunction: messengerFn });
const notifySorry = new tasks.LambdaInvoke(this, "Messenger: sorry", { lambdaFunction: messengerFn });
cook.addCatch(refund.next(notifySorry), { errors: ["OutOfIngredients"] });
const definition = authorise.next(cook).next(notifyReady);
new sfn.StateMachine(this, "OrderManager", {
definitionBody: sfn.DefinitionBody.fromChainable(definition),
tracingEnabled: true,
});
Step Functions can also call many AWS services directly (DynamoDB, SQS, SNS, EventBridge and more) without a Lambda in between — fewer cells to maintain.
Standard vs Express
| Standard | Express | |
|---|---|---|
| Duration | Up to 1 year | Up to 5 minutes |
| Execution | Exactly-once workflow steps, full history | At-least-once, high volume, cheaper per execution |
| Use | Business processes, human approvals, long waits | High-throughput event processing |
When to hire a manager
| Use orchestration (a manager) when… | Use choreography (town square) when… |
|---|---|
| The process has strict order and compensation | Humans react independently to facts |
| You need visibility of each case's progress | Adding new reactions should need no central change |
| There are timeouts, waits or human approvals | The flow is simple or open-ended |
A manager makes complex processes visible and recoverable. Use one for critical, multi-step business flows — and keep every human's own job inside that human.