Living Systems Handbook microservices as humans Bipin Singh
Lifestyles

The rat race

3 min readChapter 23 of 34By Bipin Singh

The rat race: everyone rushing, deadlines everywhere, huge crowds arriving at once. In software, think of a flash sale at 12:00, concert tickets going on sale, trading at market open, or a food-delivery app at 8 p.m. on Friday. These systems face extreme spikes, and when one part gets stressed, the stress spreads — like burnout passing through an overworked team.

How rat-race systems fail

Human burnout System failure
Too many demands at once Traffic spike exceeds capacity; throttling and timeouts
One stressed person stresses everyone they work with Cascading failure: a slow service makes its callers slow, and so on
Panic retries ("I'll just ask again… and again") Retry storms multiply load on a struggling service
Waking up slowly when suddenly needed Cold starts during a sudden spike
No time to think Synchronous chains with no buffer

Survival habits for the rat race

1. Don't let the crowd in all at once — queue them

Put a queue between the gate and the busy humans. Accept the order instantly ("you're in line, order #42"), then process at a sustainable pace. The queue is a waiting room that never loses anyone.

Customers ─► API Gateway ─► Waiter (validate, save, enqueue) ─► SQS ─► workers at a steady rate

For extreme events (ticket drops), a virtual waiting room in front of the site lets visitors in gradually.

2. Set boundaries — throttling and reserved concurrency

3. Be awake before the rush — provisioned concurrency

For latency-critical functions, keep warm copies ready so there are no cold starts when the crowd arrives. Schedule it up before known events and down afterwards.

// Inside the Waiter construct, when lifestyle === "ratRace"
const live = placeOrder.addAlias("live", {
  provisionedConcurrentExecutions: 50, // warm copies ready before the sale
});
// route the API to the alias instead of the function
props.api.addRoutes({
  path: "/orders",
  methods: [apigw.HttpMethod.POST],
  integration: new HttpLambdaIntegration("PlaceOrderLive", live),
});

(Provisioned concurrency costs money while it's on — use Application Auto Scaling schedules to raise it only around peaks.)

4. Remember answers — caching

Most rat-race reads are the same: the product page, the menu, today's prices. Cache them at the edge (CloudFront), at the API, or in memory (ElastiCache, DAX). Every cached answer is one less task for a stressed human.

5. Stop panicking — backoff and circuit breakers

6. Make memory ready for the crowd

7. Rehearse

Run load tests at and beyond expected peak in a staging environment, and raise service quotas (Lambda concurrency, API Gateway limits) in advance. The rat race punishes surprises.

A rat-race checklist

Key idea

Rat-race systems survive by controlling the pace: queue the crowd, protect the fragile, stay warm before the rush, remember common answers, and never panic-retry.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I design and build serverless systems on AWS — and love explaining them simply.

Work with me