Heartbeat and health
Doctors understand a body through its pulse, temperature, history and symptoms. You understand a running microservice the same way — through observability: logs, metrics, traces and alarms. A service without them is like a patient who can't speak.
The medical kit
| Medical idea | Observability | AWS |
|---|---|---|
| Diary — what happened, in words | Logs | CloudWatch Logs |
| Pulse, temperature — numbers over time | Metrics | CloudWatch Metrics |
| Following one thought through the body | Traces | AWS X-Ray (or OpenTelemetry) |
| Pain — something is wrong, act now | Alarms | CloudWatch Alarms → SNS/Slack/on-call |
| Medical chart on the wall | Dashboards | CloudWatch Dashboards |
| Regular check-up | Synthetic tests | CloudWatch Synthetics canaries |
Writing a useful diary: structured logs
Write logs as JSON with consistent fields so you can search them:
import { Logger } from "@aws-lambda-powertools/logger";
const logger = new Logger({ serviceName: "waiter" });
logger.info("Order placed", { orderId: order.orderId, tableNumber: order.tableNumber, total: order.total });
Include a correlation ID (for example the order ID or request ID) in every log line, and pass it along in events — so you can follow one customer's journey across many humans.
Taking the pulse: metrics
Three kinds of numbers matter:
- Technical — invocations, errors, duration, throttles, queue depth, DynamoDB throttled requests.
- Business — orders placed per minute, payments failed, average order value.
- Experience — latency percentiles (p50, p95, p99), not just averages.
import { Metrics, MetricUnit } from "@aws-lambda-powertools/metrics";
const metrics = new Metrics({ namespace: "CafeSukoon", serviceName: "waiter" });
metrics.addMetric("OrdersPlaced", MetricUnit.Count, 1);
metrics.publishStoredMetrics();
(Powertools for AWS Lambda — available for TypeScript, Python, Java and .NET — makes logging, metrics, tracing and idempotency easy and consistent.)
Following a thought: tracing
A trace shows one request's path — API Gateway → Waiter Lambda → DynamoDB → EventBridge — with the time spent in each step. Enable X-Ray tracing on API Gateway and Lambda (tracing: lambda.Tracing.ACTIVE in CDK) to find which organ is slow.
Feeling pain: alarms that matter
Alarm on symptoms customers feel, not on every twitch:
| Alarm | Why |
|---|---|
| API 5xx error rate above a threshold | Customers are failing |
| p99 latency above target | Customers are waiting |
| Messages in a dead-letter queue > 0 | Work is being lost |
| Queue age growing | Workers can't keep up |
| Business metric drops (no orders in 15 minutes at lunch) | Something silent is broken |
Too many alarms is like chronic pain — people stop noticing. Every alarm should require a human action; otherwise make it a dashboard metric.
Health checks
For serverless services there's no server to ping, but you can still:
- Expose a lightweight
/healthroute that checks critical dependencies (memory reachable, configuration loaded). - Run synthetic canaries that place a test order every few minutes and alarm if it fails.
The doctor's dashboard
One dashboard per human: traffic, errors, latency, queue depth, dead letters, and its key business metric. One dashboard for the whole society: the journey of an order across all humans.
If you can't answer "is this human healthy right now, and if not, which organ is sick?" within a minute, the service isn't production-ready yet.