Streaming and data ingestion
Task statement 3.5 covers getting data into AWS fast and at scale — clickstreams, logs, IoT telemetry, transactions. The key skill is matching the ingestion service to the required latency, ordering, replay and destination.
Batch vs streaming
| Batch | Streaming | |
|---|---|---|
| Frequency | Hourly/daily files | Continuous events |
| Latency | Minutes to hours | Milliseconds to seconds |
| Tools | DataSync, S3 uploads, Glue jobs, Transfer Family | Kinesis, MSK, Firehose |
Amazon Kinesis Data Streams
A real-time, durable stream that many consumers can read independently.
- Data is split into shards (provisioned mode: each shard ingests up to about 1 MB/s or 1,000 records/s and serves about 2 MB/s of reads) — or use on-demand mode to scale automatically.
- Records are ordered per partition key within a shard.
- Retention from 24 hours (default) up to 365 days — consumers can replay data.
- Consumers: Lambda, Managed Service for Apache Flink, Firehose, custom apps (KCL), with enhanced fan-out for dedicated throughput per consumer.
Cue: "real-time", "multiple applications process the same stream", "replay", "ordering per key".
Amazon Data Firehose (formerly Kinesis Data Firehose)
The simplest way to load streaming data into destinations:
- Destinations: S3, Redshift (via S3), OpenSearch, Splunk, HTTP endpoints and partner tools.
- Fully managed, scales automatically, no consumers to write.
- Buffers by size/time → near real-time (seconds to minutes), not instant.
- Can transform records with Lambda, convert formats (e.g. JSON → Parquet/ORC), compress and encrypt.
"Real-time processing, custom consumers, replay" → Kinesis Data Streams. "Deliver streaming data into S3/Redshift/OpenSearch with the least effort, near real time" → Data Firehose.
Amazon Managed Service for Apache Flink (formerly Kinesis Data Analytics)
Run Apache Flink applications (Java, Python, SQL) that process streams in real time — windowed aggregations, anomaly detection, joins — reading from Kinesis or MSK.
Amazon MSK (Managed Streaming for Apache Kafka)
Fully managed Apache Kafka clusters (and MSK Serverless). Choose when the company already uses Kafka, needs Kafka's ecosystem/APIs, or very large message sizes and long retention. MSK Connect runs Kafka Connect connectors.
Amazon Kinesis Video Streams
Ingest and store video (and other time-encoded data) from cameras and devices for playback, analytics and ML (e.g. with Rekognition Video).
Typical streaming pipeline
Producers (apps, IoT, logs)
│
▼
Kinesis Data Streams ──► Managed Flink (real-time aggregates) ──► DynamoDB / dashboards
│
└─► Data Firehose ──► S3 (Parquet, partitioned) ──► Glue Data Catalog ──► Athena / Redshift / QuickSight
Securing ingestion
- IAM policies on producers and consumers; VPC endpoints for private access.
- Server-side encryption with KMS; TLS in transit.
- For external clients: API Gateway (with auth) in front of Kinesis, or Cognito-issued credentials.
Sizing (task 3.5 "sizes and speeds")
- Estimate peak records/second and average record size → shards needed (provisioned) or choose on-demand.
- Firehose buffer settings trade delivery latency against file size (bigger files are better for analytics).
Exam patterns
- "Collect clickstream data and analyse it in real time with multiple consumers" → Kinesis Data Streams.
- "Load log data into S3 and convert it to Parquet without managing servers" → Data Firehose with format conversion.
- "Existing on-prem Kafka; move to a managed service" → MSK.
- "Detect anomalies in streaming data within seconds" → Managed Service for Apache Flink.