Appearance
Serverless Inference: Deploying on AWS Lambda and Beating Cold Starts
One-line definition: serverless inference runs your model on a pay-per-call, scale-to-zero platform such as AWS Lambda — there are no standing instances; the platform spins one up when a request arrives and tears it down afterward.
Why it's worth doing: if your inference traffic is low-frequency, bursty, or experimental, the money spent on a standing GPU instance is money burned — an A10 instance rents for thousands of yuan a month yet may only see a few dozen calls a day, a utilization rate under 1%. Serverless brings these scenarios down to a few cents per call and absorbs bursts automatically. This guide deploys an image classification service on AWS Lambda using a container image, focuses on defeating cold starts, and closes with the cost crossover point versus always-on serving.
1. When Serverless Inference Fits
Verdict first — four scenarios fit:
- Low-frequency workloads: a few hundred to a few thousand calls per day, with no steady traffic curve;
- Bursty traffic: flash sales, campaigns, batch jobs with peak-to-trough swings of 100× or more;
- Experimental models: models still in validation that may be discarded at any moment — not worth keeping an instance running;
- Event-driven: the model is triggered by message or object-storage events (tag images the moment they're uploaded).
Not a fit: stable high-QPS core paths (Lambda's concurrency cap plus its per-instance throughput trail always-on GPU serving; once QPS climbs, serverless gets both slower and more expensive), models that need GPUs (Lambda has none; see Section 5), and millisecond ultra-low latency (cold starts plus extra network hops).
The scenario framework maps to the "event-driven/serverless pattern" in deployment architecture patterns.
2. Architecture and Packaging: Lambda Container Images
text
API Gateway ──▶ AWS Lambda (container image) ──▶ Model inference ──▶ Return resultLambda supports custom container images (up to 10GB, with the model and dependencies packed right in), a better fit for weight-carrying models than ZIP packages (250MB cap). The Dockerfile:
dockerfile
# Must use the AWS-provided Lambda base image
FROM public.ecr.aws/lambda/python:3.11
WORKDIR /var/task
COPY requirements.txt .
# Note: the model weights live in the image too (simplest when the model is small)
COPY models/resnet18.pth models/
COPY app/ app/
# The custom runtime entry point points at the handler
CMD ["app.handler.lambda_handler"]python
# app/handler.py
import io
import torch
from PIL import Image
import torchvision.transforms as T
from torchvision.models import resnet18
# Global singleton: load once during cold start; warm requests on the same instance reuse it
_model = None
def _get_model():
global _model
if _model is None:
model = resnet18(weights=None)
model.load_state_dict(torch.load("/var/task/models/resnet18.pth", map_location="cpu"))
model.eval()
_model = model
return _model
def lambda_handler(event, context):
import base64
body = event.get("body", "")
img_bytes = base64.b64decode(body)
transform = T.Compose([T.Resize(256), T.CenterCrop(224), T.ToTensor(),
T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])
tensor = transform(Image.open(io.BytesIO(img_bytes)).convert("RGB")).unsqueeze(0)
with torch.inference_mode():
logits = _get_model()(tensor)
top5 = torch.topk(torch.softmax(logits, dim=1)[0], k=5)
return {"statusCode": 200,
"body": [{"score": float(s), "index": int(i)} for s, i in zip(top5.values, top5.indices)]}Compared with a FastAPI service: the same "module-level loading + model singleton" principle applies, but Lambda's concurrency model is "the platform scales the instance count," and each instance runs one concurrent request at a time (by default). So module-level loading pays off while the instance is being reused — you pay the loading cost once per cold start, and requests over the following tens of seconds all hit the warm path.
3. Cold Starts: The Number-One Enemy
Cold start time = container startup + runtime initialization + model loading. Reference numbers (ResNet18 on CPU):
| Phase | Time |
|---|---|
| Container/runtime startup (platform side) | 300ms–2s |
| Python + PyTorch imports | 2–5s |
| Model weight loading | 0.5–2s |
| Total cold start | ~3–8s (warm requests < 100ms) |
Four remedies, ordered from cheapest to most expensive:
- Warm-up (concurrency warm-up): a timer (e.g. an EventBridge rule firing every 2 minutes) sends a health-check request to keep instances warm — zero cost, but instances get reclaimed when real traffic doesn't arrive, so it only treats the symptom;
- Provisioned Concurrency: keep N instances ready ahead of time, removing cold starts from the path entirely — billed on the reserved amount, handing some of the savings back; fits core scenarios where stability is a must;
- Shrink the cold start itself: use Amazon Linux 2-style images, avoid exotic dependencies, lazy-load the model from S3 (don't bundle it; mount it with
--mount-type), or switch to a faster runtime (e.g. move the CPU model to ONNX Runtime, or write the handler in Rust/Go); - Raise the function timeout and memory: Lambda's timeout cap is 900s (15 minutes); the 3s default is nowhere near enough for model loading. Memory above 1024MB also raises the CPU quota (memory and CPU scale together).
Cold-Start Avalanche
A burst spins up dozens of instances at once, and every one of them pays its own 3–8s cold start — the first request's P99 can exceed 8 seconds. Mitigations: provisioned concurrency for a minimum floor, gateway-level degradation messages for cold-start requests, and going async on the business side (return a task_id first, deliver results via callback). Don't count on "cold starts disappearing once you load-test."
Event Triggers: Turn Request–Response into Event–Completion
Another common shape for low-frequency inference is event-driven — an object-storage upload or an arriving queue message triggers inference, with no need to wait for a synchronous reply. Lambda's S3/event triggers support this natively:
python
# An S3 image upload triggers tagging; results are written back to another bucket or a database
import boto3
from PIL import Image
import io
s3 = boto3.client("s3")
def lambda_handler(event, context):
for record in event.get("Records", []):
bucket = record["s3"]["bucket"]["name"]
key = record["s3"]["object"]["key"]
obj = s3.get_object(Bucket=bucket, Key=key)
predictions = predict(obj["Body"].read()) # reuse the predict from Section 3
s3.put_object(Bucket=bucket, Key=key.replace("uploads/", "tags/"),
Body=json.dumps(predictions))
return {"statusCode": 200}The benefit of event-driven inference: no gateway caller is sitting there waiting, so a slow first request no longer hurts the user experience — slow model loading doesn't matter. This is one of the most comfortable uses of serverless inference, and it matches the "event-driven pattern" in deployment architecture patterns.
Async Long Jobs: Return a task_id First, Deliver Results via Callback
When inference exceeds Lambda's synchronous timeout budget (29s at the gateway / 900s maximum for the function), switch to a two-step async flow: the request only enqueues the job (write to SQS/DB) and immediately returns a task_id; a Step Functions workflow (with retries and conditional branches if needed) then executes the inference, writes the results, and notifies the business side. Async costs about the same as sync, but it turns timeouts and retries into engineered, observable steps — the same pattern works as a lightweight front for batch inference pipelines.
4. The Cost Model: Lambda vs. Always-On GPU
Verdict first: below 1–5 QPS, Lambda is usually an order of magnitude cheaper than an always-on GPU instance; above a steady 5–10 QPS, always-on instances with autoscaling start to win. Let's do a concrete estimate:
Lambda billing ≈ invocations × (memory GB × duration seconds) × unit price (currently about $0.00001667/GB·s, plus $0.20 per million requests). An inference function with 1024MB of memory and a 2s average duration costs about 1GB × 2s × 0.00001667 ≈ $0.000033 per call — roughly ¥0.00023 per call.
| Option | Monthly cost (10K calls/month) | Monthly cost (3M calls/month) |
|---|---|---|
| AWS Lambda (1024MB, 2s average) | ~$0.5 | ~$150 |
| Always-on CPU instance (2 cores/4GB, running 24/7) | ~$30 | ~$30 (unlimited runs) |
| Always-on GPU instance (A10, running 24/7) | ~$700+ | ~$700+ |
Reframing in QPS terms makes it more intuitive: 10K calls/month ≈ 0.004 QPS on average; 3M calls/month ≈ 1.2 QPS on average. In other words, for a steady workload averaging under ~1 QPS, Lambda is still on par with an always-on CPU instance even at 3M calls/month; only sustained traffic above 1–5 QPS makes an always-on GPU genuinely economical. And don't cost just the compute — standing instances also carry maintenance, scaling, and the opportunity cost of idle GPUs; serverless bundles all of that into the unit price.
The numbers tell the story: at low frequency Lambda is 2–3 orders of magnitude cheaper; at high frequency an always-on instance amortizes to nearly free. The crossover sits at roughly "steady QPS ≈ 1–5". For the complete capacity and cost planning method, see performance optimization and capacity planning.
5. Limits and Workarounds
| Limit | Value | Workaround |
|---|---|---|
| Memory cap | 10240MB (10GB) | Model + weights must fit in memory; compress/quantize large models |
| Timeout cap | 900s (15 minutes) | Move long jobs to batch processing or Step Functions async orchestration |
| Image size cap | 10GB | Pull weights from S3 on demand, or split the model into its own service |
| ZIP package cap | 250MB (uncompressed) | Use container images (≤10GB) |
| Concurrency cap | Region-level quota (default 1000) | Provisioned concurrency + request queue to smooth peaks |
| GPU | Unavailable | If you need GPUs, switch to a container platform (below) |
| CPU quota | More memory means more CPU | Pick a higher memory tier for compute-heavy inference |
The lack of GPUs is the biggest boundary of serverless inference: Lambda has no GPU, so models beyond a few billion parameters or workloads that need real-time GPU inference must go to GPU-capable container platforms (SageMaker Serverless Inference on AWS, Google Cloud Run + GPU, Alibaba Cloud Function Compute GPU instances, or self-hosted K8s + KEDA scaling to zero with traffic). These platforms also bill per invocation/per scale-out, but they keep the GPU.
6. Platform Comparison in One Line Each
| Platform | One-line positioning |
|---|---|
| AWS Lambda | Richest ecosystem and the most trigger sources, but no GPU and the strictest limits |
| Google Cloud Run | Containers as a service, scales to zero, GPU support (beta), fast cold starts |
| Alibaba Cloud Function Compute (FC) | Strong China ecosystem, GPU instances and custom runtimes |
| Tencent Cloud SCF | Integrates with the Tencent ecosystem (WeChat, COS triggers), lightweight and easy to pick up |
| Self-hosted K8s + KEDA | Fully controllable and GPU-capable, at the cost of operating scale-to-zero and cold starts yourself |
Selection principle: first decide whether you need a GPU — if yes, choose among Cloud Run/FC/K8s; if no and you lean heavily on a cloud ecosystem, Lambda is the least friction. The platform decision framework is in how to choose frameworks and platforms.
Common Pitfalls and Troubleshooting
| Pitfall | Symptom | Fix |
|---|---|---|
| Cold-start avalanche | First requests take 5–10s during bursts | Provisioned concurrency floor + going async; see Section 3 |
| 3s default timeout | Timeout fires before the model finishes loading | Raise Timeout (at least 30s for model-loading workloads) |
| Out of memory | OutOfMemoryError loading weights | Move up a memory tier; lazy-load weights from S3; quantize |
| Handler can't find the module | Unable to import module | Make the image WORKDIR/paths match the handler declaration |
| Image over 10GB | Deployment rejected | Pull weights from S3 instead of baking everything into the image |
| Concurrency over quota | TooManyRequestsException | Provisioned concurrency + SQS request queue to smooth peaks |
| GPU model on Lambda | Runtime reports no CUDA device | Lambda has no GPU; switch to a GPU-capable container platform |
Further Reading
- Deployment architecture patterns — where serverless/event-driven patterns sit in the deployment architecture spectrum
- Performance optimization and capacity planning — the full method for cost models, QPS crossovers, and capacity estimation
- FastAPI + Docker online serving — the always-on counterpart, the other end of the cost comparison
- Serving and inference APIs — a serverless handler is still "serving"; the interface design principles carry over
- The MLOps deployment pipeline — artifact management for baking models into images or pulling them from S3
- Batch inference pipelines — often a better answer for low-frequency, high-volume workloads: run in bulk, then release
References
- AWS Lambda documentation: https://docs.aws.amazon.com/lambda/latest/dg/welcome.html
- AWS Lambda container images: https://docs.aws.amazon.com/lambda/latest/dg/images-create.html
- Google Cloud Run documentation: https://cloud.google.com/run/docs
- Alibaba Cloud Function Compute documentation: https://help.aliyun.com/zh/fc/
- KEDA (K8s autoscaling to zero): https://keda.sh/docs/