Skip to content

Serverless Inference

At a glance Serverless is the best fit for low-frequency, bursty, pay-per-call inference. Deploy an inference service on AWS Lambda (or a cloud function), tackle cold starts head-on, and compare cost models against always-on serving.

Serverless Inference: Deploying on AWS Lambda and Beating Cold Starts ​

One-line definition: serverless inference runs your model on a pay-per-call, scale-to-zero platform such as AWS Lambda — there are no standing instances; the platform spins one up when a request arrives and tears it down afterward.

Why it's worth doing: if your inference traffic is low-frequency, bursty, or experimental, the money spent on a standing GPU instance is money burned — an A10 instance rents for thousands of yuan a month yet may only see a few dozen calls a day, a utilization rate under 1%. Serverless brings these scenarios down to a few cents per call and absorbs bursts automatically. This guide deploys an image classification service on AWS Lambda using a container image, focuses on defeating cold starts, and closes with the cost crossover point versus always-on serving.

1. When Serverless Inference Fits ​

Verdict first — four scenarios fit:

  1. Low-frequency workloads: a few hundred to a few thousand calls per day, with no steady traffic curve;
  2. Bursty traffic: flash sales, campaigns, batch jobs with peak-to-trough swings of 100× or more;
  3. Experimental models: models still in validation that may be discarded at any moment — not worth keeping an instance running;
  4. Event-driven: the model is triggered by message or object-storage events (tag images the moment they're uploaded).

Not a fit: stable high-QPS core paths (Lambda's concurrency cap plus its per-instance throughput trail always-on GPU serving; once QPS climbs, serverless gets both slower and more expensive), models that need GPUs (Lambda has none; see Section 5), and millisecond ultra-low latency (cold starts plus extra network hops).

The scenario framework maps to the "event-driven/serverless pattern" in deployment architecture patterns.

2. Architecture and Packaging: Lambda Container Images ​

text
API Gateway ──▶ AWS Lambda (container image) ──▶ Model inference ──▶ Return result

Lambda supports custom container images (up to 10GB, with the model and dependencies packed right in), a better fit for weight-carrying models than ZIP packages (250MB cap). The Dockerfile:

dockerfile
# Must use the AWS-provided Lambda base image
FROM public.ecr.aws/lambda/python:3.11
WORKDIR /var/task

COPY requirements.txt .
# Note: the model weights live in the image too (simplest when the model is small)
COPY models/resnet18.pth models/
COPY app/ app/

# The custom runtime entry point points at the handler
CMD ["app.handler.lambda_handler"]
python
# app/handler.py
import io
import torch
from PIL import Image
import torchvision.transforms as T
from torchvision.models import resnet18

# Global singleton: load once during cold start; warm requests on the same instance reuse it
_model = None

def _get_model():
    global _model
    if _model is None:
        model = resnet18(weights=None)
        model.load_state_dict(torch.load("/var/task/models/resnet18.pth", map_location="cpu"))
        model.eval()
        _model = model
    return _model

def lambda_handler(event, context):
    import base64
    body = event.get("body", "")
    img_bytes = base64.b64decode(body)

    transform = T.Compose([T.Resize(256), T.CenterCrop(224), T.ToTensor(),
                           T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])
    tensor = transform(Image.open(io.BytesIO(img_bytes)).convert("RGB")).unsqueeze(0)
    with torch.inference_mode():
        logits = _get_model()(tensor)
    top5 = torch.topk(torch.softmax(logits, dim=1)[0], k=5)
    return {"statusCode": 200,
            "body": [{"score": float(s), "index": int(i)} for s, i in zip(top5.values, top5.indices)]}

Compared with a FastAPI service: the same "module-level loading + model singleton" principle applies, but Lambda's concurrency model is "the platform scales the instance count," and each instance runs one concurrent request at a time (by default). So module-level loading pays off while the instance is being reused — you pay the loading cost once per cold start, and requests over the following tens of seconds all hit the warm path.

3. Cold Starts: The Number-One Enemy ​

Cold start time = container startup + runtime initialization + model loading. Reference numbers (ResNet18 on CPU):

PhaseTime
Container/runtime startup (platform side)300ms–2s
Python + PyTorch imports2–5s
Model weight loading0.5–2s
Total cold start~3–8s (warm requests < 100ms)

Four remedies, ordered from cheapest to most expensive:

  1. Warm-up (concurrency warm-up): a timer (e.g. an EventBridge rule firing every 2 minutes) sends a health-check request to keep instances warm — zero cost, but instances get reclaimed when real traffic doesn't arrive, so it only treats the symptom;
  2. Provisioned Concurrency: keep N instances ready ahead of time, removing cold starts from the path entirely — billed on the reserved amount, handing some of the savings back; fits core scenarios where stability is a must;
  3. Shrink the cold start itself: use Amazon Linux 2-style images, avoid exotic dependencies, lazy-load the model from S3 (don't bundle it; mount it with --mount-type), or switch to a faster runtime (e.g. move the CPU model to ONNX Runtime, or write the handler in Rust/Go);
  4. Raise the function timeout and memory: Lambda's timeout cap is 900s (15 minutes); the 3s default is nowhere near enough for model loading. Memory above 1024MB also raises the CPU quota (memory and CPU scale together).

Cold-Start Avalanche

A burst spins up dozens of instances at once, and every one of them pays its own 3–8s cold start — the first request's P99 can exceed 8 seconds. Mitigations: provisioned concurrency for a minimum floor, gateway-level degradation messages for cold-start requests, and going async on the business side (return a task_id first, deliver results via callback). Don't count on "cold starts disappearing once you load-test."

Event Triggers: Turn Request–Response into Event–Completion ​

Another common shape for low-frequency inference is event-driven — an object-storage upload or an arriving queue message triggers inference, with no need to wait for a synchronous reply. Lambda's S3/event triggers support this natively:

python
# An S3 image upload triggers tagging; results are written back to another bucket or a database
import boto3
from PIL import Image
import io

s3 = boto3.client("s3")

def lambda_handler(event, context):
    for record in event.get("Records", []):
        bucket = record["s3"]["bucket"]["name"]
        key = record["s3"]["object"]["key"]
        obj = s3.get_object(Bucket=bucket, Key=key)
        predictions = predict(obj["Body"].read())   # reuse the predict from Section 3
        s3.put_object(Bucket=bucket, Key=key.replace("uploads/", "tags/"),
                      Body=json.dumps(predictions))
    return {"statusCode": 200}

The benefit of event-driven inference: no gateway caller is sitting there waiting, so a slow first request no longer hurts the user experience — slow model loading doesn't matter. This is one of the most comfortable uses of serverless inference, and it matches the "event-driven pattern" in deployment architecture patterns.

Async Long Jobs: Return a task_id First, Deliver Results via Callback ​

When inference exceeds Lambda's synchronous timeout budget (29s at the gateway / 900s maximum for the function), switch to a two-step async flow: the request only enqueues the job (write to SQS/DB) and immediately returns a task_id; a Step Functions workflow (with retries and conditional branches if needed) then executes the inference, writes the results, and notifies the business side. Async costs about the same as sync, but it turns timeouts and retries into engineered, observable steps — the same pattern works as a lightweight front for batch inference pipelines.

4. The Cost Model: Lambda vs. Always-On GPU ​

Verdict first: below 1–5 QPS, Lambda is usually an order of magnitude cheaper than an always-on GPU instance; above a steady 5–10 QPS, always-on instances with autoscaling start to win. Let's do a concrete estimate:

Lambda billing ≈ invocations × (memory GB × duration seconds) × unit price (currently about $0.00001667/GB·s, plus $0.20 per million requests). An inference function with 1024MB of memory and a 2s average duration costs about 1GB × 2s × 0.00001667 ≈ $0.000033 per call — roughly ¥0.00023 per call.

OptionMonthly cost (10K calls/month)Monthly cost (3M calls/month)
AWS Lambda (1024MB, 2s average)~$0.5~$150
Always-on CPU instance (2 cores/4GB, running 24/7)~$30~$30 (unlimited runs)
Always-on GPU instance (A10, running 24/7)~$700+~$700+

Reframing in QPS terms makes it more intuitive: 10K calls/month ≈ 0.004 QPS on average; 3M calls/month ≈ 1.2 QPS on average. In other words, for a steady workload averaging under ~1 QPS, Lambda is still on par with an always-on CPU instance even at 3M calls/month; only sustained traffic above 1–5 QPS makes an always-on GPU genuinely economical. And don't cost just the compute — standing instances also carry maintenance, scaling, and the opportunity cost of idle GPUs; serverless bundles all of that into the unit price.

The numbers tell the story: at low frequency Lambda is 2–3 orders of magnitude cheaper; at high frequency an always-on instance amortizes to nearly free. The crossover sits at roughly "steady QPS ≈ 1–5". For the complete capacity and cost planning method, see performance optimization and capacity planning.

5. Limits and Workarounds ​

LimitValueWorkaround
Memory cap10240MB (10GB)Model + weights must fit in memory; compress/quantize large models
Timeout cap900s (15 minutes)Move long jobs to batch processing or Step Functions async orchestration
Image size cap10GBPull weights from S3 on demand, or split the model into its own service
ZIP package cap250MB (uncompressed)Use container images (≤10GB)
Concurrency capRegion-level quota (default 1000)Provisioned concurrency + request queue to smooth peaks
GPUUnavailableIf you need GPUs, switch to a container platform (below)
CPU quotaMore memory means more CPUPick a higher memory tier for compute-heavy inference

The lack of GPUs is the biggest boundary of serverless inference: Lambda has no GPU, so models beyond a few billion parameters or workloads that need real-time GPU inference must go to GPU-capable container platforms (SageMaker Serverless Inference on AWS, Google Cloud Run + GPU, Alibaba Cloud Function Compute GPU instances, or self-hosted K8s + KEDA scaling to zero with traffic). These platforms also bill per invocation/per scale-out, but they keep the GPU.

6. Platform Comparison in One Line Each ​

PlatformOne-line positioning
AWS LambdaRichest ecosystem and the most trigger sources, but no GPU and the strictest limits
Google Cloud RunContainers as a service, scales to zero, GPU support (beta), fast cold starts
Alibaba Cloud Function Compute (FC)Strong China ecosystem, GPU instances and custom runtimes
Tencent Cloud SCFIntegrates with the Tencent ecosystem (WeChat, COS triggers), lightweight and easy to pick up
Self-hosted K8s + KEDAFully controllable and GPU-capable, at the cost of operating scale-to-zero and cold starts yourself

Selection principle: first decide whether you need a GPU — if yes, choose among Cloud Run/FC/K8s; if no and you lean heavily on a cloud ecosystem, Lambda is the least friction. The platform decision framework is in how to choose frameworks and platforms.

Common Pitfalls and Troubleshooting ​

PitfallSymptomFix
Cold-start avalancheFirst requests take 5–10s during burstsProvisioned concurrency floor + going async; see Section 3
3s default timeoutTimeout fires before the model finishes loadingRaise Timeout (at least 30s for model-loading workloads)
Out of memoryOutOfMemoryError loading weightsMove up a memory tier; lazy-load weights from S3; quantize
Handler can't find the moduleUnable to import moduleMake the image WORKDIR/paths match the handler declaration
Image over 10GBDeployment rejectedPull weights from S3 instead of baking everything into the image
Concurrency over quotaTooManyRequestsExceptionProvisioned concurrency + SQS request queue to smooth peaks
GPU model on LambdaRuntime reports no CUDA deviceLambda has no GPU; switch to a GPU-capable container platform

Further Reading ​

References ​