Appearance
Security, Privacy, and Compliance
One-line definition: model serving security is a protection system built around two assets — your models and your data — outward-facing defenses against attackers (model theft, prompt injection, adversarial examples), inward-facing defenses against leaks (user data, training data, internal policies), plus the regulatory obligations (GDPR, cross-border data transfer, licensing).
Industry insight: model serving is a new species with a built-in attack surface — an ordinary API is attacked through your code's vulnerabilities, while a model API is attacked through the model itself. An attacker doesn't need to breach your servers: repeatedly querying the inference endpoint is enough to reverse-engineer the model's capabilities. A single "ignore all previous instructions" fed to an LLM can slip past its guardrails. A carefully crafted adversarial example can make an image classifier see a "panda" as a "gibbon." For deployment engineers, security is the last step before launch — and the one most often skipped — because it shows no visible payoff until the day something goes wrong.
1. The Deployment Threat Surface
| Threat | Attack method | Consequence | Exposure |
|---|---|---|---|
| Model theft | Repeated querying, distillation attacks (training a clone on production outputs) | Loss of model IP | Public inference APIs |
| Prompt injection | Smuggling instructions into user input to steer the LLM | Guardrail bypass, information leaks | LLM APIs |
| Adversarial examples | Perturbations invisible to the human eye | Wrong outputs, moderation bypass | Any model API |
| Data poisoning | Contaminating training/fine-tuning data | Manipulated model behavior (backdoors) | Training pipelines |
| Supply chain attacks | Malicious weights, backdoored dependencies | Arbitrary code execution (pickle) | Model loading, dependencies |
Model theft is a real threat
Research has shown that 64,000 queries are enough to distill a 1-billion-parameter model (Knockoff Nets, 2018), and a 10B model can be stripped of its usable capability with a few tens of thousands of queries. For high-value models, rate limiting + input obfuscation + watermarking are the baseline defenses.
2. The API Security Baseline
1. Authentication and Authorization
| Scheme | Strength | Best for |
|---|---|---|
| API key | Weak (can leak) | Internal services, low-value assets |
| OAuth 2.0 / JWT | Medium (revocable, auditable) | Business-facing clients |
| mTLS | Strong (mutual certificates) | High-security internal networks |
The minimum bar: every online inference endpoint must require authentication, and anonymous access is forbidden. Keys must support per-caller quota isolation (paired with rate limiting).
2. Rate Limiting and Abuse Prevention
- Token-bucket rate limits per user/IP — they stop DDoS and model theft alike (an attacker needs huge query volumes);
- Request body size limits (so oversized payloads can't blow up serialization);
- Per-caller audit logs: who, when, which model, what status codes — audit logs are the only evidence you'll have when it's time to assign blame.
3. Input Validation and Output Filtering
- Input: schema validation (fields, types, lengths), content length caps, suspicious-pattern detection (e.g., injection-signature strings);
- Output (LLM): content moderation filters (block disallowed content), sensitive-information masking, confidence thresholds.
4. Error Message Discipline
Never leak framework versions, stack traces, or CUDA details to callers — return a uniform code + message + trace_id (see the error code design in Serving and Inference APIs).
3. Network Isolation and Deployment Boundaries
text
Public internet
│ (only the gateway is exposed)
▼
API gateway (WAF + rate limiting + authentication)
│ (private VPC network)
▼
Inference cluster (no public IP, reachable only from the internal network)
│
▼
Model storage / feature store (private, readable only by service accounts)- Inference services never face the public internet: business clients reach them through the gateway or the internal network;
- Model files and credentials live in a secrets manager (e.g., Vault/KMS) — never baked into images or committed to code repositories;
- Least privilege: service accounts can read models and write logs, and nothing more.
4. Privacy Protection
| Technique | Strength | Cost | Best for |
|---|---|---|---|
| PII detection and masking | Medium | Low | IDs/phone numbers/addresses in inputs and outputs |
| Local/edge inference | High (data never leaves the device) | High (deployment cost) | Strong privacy requirements |
| Federated learning | High | High | Collaborative training across distributed data |
| Differential privacy | Medium-high | Medium | Publishing statistics |
| Homomorphic encryption | Strong in theory | Extremely high (100–1000× slower, impractical) | Very few scenarios |
Don't fall for homomorphic encryption
Fully homomorphic encryption (FHE) promises "inference directly on ciphertext," but today's overhead is 2–3 orders of magnitude above plaintext inference — impractical in engineering terms. Production systems are better served by "data masking + local inference + encryption in transit."
Privacy engineering practices:
- Log tiering: request logs are masked by default (PII stripped), with a "full detail" audit mode available on demand;
- Data retention: an explicit retention period for inference logs (e.g., 30 days) with automatic purging at expiry;
- Encryption in transit and at rest: TLS end to end + sensitive feature columns encrypted at rest.
5. Compliance Essentials
1. Data Compliance
- GDPR / PIPL: processing personal data requires notice, consent, and deletability — the user's "right to be forgotten" must actually work, including deleting their inference data;
- Cross-border data transfer: the cross-border flow of training and inference data is regulated, and cross-border services require a formal assessment (under China's Measures for Security Assessment of Data Export);
- Regional compliance: data sovereignty rules (Europe requires data to stay in the EU; China requires critical data to be processed domestically) directly determine where you deploy in the cloud.
2. Model and Open-Source Compliance
- Weight license differences: the same architecture with different weights (e.g., Llama 2 community license vs. commercial license vs. non-commercial license) carries wildly different usage rights — "open-source code" does not mean "open-source weights";
- Open-source software compliance: component licenses (Apache-2.0, MIT, commercial terms) for things like ONNX Runtime and TensorRT need legal review;
- Training data compliance: the better the model, the more sensitive its training data sources — copyright and likeness rights issues keep growing.
The minimum compliance step: before launch, have legal/compliance walk through a data-flow diagram (whose data, passes through whom, stored where, retained how long) and put it in writing.
6. LLM-Specific Security
| Risk | Defense |
|---|---|
| Prompt injection (direct/indirect) | Input sanitization, clear instruction boundaries (system prompt delimiters), built-in LLM guardrails + external policy filters |
| Jailbreaks | Redundant multi-model detection, output moderation, adversarial red-team testing |
| Data leaks (sensitive data pasted into prompts) | Input PII detection, rules barring critical data from prompts, audit logs |
| Indirect injection (instructions smuggled in via web pages/documents) | Source isolation and labeling between "external content" and "user instructions" |
Field lesson: put your guardrails outside the model (input/output filters at the gateway layer), because the model's own guardrails can be bypassed. Also run red-team exercises regularly (attack your own defenses to measure how well they hold).
7. Security Checklist (Run Through Before Launch)
text
□ Authentication: no anonymous access; key/JWT/mTLS in place
□ Authorization: per-caller quotas, 429 on breach
□ Rate limiting: token bucket per user/IP, against scraping and theft
□ Input validation: schema, length, injection-signature checks
□ Output filtering: LLM content moderation, PII masking
□ Error discipline: no stack traces/versions/frameworks leaked
□ Network isolation: no public IP on services, VPC-internal
□ Secret management: KMS/Vault, never in code repositories
□ Audit logs: who called, when, with what result, stored masked
□ Data retention: explicit retention periods for logs/requests
□ Compliance review: data-flow diagram cleared by legal
□ Model provenance: weight hash verification, trusted supply chain
□ Monitoring and alerting: abnormal call patterns (high-frequency queries from one key = theft signal) — see /concepts/monitoringTrade-offs
| Decision | Options | How to choose |
|---|---|---|
| Auth strength | API key vs. OAuth vs. mTLS | Keys for internal low-value services; OAuth/mTLS for external high-value ones |
| Defense location | In-model guardrails vs. external filtering | External first; in-model guardrails are only the second line |
| Data residency | Cloud (convenient) vs. on-prem (private) | Go on-prem for strong privacy/compliance requirements |
| Log detail | Full audit vs. masked sampling | Masked sampling by default; full audit for critical systems |
| Privacy technology | Masking + local vs. homomorphic encryption | Stay away from homomorphic encryption; the cost-benefit doesn't work |
In one sentence: model serving security = know your threats (theft/injection/adversarial) + hold the baseline (auth, rate limiting, isolation) + manage your data (masking, compliance). Security earns you no KPIs, but a single leak can wipe out a year of optimization.
Further Reading
- Serving and Inference APIs — the foundations of API contracts, error discipline, and lifecycle
- Monitoring and Observability — spotting abnormal call patterns and security alerts
- MLOps Deployment Pipelines — process guarantees for supply chain and model provenance
- Common Pitfalls and Anti-Patterns — classic incidents caused by security oversights
- Model Formats and Conversion — SafeTensors and weight supply chain safety
References
- OWASP Top 10 for LLM Applications (the LLM app security checklist)
- Knockoff Nets: Stealing Functionality of Black-Box Models (the model theft paper, CVPR 2019)
- GDPR (the EU General Data Protection Regulation)
- Measures for Security Assessment of Data Export (Cyberspace Administration of China)
- Hugging Face model licensing guide (differences between model licenses)