Skip to content

Security, Privacy, and Compliance

At a glance Once a model goes live it is exposed to the public internet, and the security risks follow: model theft, prompt injection, adversarial attacks, data leaks. This article sets out a security baseline for inference services: authentication and authorization, input defenses, privacy protection, and compliance essentials.

Security, Privacy, and Compliance ​

One-line definition: model serving security is a protection system built around two assets — your models and your data — outward-facing defenses against attackers (model theft, prompt injection, adversarial examples), inward-facing defenses against leaks (user data, training data, internal policies), plus the regulatory obligations (GDPR, cross-border data transfer, licensing).

Industry insight: model serving is a new species with a built-in attack surface — an ordinary API is attacked through your code's vulnerabilities, while a model API is attacked through the model itself. An attacker doesn't need to breach your servers: repeatedly querying the inference endpoint is enough to reverse-engineer the model's capabilities. A single "ignore all previous instructions" fed to an LLM can slip past its guardrails. A carefully crafted adversarial example can make an image classifier see a "panda" as a "gibbon." For deployment engineers, security is the last step before launch — and the one most often skipped — because it shows no visible payoff until the day something goes wrong.

1. The Deployment Threat Surface ​

ThreatAttack methodConsequenceExposure
Model theftRepeated querying, distillation attacks (training a clone on production outputs)Loss of model IPPublic inference APIs
Prompt injectionSmuggling instructions into user input to steer the LLMGuardrail bypass, information leaksLLM APIs
Adversarial examplesPerturbations invisible to the human eyeWrong outputs, moderation bypassAny model API
Data poisoningContaminating training/fine-tuning dataManipulated model behavior (backdoors)Training pipelines
Supply chain attacksMalicious weights, backdoored dependenciesArbitrary code execution (pickle)Model loading, dependencies

Model theft is a real threat

Research has shown that 64,000 queries are enough to distill a 1-billion-parameter model (Knockoff Nets, 2018), and a 10B model can be stripped of its usable capability with a few tens of thousands of queries. For high-value models, rate limiting + input obfuscation + watermarking are the baseline defenses.

2. The API Security Baseline ​

1. Authentication and Authorization ​

SchemeStrengthBest for
API keyWeak (can leak)Internal services, low-value assets
OAuth 2.0 / JWTMedium (revocable, auditable)Business-facing clients
mTLSStrong (mutual certificates)High-security internal networks

The minimum bar: every online inference endpoint must require authentication, and anonymous access is forbidden. Keys must support per-caller quota isolation (paired with rate limiting).

2. Rate Limiting and Abuse Prevention ​

  • Token-bucket rate limits per user/IP — they stop DDoS and model theft alike (an attacker needs huge query volumes);
  • Request body size limits (so oversized payloads can't blow up serialization);
  • Per-caller audit logs: who, when, which model, what status codes — audit logs are the only evidence you'll have when it's time to assign blame.

3. Input Validation and Output Filtering ​

  • Input: schema validation (fields, types, lengths), content length caps, suspicious-pattern detection (e.g., injection-signature strings);
  • Output (LLM): content moderation filters (block disallowed content), sensitive-information masking, confidence thresholds.

4. Error Message Discipline ​

Never leak framework versions, stack traces, or CUDA details to callers — return a uniform code + message + trace_id (see the error code design in Serving and Inference APIs).

3. Network Isolation and Deployment Boundaries ​

text
Public internet
  │ (only the gateway is exposed)
  ▼
API gateway (WAF + rate limiting + authentication)
  │ (private VPC network)
  ▼
Inference cluster (no public IP, reachable only from the internal network)
  │
  ▼
Model storage / feature store (private, readable only by service accounts)
  • Inference services never face the public internet: business clients reach them through the gateway or the internal network;
  • Model files and credentials live in a secrets manager (e.g., Vault/KMS) — never baked into images or committed to code repositories;
  • Least privilege: service accounts can read models and write logs, and nothing more.

4. Privacy Protection ​

TechniqueStrengthCostBest for
PII detection and maskingMediumLowIDs/phone numbers/addresses in inputs and outputs
Local/edge inferenceHigh (data never leaves the device)High (deployment cost)Strong privacy requirements
Federated learningHighHighCollaborative training across distributed data
Differential privacyMedium-highMediumPublishing statistics
Homomorphic encryptionStrong in theoryExtremely high (100–1000× slower, impractical)Very few scenarios

Don't fall for homomorphic encryption

Fully homomorphic encryption (FHE) promises "inference directly on ciphertext," but today's overhead is 2–3 orders of magnitude above plaintext inference — impractical in engineering terms. Production systems are better served by "data masking + local inference + encryption in transit."

Privacy engineering practices:

  1. Log tiering: request logs are masked by default (PII stripped), with a "full detail" audit mode available on demand;
  2. Data retention: an explicit retention period for inference logs (e.g., 30 days) with automatic purging at expiry;
  3. Encryption in transit and at rest: TLS end to end + sensitive feature columns encrypted at rest.

5. Compliance Essentials ​

1. Data Compliance ​

  • GDPR / PIPL: processing personal data requires notice, consent, and deletability — the user's "right to be forgotten" must actually work, including deleting their inference data;
  • Cross-border data transfer: the cross-border flow of training and inference data is regulated, and cross-border services require a formal assessment (under China's Measures for Security Assessment of Data Export);
  • Regional compliance: data sovereignty rules (Europe requires data to stay in the EU; China requires critical data to be processed domestically) directly determine where you deploy in the cloud.

2. Model and Open-Source Compliance ​

  • Weight license differences: the same architecture with different weights (e.g., Llama 2 community license vs. commercial license vs. non-commercial license) carries wildly different usage rights — "open-source code" does not mean "open-source weights";
  • Open-source software compliance: component licenses (Apache-2.0, MIT, commercial terms) for things like ONNX Runtime and TensorRT need legal review;
  • Training data compliance: the better the model, the more sensitive its training data sources — copyright and likeness rights issues keep growing.

The minimum compliance step: before launch, have legal/compliance walk through a data-flow diagram (whose data, passes through whom, stored where, retained how long) and put it in writing.

6. LLM-Specific Security ​

RiskDefense
Prompt injection (direct/indirect)Input sanitization, clear instruction boundaries (system prompt delimiters), built-in LLM guardrails + external policy filters
JailbreaksRedundant multi-model detection, output moderation, adversarial red-team testing
Data leaks (sensitive data pasted into prompts)Input PII detection, rules barring critical data from prompts, audit logs
Indirect injection (instructions smuggled in via web pages/documents)Source isolation and labeling between "external content" and "user instructions"

Field lesson: put your guardrails outside the model (input/output filters at the gateway layer), because the model's own guardrails can be bypassed. Also run red-team exercises regularly (attack your own defenses to measure how well they hold).

7. Security Checklist (Run Through Before Launch) ​

text
□ Authentication: no anonymous access; key/JWT/mTLS in place
□ Authorization: per-caller quotas, 429 on breach
□ Rate limiting: token bucket per user/IP, against scraping and theft
□ Input validation: schema, length, injection-signature checks
□ Output filtering: LLM content moderation, PII masking
□ Error discipline: no stack traces/versions/frameworks leaked
□ Network isolation: no public IP on services, VPC-internal
□ Secret management: KMS/Vault, never in code repositories
□ Audit logs: who called, when, with what result, stored masked
□ Data retention: explicit retention periods for logs/requests
□ Compliance review: data-flow diagram cleared by legal
□ Model provenance: weight hash verification, trusted supply chain
□ Monitoring and alerting: abnormal call patterns (high-frequency queries from one key = theft signal) — see /concepts/monitoring

Trade-offs ​

DecisionOptionsHow to choose
Auth strengthAPI key vs. OAuth vs. mTLSKeys for internal low-value services; OAuth/mTLS for external high-value ones
Defense locationIn-model guardrails vs. external filteringExternal first; in-model guardrails are only the second line
Data residencyCloud (convenient) vs. on-prem (private)Go on-prem for strong privacy/compliance requirements
Log detailFull audit vs. masked samplingMasked sampling by default; full audit for critical systems
Privacy technologyMasking + local vs. homomorphic encryptionStay away from homomorphic encryption; the cost-benefit doesn't work

In one sentence: model serving security = know your threats (theft/injection/adversarial) + hold the baseline (auth, rate limiting, isolation) + manage your data (masking, compliance). Security earns you no KPIs, but a single leak can wipe out a year of optimization.

Further Reading ​

References ​