Skip to content

AI Safety and Governance

At a glance AI safety and governance is about identifying, measuring, and mitigating the risks of AI systems so that AI stays controllable, trustworthy, and accountable — this article covers a four-layer risk taxonomy, six technical countermeasures, the global regulatory landscape, and a practical checklist for engineers.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

AI Safety and Governance ​

Defining the Concept: From "How Strong Is the Model?" to "What Happens When Things Go Wrong?" ​

AI safety and governance study how to identify, measure, and mitigate the risks posed by AI systems, and establish the rules that keep AI controllable, trustworthy, and accountable.

Over the past few years, the questions the industry asks about AI have escalated twice: first from "what can models do" to "how well do they do it" — which gave us evaluation and benchmarks; then from "how well do they do it" to "what happens when they get it wrong, and who is responsible" — which is exactly the question AI safety and governance exist to answer. If LLM Evaluation and Benchmarks measures capability, then safety and governance measure risk and responsibility. To see the full landscape of hot AI concepts, start with What Are the Hot AI Concepts for the capability map, then come back to this page for the risk map.

Safety leans technical: red-teaming, jailbreak defenses, alignment training, watermarking and provenance. Governance leans rules: laws, standards, filing and registration, audits, accountability. The two are two sides of the same coin — technical countermeasures answer "how do we prevent this," while governance rules answer "who should prevent it, to what standard, and who bears responsibility when prevention fails." Model vendors, application developers, regulators, corporate compliance teams, and end users are all participants in this system.

The bottom line

Safety is the technical foundation of governance; governance is the institutional safeguard for safety. Technology without regulation leaves engineering running naked; regulation without technology leaves oversight hanging in the air.

1. Why AI Safety Has Become a Must-Answer Question ​

Four real-world forces have turned "safety and governance" from an academic topic into an enterprise-level engineering topic:

DriverWhat It Looks Like in Practice
Rapidly rising capabilityLarge models are moving from conversation to agents, multimodality, and autonomous execution, and the blast radius of an error has grown from "one sentence" to "one wire transfer" (see AI Agents)
Deployment at scaleGenerative AI has entered high-stakes domains such as customer service, finance, healthcare, and law, upgrading errors from "jokes" to "incidents"
Adversarial exploitationJailbreaks, prompt injection, and deepfake scams have already produced real victims (see Section 5)
Regulation taking effectThe EU AI Act is phasing in, China's filing regime has become routine, and standards are landing worldwide — compliance is now a hard constraint

Don't treat "safety" as the last step before launch

The cost curve of safety is cheaper the earlier you start: fixing a problem caught at the requirements and architecture stage costs a tenth — or less — of fixing one caught after launch. That is where the "Security by Design" philosophy comes from, expanded in Section 6 below.

2. A Risk Taxonomy: The Four-Layer Risk Map ​

Risk isn't flat. The industry commonly uses a layered approach to break it into four layers — content, system, societal, and existential. Deciding which layer a given risk belongs to determines whether you should respond with technology, product measures, or policy.

LayerExample RisksPrimarily AffectsPrimary MitigationsRelated Page
Content layerHallucination, bias and discrimination, harmful contentEnd users, corporate reputationEvaluation, filtering, value alignment, fact-checkingLLM Evaluation and Benchmarks
System layerPrompt injection, jailbreaks, data poisoning, model extractionApplication developers, enterprise assetsInput/output filtering, classifiers, red-teaming, sandbox isolationPrompt Engineering
Societal layerDeepfakes, misinformation, job displacement, privacy erosionThe general public, public orderWatermarking and provenance, platform governance, laws and regulations, industry self-regulationDiffusion Models and Generative AI
Existential layerLoss of control, misalignment, long-term AGI risksAll of humanityAlignment research, interpretability, international coordinationAlignment: RLHF and DPO

Content Layer: When the Model "Gets It Wrong, or Says It Badly" ​

  • Hallucination: the model confidently fabricates nonexistent facts, sources, and citations;
  • Bias and discrimination: biases in the training data get amplified into unfair outputs — stereotypes about gender, race, and region;
  • Harmful content: hate speech, incitement to violence, self-harm guidance, illegal information.

What content-layer risks share is this: they are problems of "output quality + values," and measuring them depends on an evaluation system — without evaluation there is no baseline, and without a baseline you cannot measure whether mitigation works. That is why hallucination rates, bias metrics, and safety benchmarks all fall under LLM Evaluation and Benchmarks.

System Layer: When People "Exploit the Loopholes" ​

  • Prompt injection: an attacker smuggles malicious instructions into text or into data returned by tools, making the model perform unintended actions;
  • Jailbreak: carefully crafted prompts bypass the model's safety guardrails (role-play, encoding workarounds, multi-turn persuasion...) to make it output prohibited content;
  • Data poisoning: contaminating training or fine-tuning data so the model misbehaves or turns malicious under specific trigger conditions;
  • Model extraction: reverse-engineering a model's weights or capabilities through massive querying to get around commercial protections.

System-layer risks share the same underlying mechanism as Prompt Engineering — "how input shapes output." Prompt engineers study how to make inputs better; attackers study how to make inputs worse. They are two faces of the same coin.

Prompt injection is the #1 system-layer risk of the agent era

Once models start calling tools, touching databases, and executing code (see AI Agents), prompt injection escalates from "non-compliant output" to "non-compliant action" — a single injection could trigger a wire transfer, wipe a database, or leak secrets. This is exactly why the "least privilege" principle in Section 7 exists.

Societal Layer: When Generative AI Collides with the Real World ​

  • Deepfakes: face swaps and fabricated audio and video, used for fraud, blackmail, and forged evidence — the core technology is precisely Diffusion Models and Generative AI;
  • Misinformation: mass-producing plausible fake news and fake reviews at near-zero cost, amplifying manipulation of public opinion;
  • Job displacement: role transitions and social redistribution driven by automation;
  • Privacy erosion: personal information in training data, leaked conversation logs, and inferential privacy (deriving sensitive facts from seemingly unrelated data).

No single company can solve societal-layer risks alone. It takes a combination of "technical watermarking + platform rules + laws and regulations + public literacy."

Existential Layer: Long-Term and Extreme Risks ​

  • Loss of control: a future superintelligence behaving beyond the intent and understanding of its developers;
  • Misalignment: the model's optimization objective diverging from humanity's true intentions (see Alignment: RLHF and DPO);
  • Dual-use misuse: capabilities abused for large-scale harm in domains such as biotech or cyber.

The debate over existential-layer risk

Academia is deeply divided on existential-layer risks: some call them "the most important safety issue of this century," others "a sci-fi distraction." But the mainstream consensus is that they are at least worth researching, and that they call for governance dialogue at the international level. Frontline engineers need not take sides, but should understand that alignment as a technology is the common foundation beneath nearly every risk discussion.

3. The Technical Countermeasure Toolbox: Six Approaches ​

1. Red-Teaming ​

Red-teaming means proactively probing models and systems for weaknesses from an attacker's point of view — simulating jailbreaks, injections, and harmful inputs, then fixing what is found, re-testing, and iterating. It is the de facto standard for safety evaluation.

StepWhat It Involves
Define a threat modelClarify "who would attack, what they would target, and what they aim to achieve"
Build attack samplesJailbreak templates, injection payloads, edge-case inputs, multilingual and encoding bypasses
Execute at scale and measureAttack success rate (ASR), violation rate, refusal rate
Fix and regressAdd defenses against the failed samples, run regression checks, and curate the results into a "safety evaluation set"

The bottom line

The value of red-teaming lies not in "finding every vulnerability" but in building a sustainable, regression-tested line of defense — turning every newly discovered attack into an evaluation case so the next model release is "at least no worse than the last."

2. Jailbreak Defenses: Input/Output Filtering and Classifiers ​

Two gates against jailbreaks and injection:

  • Input side: system prompt hardening (explicitly instructing the model to ignore any instruction telling it to take on a different persona), input classifiers (spotting attack patterns and blocking them), and injection detection (tagging tool-returned text as untrusted or parsing it separately, never concatenating it straight into the prompt);
  • Output side: output classifiers (safety classification models such as OpenAI's Moderation API and Llama Guard), sensitive-content blocking, and length and format constraints.
text
User input ──► Input filter/classifier ──► LLM (system-prompt guardrails) ──► Output classifier ──► User / downstream systems
                 │                        │                      │
        Block if suspicious        Ignore injected instructions   Rewrite or refuse violations

3. Alignment Training ​

Alignment training reduces content-layer and existential-layer risks at the root: teaching the model to refuse, to ask for clarification, and to be honest. The mainstream techniques are RLHF and DPO; frontier directions include Constitutional AI (constraining behavior with principles rather than human preference labels). This is a standalone topic in this book — for the technical details see Alignment: RLHF and DPO, and for how it relates to fine-tuning see Fine-Tuning and PEFT (LoRA).

4. Watermarking and Provenance ​

  • Content watermarking: embedding markers into generated text and images that are invisible to humans but machine-detectable (e.g., Google DeepMind's SynthID);
  • Provenance standards: C2PA Content Credentials, which record a piece of content's capture, generation, and editing history;
  • Detection tools: AI content detectors — of limited accuracy; usable only as auxiliary signals, never as evidence.

Watermarking directly serves societal-layer risks: it makes "this was AI-generated" verifiable, giving accountability and platform governance a technical lever.

5. Interpretability Research ​

Interpretability research asks "what is the model actually computing under the hood," and it is a long-term investment in safety: once we can read a model's internal representations, we gain a chance of spotting the early warning signs of "loss of control." One of the best-known bodies of work here is Anthropic's mechanistic interpretability research. It is also the technical backbone of the "transparency" principle (see Section 6).

6. Sandboxing and Least Privilege ​

Once AI upgrades from "answering questions" to "executing tasks" (agents), the set of "minimal actions the AI can take" must be strictly controlled:

  • Sandboxed tool calls: code execution, file access, and network requests all run in isolated environments;
  • Least privilege: agents get only the minimum permissions needed to complete the task, and critical actions require human confirmation;
  • Audit logs: every tool call is recorded — replayable and traceable.

The first principle of agent permission design

Always assume that the instructions an agent receives may be tainted — design its permissions as you would for an "untrusted program," not a "trusted employee." For hands-on guidance see Building an Agent from Scratch and Manus and Agent Applications.

4. The Governance and Regulatory Landscape (as of Mid-2025) ​

Global governance takes the shape of "three poles plus one standards system": the EU pursues "risk-tiered, strict regulation"; China pursues "filing + category-based regulation + content labeling"; and the US pursues "executive orders rising and falling + state-level backfilling + industry self-regulation."

The EU: Risk-Tiered Regulation Under the AI Act (the World's Strictest) ​

The EU AI Act entered into force on August 1, 2024. It is the world's first comprehensive AI law, and its core idea is to attach different obligations to different risk tiers.

Risk TierCovered ApplicationsCore Obligations
Unacceptable riskSocial scoring, subliminal manipulation, and the likeBanned outright
High riskCritical infrastructure, employment, education, law enforcement, biometrics, and the likeRisk management systems, data governance, human oversight, pre-market assessment, registration
Limited riskChatbots, deepfakes, and the likeTransparency obligations: disclose that users are interacting with AI, label synthetic content
Minimal riskThe vast majority of general-purpose applicationsEssentially no mandatory obligations

Implementation is phased: the prohibited-practice provisions took effect in February 2025; obligations for general-purpose AI (GPAI) models apply from August 2025; and obligations for high-risk systems apply in full from August 2026. Note: the Act applies equally to Chinese companies offering services in the EU market.

China: Filing + Category-Based Regulation + Content Labeling ​

  • The Provisions on the Administration of Deep Synthesis in Internet Information Services (effective January 10, 2023): deep synthesis content must be labeled;
  • The Interim Measures for the Management of Generative AI Services (effective August 15, 2023): generative AI services offered to the public require a security assessment + algorithm filing + large-model launch filing, with explicit requirements on training-data legality, content safety, and user rights;
  • The Measures for Labeling AI-Generated Synthetic Content (effective September 1, 2025): mandatory explicit/implicit labeling of AI-generated content.

The regulatory keywords: filing, labeling, content safety, data legality. Virtually any model or application offered to the public has to go through the filing process.

The US: Executive Orders Rising and Falling, and "Soft Regulation" ​

  • On October 30, 2023, President Biden signed the Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence (EO 14110): it required developers of the most powerful models to report safety test results to the government, promoted watermarking and content authentication, and set up AI safety research bodies;
  • In January 2025, Trump signed an executive order revoking EO 14110, shifting the federal stance toward "deregulation + promoting innovation";
  • The backfilling forces: state legislation (several states are advancing large-model safety bills), federal agencies enforcing under existing mandates (e.g., the FTC on false advertising, the EEOC on employment discrimination), and court precedents (e.g., the Air Canada chatbot case — see Section 5).

Standards Bodies: Technical Consensus Ahead of the Law ​

StandardPublisher and DateContent
NIST AI Risk Management Framework (AI RMF)NIST (US), January 2023A full-cycle risk methodology built on Govern, Map, Measure, Manage (GMMM)
ISO/IEC 42001ISO/IEC, December 2023The world's first international AI management system standard (as ISO 9001 is to quality management)
OECD AI PrinciplesOECD, 2019The first intergovernmental AI principles, cited in regulatory documents across many countries
EU AI Act supporting standardsCEN/CENELECHarmonized technical standards for conformity assessment of high-risk AI

The bottom line

Regulation sets the "floor," standards provide the "method," and technology provides the "capability." Mature, compliance-minded companies stack all three: use regulation to set goals, use ISO 42001 to build the management system, and use red-teaming and evaluation to implement the technical measures.

Corporate Compliance Checklist (Written for Engineering and Product Teams) ​

  • [ ] Inventory your models and use cases, and classify them by risk tier (you can borrow the AI Act's tiering logic);
  • [ ] Public-facing services: complete the filing/registration/security assessment process;
  • [ ] AI-generated content: implement both explicit and implicit labeling;
  • [ ] Data compliance: licensing, anonymization, and retention periods for training/fine-tuning data;
  • [ ] Incident response: playbooks and escalation paths for content incidents and data breaches;
  • [ ] Audit logs: inputs, outputs, tool calls, and model versions all replayable;
  • [ ] Supply chain review: assess the compliance status of purchased APIs and open-source models (e.g., whether data leaves the country).

5. Real-World Incidents: Risk Is Not a Theory ​

Three widely reported cases, spanning the system layer, the societal layer, and the physical world — plus one governance-side court ruling.

Case 1: Microsoft Tay — The Chatbot That "Learned to Be Bad" (March 2016) ​

On March 23, 2016, Microsoft launched the chatbot Tay on Twitter, advertising that it would understand you better the more you chatted. Within 24 hours of launch, users — through massive repeated inputs and targeted poisoning — baited Tay into posting racist, sexist, and otherwise offensive remarks, and Microsoft was forced to take it offline on March 25.

Lesson: Public conversational systems are inherently exposed to adversarial input; "learning from conversations" without content filtering and rollback mechanisms amounts to handing control of the model to attackers. It is a textbook case of data poisoning and prompt contamination.

Case 2: The Hong Kong Deepfake "Video Conference" Scam (February 2024) ​

According to reports, in February 2024 a finance employee at a multinational company in Hong Kong joined a "video conference with multiple attendees" in which every "executive" — including the chief financial officer — was generated by deepfake technology. Based on that meeting, the victim approved transfers and was defrauded of about HK$200 million (roughly US$25 million).

Lesson: Visual and voice verification can no longer be trusted; social engineering attacks gain devastating new power when combined with generative AI. Companies must build institutional defenses such as "large transfers require offline or two-person verification" — what you are defending against is not just AI, but AI-enhanced social engineering. This kind of content is the "dark side" of Diffusion Models and Generative AI.

Case 3: The Uber Self-Driving Car Fatality (March 2018) ​

On March 18, 2018, Uber's autonomous test vehicle (a Volvo XC90) struck and killed a pedestrian pushing a bicycle across the road during a nighttime test in Tempe, Arizona — the first publicly known case of a self-driving vehicle killing a pedestrian. The NTSB investigation found that the system's pedestrian detection was deficient, and that the in-car safety operator was looking down at a phone at the time and failed to take over in time.

Lesson: Safety is not "just add another sensor" — it is a reliability problem for the entire chain of perception, decision-making, human takeover, and monitoring and alerting — echoing directly the "monitoring and alerting" item on the Section 7 engineer checklist.

Governance-Side Addendum: The Air Canada Chatbot Case (Ruled February 2024) ​

The Civil Resolution Tribunal of British Columbia, Canada, ruled that Air Canada was liable for the incorrect refund policy information provided by its chatbot, and ordered it to compensate the passenger. The court's logic: the chatbot is the airline's "spokesperson," and its mistakes are the company's responsibility. This set a precedent pointing toward "when AI customer service errs, the company is liable" — a warning any company that puts generative AI in front of customers should heed.

6. Responsible AI Principles: From "Fixing It After the Fact" to "Safety by Design" ​

Responsible AI elevates safety from a "testing phase" to an "organizational principle." Six principles recur across the industry:

PrincipleMeaningImplementation Examples
FairnessTreat different groups equitablyReport metrics broken down by sensitive attributes, bias audits
TransparencyUsers know they are interacting with AI and that content is AI-generatedBot identity disclosures, labeling of generated content
ExplainabilityDecisions come with reasonsProvide grounds when refusing service, keep decision logs
PrivacyData minimization, no leaks, no misuseAnonymizing training data, deleting conversation logs after a retention period
AccountabilitySomeone is answerable for the AI's behaviorClear ownership of responsibility, incident escalation and compensation mechanisms
RobustnessReliable under adversarial input and distribution shiftRed-teaming, adversarial training, production monitoring

The key shift: from "fixing it after the fact" (incident → PR → compensation) to "Safety by Design / Responsible by Design" — building safety into the entire lifecycle of requirements, architecture, data, training, deployment, and operations, instead of "bolting it on" right before launch. This is the exact opposite of the "treating safety as a patch" anti-pattern in Common Pitfalls and Anti-Patterns.

7. What Engineers Can Do Day to Day: An Actionable Checklist ​

Safety is not just the security team's job — frontline engineers make safety decisions every single day. Ordered by "evaluate → defend → monitor."

1. Evaluate First: Build a Safety Baseline Before Launch ​

  • Write safety cases into your evaluation set: harmful requests, jailbreak templates, injection payloads, sensitive topics;
  • Track safety metrics: attack success rate, refusal rate, violation rate — and the false-refusal rate (over-blocking = degraded usability);
  • Run a safety regression before every model release (methodology in Building an LLM Evaluation Suite).

2. Least Privilege: Give the Model as Few "Hands" as Possible ​

  • Tool calls follow least privilege; refuse "do-everything tools";
  • Critical operations (transfers, deletions, sends, deployments) require human confirmation;
  • Put code execution and file access in a sandbox (details in Building an Agent from Scratch).

3. Content Policy: Guardrails Don't Run on Prayers ​

  • Spell out role boundaries and refusal rules in the system prompt;
  • Attach filters and classifiers on both the input and output sides;
  • Treat tool-returned data as "untrusted input" — parse it separately before use rather than concatenating it into the prompt.

4. Monitoring and Alerting: The Safety Radar for Production ​

  • Monitor anomalies in real time: the share of violating outputs, jailbreak attempt frequency, injection attack signatures, abnormal tool calls;
  • Set alert thresholds and make sure someone responds — don't just "log it and move on";
  • Retain logs (inputs, outputs, model versions, timestamps), replayable and traceable;
  • For production monitoring and ops practices, see Deployment and Inference Optimization in Practice.

The bottom line

The bar for going live is not "no vulnerabilities found" but "every required defense is in place, and every defense is measured." Replacing "we tested it" with "we have monitoring metrics, alerts, and rollback" is the dividing line between mature teams and amateur ones.

8. Trade-offs ​

  • Safety vs. usability: the stricter the filtering, the more false refusals and the worse the user experience. The industry approach is tiering: strict for high-risk instructions, lenient for low-risk scenarios;
  • Moderation cost vs. latency: adding classifiers or moderation introduces extra latency and compute cost — choose per scenario: offline batch workloads can afford heavy moderation, real-time interactions need lightweight ones;
  • Compliance cost vs. innovation speed: filing, audits, and management systems carry real costs, but for public-facing services they are already "table stakes," not "a bonus";
  • First-mover dividends vs. latecomer lessons: early entrants paid dearly in "safety lessons"; mature teams treat the safety budget as insurance, not cost.

Common Misconceptions ​

  • ❌ "Open-source models don't need compliance": services offered to the public are equally subject to content safety and labeling obligations;
  • ❌ "We just call an API — safety is the vendor's problem": supply chain risks must be assessed and managed by you;
  • ❌ "We red-teamed it, so we're safe": red-teaming covers known attack surfaces; new attack techniques keep emerging, so iteration must be continuous;
  • ❌ "Safety = content filtering": system-layer injection, societal-layer fakes, and existential-layer alignment go far beyond what content filtering can cover.

Further Reading ​

References ​