siddhant

Knowledge / Artificial Intelligence

AI Safety and Governance

Technical and organisational approaches to reducing AI failures, misuse and unacceptable risk.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

Scope

AI safety is a broad engineering and governance problem rather than a single algorithm. It includes preventing predictable failures, evaluating systems against realistic adversarial conditions, managing misuse and reducing the chance that powerful systems produce unacceptable consequences.

02

Technical risks

  • Robustness failures under distribution shift
  • Unreliable or fabricated outputs
  • Reward misspecification
  • Unsafe exploration
  • Security vulnerabilities
  • Unexpected interactions between automated components

03

Alignment

Alignment refers broadly to making a system's behaviour correspond to the intentions and constraints its designers and users actually care about. The difficulty is that objectives are often incomplete proxies for what humans want.

Current approaches include preference-based training, rule-based constraints, adversarial evaluation, red-teaming and methods for monitoring or interpreting model behaviour.

04

Evaluation

Evaluation should combine benchmark performance with task-specific testing, stress testing, adversarial testing and real-world observation. A benchmark can demonstrate capability on a defined distribution but cannot establish general reliability by itself.

05

Risk management

NIST's AI Risk Management Framework provides a voluntary framework for organisations developing, deploying or using AI systems. Its core functions are Govern, Map, Measure and Manage. The framework treats risk management as a continuous lifecycle activity rather than a one-time certification.

References

  1. Elham Tabassi. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, 2023.
    https://www.nist.gov/itl/ai-risk-management-framework
  2. Stanford Institute for Human-Centered Artificial Intelligence. The 2026 AI Index Report.
    https://hai.stanford.edu/ai-index/2026-ai-index-report

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.