01
Scope
AI safety is a broad engineering and governance problem rather than a single algorithm. It includes preventing predictable failures, evaluating systems against realistic adversarial conditions, managing misuse and reducing the chance that powerful systems produce unacceptable consequences.
02
Technical risks
- Robustness failures under distribution shift
- Unreliable or fabricated outputs
- Reward misspecification
- Unsafe exploration
- Security vulnerabilities
- Unexpected interactions between automated components
03
Alignment
Alignment refers broadly to making a system's behaviour correspond to the intentions and constraints its designers and users actually care about. The difficulty is that objectives are often incomplete proxies for what humans want.
Current approaches include preference-based training, rule-based constraints, adversarial evaluation, red-teaming and methods for monitoring or interpreting model behaviour.
04
Evaluation
Evaluation should combine benchmark performance with task-specific testing, stress testing, adversarial testing and real-world observation. A benchmark can demonstrate capability on a defined distribution but cannot establish general reliability by itself.
05
Risk management
NIST's AI Risk Management Framework provides a voluntary framework for organisations developing, deploying or using AI systems. Its core functions are Govern, Map, Measure and Manage. The framework treats risk management as a continuous lifecycle activity rather than a one-time certification.