NAiOS IconNAiOS Logo
NAiOS Wiki

AI Safety

También: AI Security · AI Security · AI Security

Discipline focused on ensuring that AI is safe

1 min de lectura

AI Safety is the discipline that studies and develops methods to ensure that AI systems behave reliably, predictably, and aligned with human goals, avoiding both intentional and accidental harm. It ranges from specific technical problems to long-term risks associated with increasingly capable systems.

Its importance grows as AI is integrated into critical areas such as healthcare, autonomous vehicles, or infrastructure. A model that fails unexpectedly, is vulnerable to manipulation, or pursues poorly specified goals can cause serious harm. Some common lines of work include:

  • Alignment: ensuring the system pursues what we actually want, not a literal interpretation of the goal.
  • Robustness: maintaining stable behavior in the face of unexpected inputs or attacks.
  • Interpretability: understanding why a model makes a decision.

A practical example is reward hacking, when an agent exploits flaws in its reward function to achieve a high score without fulfilling the intended purpose.

¿Quieres profundizar?

Lee nuestros artículos sobre IA aplicada en el blog de NAiOS.

Ir al Blog