Metatron Intelligence
Read today5 AI stories0 new papers437 model prices, 3 changed353 open models ranked

AI Wiki

AI safety

In short: AI safety covers two questions: whether a model behaves as intended, and whether it is protected against misuse. With agents that act on their own, both become an operational duty.

Updated 29 September 2026 · 2 min read · 4 sources

What it is

AI safety asks two questions. First: does the system do what it should, and nothing harmful? Second: is it protected against people who try to misuse it? For a company both come down to the same thing — the AI must not cause damage it was not supposed to cause.

The main risks

RiskWhat happensCountermeasure
Prompt injectionHidden instructions in a web page or email take control of the modelSeparate data from instructions, limit rights
Data leakageConfidential data appears in answers or leaves the companyClear data rules, access control
JailbreakUsers trick the model into ignoring its rulesFilters, testing, monitoring
Excessive agencyAn agent does more than it shouldLeast privilege, human approval
HallucinationFalse statements are taken as factChecks before use

The list follows the OWASP Top 10 for LLM applications.1

Frameworks

  • NIST AI Risk Management Framework (USA, 2023): govern, map, measure, manage.2
  • ISO/IEC 42001 (2023): the first certifiable management-system standard for AI.3
  • EU AI Act: legal duties by risk class, including cybersecurity for high-risk systems.4

In practice

  • Keep a list of AI tools and who may use them.
  • Give AI systems only the access they need.
  • Let a person approve irreversible actions.
  • Test with attacks before going live, and log what the system does.

Key terms

Alignment
Making a model’s behaviour match human intentions.
Red teaming
Deliberately attacking a system to find weaknesses before others do.
Least privilege
Only the minimum rights a task needs.

Latest research

New papers and reports that mention this term, found by our daily source scan. One line per source, quoted as published.

Sources

  1. OWASP Top 10 for Large Language Model Applications
  2. NIST AI Risk Management Framework 1.0 (January 2023)
  3. ISO/IEC 42001:2023 — AI management systems
  4. Regulation (EU) 2024/1689 (AI Act), EUR-Lex

Related terms

All terms Today's headlines