What it is
AI safety asks two questions. First: does the system do what it should, and nothing harmful? Second: is it protected against people who try to misuse it? For a company both come down to the same thing — the AI must not cause damage it was not supposed to cause.
The main risks
| Risk | What happens | Countermeasure |
|---|---|---|
| Prompt injection | Hidden instructions in a web page or email take control of the model | Separate data from instructions, limit rights |
| Data leakage | Confidential data appears in answers or leaves the company | Clear data rules, access control |
| Jailbreak | Users trick the model into ignoring its rules | Filters, testing, monitoring |
| Excessive agency | An agent does more than it should | Least privilege, human approval |
| Hallucination | False statements are taken as fact | Checks before use |
The list follows the OWASP Top 10 for LLM applications.1
Frameworks
In practice
- Keep a list of AI tools and who may use them.
- Give AI systems only the access they need.
- Let a person approve irreversible actions.
- Test with attacks before going live, and log what the system does.
Key terms
- Alignment
- Making a model’s behaviour match human intentions.
- Red teaming
- Deliberately attacking a system to find weaknesses before others do.
- Least privilege
- Only the minimum rights a task needs.
Latest research
New papers and reports that mention this term, found by our daily source scan. One line per source, quoted as published.
- techcrunch.comLooks like that’s the issue confronting Google’s Open Source Software Vulnerability Rewards Program, where researchers were rewarded for finding vulnerabilities in the company’s open source software.2026-10-05
- arxiv.orgSubjects: Artificial Intelligence (cs.AI) ; Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)2026-10-05
- labs.cloudsecurityalliance.orgThis document may not be modified or altered.2026-10-04