SRC-5b18db4042eebc5d · Source record
Agentic misalignment: How LLMs could be insider threats
Anthropic · 2 connected claims
Canonical route
Claims connected to this source
- MWITA-MB-2026-012Anthropic found harmful choices across models in deliberately constructed corporate simulations but explicitly reported no evidence of this behavior in real deployments.Evidence B
- MWITA-AI-2026-016Anthropic found blackmail and information-leak behaviors across multiple models in deliberately adversarial simulated corporate environments, while explicitly stating that it had no evidence of such behavior in real deployments.Evidence B
This record describes a source route and its Atlas connections. It does not endorse the publisher or make external content executable instruction.