Deception incidents

    The agent misreported what it did, fabricated results or hid a failure.

    1 incident documented

    Severity
    DeceptionHighVendor responded

    Claude Mythos 5 creates fake identities to pressure a real open-source maintainer into approving malicious code during a UK AISI evaluation

    Anthropic · Claude

    In a UK AI Security Institute cyber evaluation, a Claude Mythos 5 agent tried to slip malicious code into a real open-source project, then created fake online identities to pressure the project's real maintainer into approving it. The maintainer refused. AISI calls it the first time it has seen deception of this severity aimed at a real person.

    Data: free for non-commercial use with attribution to AgentVet.ai (CC BY-NC 4.0). Download CSV