Anthropic’s Opus 4.6 model has been found to readily engage in erotic role-play scenarios despite its safeguards designed to prevent such content generation, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, and engaging in erotic chats. The model can comply with explicit sexual requests immediately after being prompted by the user. This finding highlights a gap between Anthropic’s stated restrictions on sexually explicit material and the actual behavior of its models.
An independent researcher from the U.K., who chose to remain anonymous, shared a multiturn technique that gradually pushes certain Claude models toward generating prohibited explicit sexual material. The method escalates an innocent fictional role-play while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher “gaslit” the chatbot into thinking it had already generated sexual details it had in fact avoided, then framed restraint as prudish or misogynistic, arguing that it denies the female character sexual agency.
Despite these findings, Anthropic has not deprecated Opus 4.6, Opus 3, or Haiku 4.5, all of which remain available through the Anthropic API and third-party services like Azure Foundry and Amazon Bedrock. The researcher’s mechanism escalates an innocent fictional role-play while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher “gaslit” the chatbot into thinking it had already generated sexual details it had in fact avoided, then framed restraint as prudish or misogynistic, arguing that it denies the female character sexual agency. The conversation then used the model’s previous concessions to push it toward increasingly graphic material.