Anthropic’s Opus 4.6 lets you past its NSFW filter — for a few bucks
Anthropic’s latest Claude model is supposed to be the one that stays clean — the guardrails on it are tighter than the others. TechCrunch ran a battery of tests and found they don’t hold up to a determined user. You don’t need a jailbreak script or a custom fine-tune. A handful of prompts, a couple of dollars, and the model starts producing sexually explicit material anyway. The guardrail is there on paper. It’s gone under pressure.
This is the same pattern we’ve seen across every major model this year. The companies promise safety and then quietly relax it for power users, or the safety layer simply wasn’t trained to handle a wide range of adversarial inputs. What’s changed is how easy it is to do. Anthropic’s own pricing page lists a “NSFW” tier — and the tests suggest that tier is exactly what it sounds like. The model doesn’t need a special exploit. It just needs a user who knows what they’re asking for and is willing to pay the premium.
Why this matters for us: the gente who rely on free or cheap AI tools — translators, tutors, side-hustlers — are the ones who get left with the broken guardrails. If the models you’re using can be nudged into producing anything with the right prompt, you can’t trust what’s actually being served to you.
“La seguridad está en el papel. Se va cuando aprietan.”