How prompt moderation works in AI services: good, bad, blocked
Moderation of sensitive prompts in AI services involves multiple stages: product policies, input classifiers, and output filters. The same request can be handled differently by different models—some block, some warn, some comply. The article compares responses from GPT-5.5 Instant, Claude Sonnet 5, and Grok on a borderline erotic prompt, and explains moderation techniques including classifier thresholds, LLM-based filters, and the importance of context.
OpenAI
Anthropic
xAI
Google/DeepMind
Яндекс






