# The LLM Jailbreak and Moderation Arms Race

Attackers keep finding encodings that pass input moderation while the model reconstructs the blocked intent, forcing safety controls to move from input filtering toward output and trace inspection.

- Conviction: 40 / 100 (forming)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-07-31T00:00:00.000Z
- Last updated: 2026-07-31T06:02:07.324Z
- Canonical: https://polylog.news/ai/trends/llm-moderation-arms-race
- Publisher: Polylog
- Affected regions: Global

## Recent evidence

- [confirms] New Jailbreak Uses Dual-Layer Encoding to Reconstruct Blocked Prompts Past Moderation (2026-07-31): The RoguePrompt jailbreak uses dual-layer encoding so the model itself reconstructs a disallowed request during generation, bypassing the input-moderation layer — a direct instance of encodings defeating input filtering and pushing controls toward output/trace inspection.
