The LLM Jailbreak and Moderation Arms Race
Attackers keep finding encodings that pass input moderation while the model reconstructs the blocked intent, forcing safety controls to move from input filtering toward output and trace inspection.
forming · confidence 40 · Emerging (watchlist) · tracking since July 31, 2026 · updated July 31, 2026
Why the conviction moved
- Jul 31Strengthened +4
The RoguePrompt jailbreak uses dual-layer encoding so the model itself reconstructs a disallowed request during generation, bypassing the input-moderation layer — a direct instance of encodings defeating input filtering and pushing controls toward output/trace inspection.
Source trail
Supporting · July 31, 2026
New Jailbreak Uses Dual-Layer Encoding to Reconstruct Blocked Prompts Past Moderation
The RoguePrompt jailbreak uses dual-layer encoding so the model itself reconstructs a disallowed request during generation, bypassing the input-moderation layer — a direct instance of encodings defeating input filtering and pushing controls toward output/trace inspection.
arXiv (cs.CR)
Unlock full source trail, score history, and daily updates.
Unlock Trends