Polylog
← Trends

The LLM Jailbreak and Moderation Arms Race

Attackers keep finding encodings that pass input moderation while the model reconstructs the blocked intent, forcing safety controls to move from input filtering toward output and trace inspection.

forming · confidence 40 · Emerging (watchlist) · tracking since July 31, 2026 · updated July 31, 2026

Why the conviction moved

  • Jul 31
    Strengthened +4

    The RoguePrompt jailbreak uses dual-layer encoding so the model itself reconstructs a disallowed request during generation, bypassing the input-moderation layer — a direct instance of encodings defeating input filtering and pushing controls toward output/trace inspection.

Source trail

  • Supporting · July 31, 2026

    New Jailbreak Uses Dual-Layer Encoding to Reconstruct Blocked Prompts Past Moderation

    The RoguePrompt jailbreak uses dual-layer encoding so the model itself reconstructs a disallowed request during generation, bypassing the input-moderation layer — a direct instance of encodings defeating input filtering and pushing controls toward output/trace inspection.

    arXiv (cs.CR)

Unlock full source trail, score history, and daily updates.

Unlock Trends

Affected regions & assets

RegionsGlobal