active defense and adversarial agents: guardrail triggers
the previous post briefly touched on active defense in the scope of AI agents and LLMs. active defense is the practice of placing traps and tripwires that force an attacker (or an automated agent) to reveal itself or interrupt its own workflow rather than …