active defense and adversarial agents: service sandbagging
in previous posts we have looked at guardrail triggers, token burn, hostname beacons, malicious skills, context bombs, and exploding search space. this time: service sandbagging.
active defense is planting traps that force an agent to reveal itself or change …