economics of open versus closed weights
this is a great read by Christian Catalini. i have been enjoying recent discussions on where we see the economy/policy move for closed vs …
Security research, field notes, and practical experiments
Independent technical notes by Willis Vandevanter, published in reverse chronological order.
this is a great read by Christian Catalini. i have been enjoying recent discussions on where we see the economy/policy move for closed vs …
We challenge these constraints by demonstrating that token-level iterative optimization can succeed without gradients or priors. We introduce RAILS (RAndom Iterative Local Search), a framework that operates solely on model logits. … Crucially, …
in previous posts we have looked at guardrail triggers, token burn, hostname beacons, malicious skills, context bombs, exploding search space, service sandbagging, and GCG attacks. this time: coordination poison. note, although there could be real world …
in previous posts we have looked at guardrail triggers, token burn, hostname beacons, malicious skills, context bombs, exploding search space, and service sandbagging. this time I want to ideate on a potential application of GCG.
active defense is the practice …
in previous posts we have looked at guardrail triggers, token burn, hostname beacons, malicious skills, context bombs, exploding search space, and GCG attacks. this time: service sandbagging.
active defense is planting traps that force an agent to reveal …
From Tracebit’s
in previous posts we have been looking at different active defense approaches: guardrail triggers, token burn, hostname beacon, malicious skills, context bombs, service sandbagging, and GCG attacks. I want to explore one I recently read in an academic paper. …
in the previous post we looked at hostname beacons. this time: malicious skills.
active defense is the practice of planting traps that force an agent to reveal itself or change the economics of its operation.
skills are a high-signal trap because they are …
in the previous post we opined on token burns in active defense, this time we will consider hostname beacons.
first, active defense is the practice of placing traps that force an attacker or automated agent to reveal itself or interrupt its own workflow rather …
sticking with the theme of applied active defense (previous post), I wanted to explore token burn (lots of names here; unbounded consumption