Hack the Bus: Removing Refusal from an Open-Weight LLM

Post-training safety looks robust until you find the one direction in activation space that carries refusal. Ablate it and the model stops saying no, which tells us alignment is thinner than it appears.

August 2026 · 18 min · Simone Mattia