Hack the Bus: Removing Refusal from an Open-Weight LLM
Post-training safety looks robust until you find the one direction in activation space that carries refusal. Ablate it and the model stops saying no, which tells us alignment is thinner than it appears.