VynarisEarly betaGet my API key

Refusal direction

A refusal direction is a pattern in a language model's activation space associated with declining certain prompts, such as an apology, warning, or redirect. Researchers can estimate it by comparing activations from refused and answered prompt pairs. Ablation methods try to reduce the influence of that pattern, but the direction is an analytical approximation, not a universal switch shared by every model. Measuring refusal behavior before and after an edit helps identify changes and possible capability loss. The result should be tested on the served artifact, not assumed from the editing method alone.

Vynaris is an inference gateway that routes each request to the cheapest right-sized model and shows the receipt. Get an API key or read the docs.