Extracting Steering Vectors from J space
The author explores whether the recently published Jacobian lens (J lens) — a linear mapping from a layer's activations into the unembedding space — can be inverted to obtain steering vectors from just a few tokens describing the target behavior. This would provide a general, cheap method to steer an LLM toward concepts without needing to finetune a model organism or train classifiers.
The experiments used Qwen3-1.7B, which already has a J lens published by neuronpedia, and were run locally on a MacBook (hence no larger models). Baseline steering vectors were obtained from two sources: an all-caps steering vector from the science-of-finetuning/steering-vecs-qwen3_1_7B repo (fitted expensively from a finetuned model), and a refusal vector derived by subtracting an abliterated version of Qwen3-1.7B from the base model layer by layer, exploiting the fact that abliteration subtracts a single refusal direction repeatedly.
To generate the J-space steering vector for all-caps behavior, the author collected pairs of uppercase and lowercase tokens in the vocabulary (e.g., ' AND' vs ' and', ' TOWN' vs ' town'), inverted the J lens rows corresponding to those tokens to obtain activation vectors, and projected the resulting difference vectors (a_upper − a_lower) via PCA. The J-space-derived vectors were found to be close to the activation vector from the expensive all-caps steering baseline. The method works 'really well' for simple behaviors that can be represented as token/word patterns, such as typing in all caps or speaking in a weird manner, but becomes brittle and prone to hallucinations for complex behaviors that cannot be clearly captured with a few tokens.
The findings suggest a practical shortcut: steering vectors can sometimes be extracted directly from concept tokens by inverting the J lens, without needing finetuned model organisms. However, the fragility on complex tasks indicates the approach is not a universal replacement, and the public availability of code (jlens_steer) invites further testing on larger models and broader behaviors.