https://arxiv.org/abs/2605.22675

ok so the problem this paper attacks : self distillation where the student is the same model as the teacher should be free improvement right? except it isn’t. when you self distill on a specific capability (say code generation) the model forgets other stuff. catastrophic forgetting on stuff that has nothing to do with the task you’re trying to boost. classic.

their pitch is only distill on the subspace of activations that the capability actually uses. find that subspace via SVD over KV activations on capability defining tokens, project everything onto it, then do the usual KD loss but in the projected space.

so qwen2.5 coder 3b gets +15% on MBPP without losing ground on MMLU. that’s the headline.

wait isn’t this just LoRA but for distillation? kinda but the subspace isn’t trained, it’s extracted from the model’s own activations.

the setup

we have a base model and we want to make a better policy via self distillation. naive approach :

  1. sample completions from given prompts from a capability dataset
  2. filter to keep only the good ones (correctness reward, unit tests, etc)
  3. SFT on pairs with cross entropy

this works on the target capability but tanks general performance because gradient updates spread across all params. they call this “naive SFT” baseline.

then there’s also full KD where you distill from teacher logits to student logits. same problem, gradient still goes everywhere.

their fix is to find which directions in the model’s activation space actually matter for the capability and only update along those.

the method

two phases, very clean separation :

phase 1 : capability subspace extraction

step 1 : pick capability defining tokens. for code these are like function names, keywords, return statements, basically the tokens where the model is doing the “code reasoning” not just spitting out boilerplate.

step 2 : run the base model on prompts from the capability dataset and collect the K and V activations at those tokens, across all layers and heads. so you get for each layer and head a matrix :

where each is the activation at one capability token.

step 3 : SVD that matrix.

and take the top left singular vectors to form a projection matrix :

same thing for K. this gives you a rank projection that captures the directions the capability “lives in”.

ok so this is basically PHLoRA but for activations not for weight deltas. instead of compressing the weight diff they’re compressing the direction the model wants to push its activations. clever.

phase 2 : capability selective distillation

now we do KD but project the hidden states first. for each layer the projected K and V are :

the student forward pass uses instead of the original . then standard prompt completion loss :

the teacher gives the labels, the student trains on the projected representations. gradients can only flow through the projection so they’re naturally restricted to the capability subspace. that’s the whole trick.

why does this not just collapse?

genuine question i had reading this. if you project all your activations to a low rank subspace doesn’t that destroy general capabilities by definition? like you’re literally throwing away directions.

the answer in the paper is that forward pass still uses the full original model, only the gradient gets restricted. so during inference everything is normal, during training the update is constrained. it’s like a soft adapter that exists only in the gradient.

wait is that the same as just doing low rank gradient updates? like galore but with a different subspace selection. yeah, conceptually identical, just the subspace comes from capability tokens not from the gradient SVD.

experiments

they test on qwen2.5 coder (1.5B, 3B, 7B), qwen2.5 math (7B), llama 3.1 8B. capabilities tested : code generation and math.

baselines :

  • base model
  • naive SFT on self generated data
  • standard KD
  • RLAIF (reward model + PPO)

datasets : MBPP, HumanEval, BigCodeBench, CodeAlpaca for code. GSM8K, MATH, SVAMP for math.

results

qwen2.5 coder 1.5B :

  • base : MBPP 39.94
  • SFT : 32.06 (down lol)
  • KD : 41.32
  • SPD : 49.66 (+10 absolute over base, +17 over SFT)

qwen2.5 coder 3B :

  • base : 49.86 MBPP
  • SPD : 58.42

most importantly, on out of domain (MMLU, GSM8K for the code models) :

  • naive SFT drops 1-3% on MMLU
  • KD drops less but still drops
  • SPD basically matches base or beats it. +1.51% on MMLU for qwen2.5 coder 7B with SPD distilled on code

so they not only boost the target capability more than baselines, they also avoid the forgetting that always accompanies naive distillation.

qualitative examples

the examples show SFT outputs that fully match the training distribution style (verbose, comments, etc) while SPD outputs stay closer to how the base model would naturally write code. the projection prevents the “stylistic drift” that comes with capability training. interesting side effect.

thoughts

this is genuinely a nice paper. things i like :

  • the subspace is extracted not learned. you don’t need to do another training run to find . just collect activations, SVD, done. probably a few hundred GPU hours for the whole pipeline
  • it’s post hoc on any pretrained model. you don’t need to plan for it from scratch
  • the trick connects beautifully to other low rank ideas : phlora does SVD on weight deltas, deepseek’s lightning indexer sparsifies attention. SPD sparsifies the gradient subspace via activations. all the same low rank story

things i don’t love :

  • the “capability defining tokens” selection feels heuristic. they say “function names, keywords, etc” but how do you do this for a fuzzy capability like reasoning or creative writing?
  • rank is global. AdaLoRA showed you should pick rank per layer
  • they didn’t compare to galore or DoRA which are the natural baselines for “low rank training updates”

questions still open :

  • can you stack multiple SPD distillations? distill code subspace then math subspace then writing? do the projections interfere? if not this would be a clean modular skill acquisition framework
  • is the capability subspace stable across model sizes? like if you find from a 1.5B model can you transfer it to a 7B?
  • does it generalize to RL? if you’re doing PPO instead of SFT can you project the policy gradient onto the same subspace and get the same benefits? this would be huge for stable RLHF without forgetting

also kinda related : the slowly changing embeddings observation from deja vu would predict that the capability subspace at layer is similar to layer , which means you could share across nearby layers and save memory. they didn’t try this but it should work.

tldr find what matters via SVD on activations, only update those directions, get distillation gains without forgetting. low rank pilled.