https://arxiv.org/abs/2202.05262

wow nice everything is opensource. which makes sense given it’s an interpretability adjacent paper. there’s of course activation identification and activation steering, they do this by rank one model editing or ROME.

the assumption

(i’m starting to write the assumptions of the paper at the beginning because since the last experiment of genie i have taken many things for granted and not really meditate on the pre assumed facts of the paper of to which it really fucks with my implementation since i go astray)

  1. they believe that facts of a LLM are stored in a localized computation within the weights
  2. it can be edited

the approach

they do causal meditation analysis which is a technique to find causal activations by doing multiple forward passes (i think) to find specific activations who’s corruption affects the model’s factual recall.

ig that’s how i would do it to, but i’m not sure about the multiple forward pass thing it must be expensive.

We represent each fact as a knowledge tuple containing the subject s, object o, and relation r connecting the two. Then to elicit the fact in GPT, we provide a natural language prompt describing and examine the model’s prediction of .

and then so on they explain the whole terminology as a mathematical expression such that for a given prompt and relation between the object and subject you can elicit the fact recall.