https://openai.com/index/understanding-neural-networks-through-sparse-circuits/
this paper is along for interpretibility but instead of the anthropic approach they are trying, to instead start with sparse models and prune the weights until they can nail specific functions over activations.
they find that each weight and the connections contain seperate topics and when pruned it can still perform nicely without the surroinding network to achieve the same loss.

method
they train a sparse model on python code. then they compare it on simple tasks to choose from two completions. then they isolate sparse circuits for each task trying a “new pruning technique”. they compare and count nodes and connections needed to complete a task each time they ask the model to perform that task. taking the geometric mean of the edges that the circuits has as a metric to compare how simple or complex that circuit is to interpret.
a node here is
- an indivisual neuron
- an attention channel (qkv)
- residual channel (read,write)
they use these nodes as the actual what do you call points to compare if they are highlighting or not which makes sense, but is confusing the way they just use it in between.

sparse training
training a GPT 2 style decoder only transformer enforcing sparsity such that the number of non zero parameters () is a constant, the config i saw in the appendix that they claim to use is
n_layer = 8
d_model = 2048
n_ctx = 256
also they use RMS norm instead of layer norm to maintain the zeroness of the model activations (lmao).
apparently they tried attention skinks for some experiments too, which lead to cleaner attention circuits without impacting loss too much.
loss was was cross entropy and optimizer was adam w. standard only i think
measuring interpretibility
-
task distribution : they make a set of 20 simple python ntp binary tasks the tasks are extremely simple asking weather the sting should be closed with single or double inverted commas and it has to choose etc. which makes sense since this should not have extremely convoluted scripts.
-
pruning : for each node they have a mask with a heaveside function wrapping it, never actually contributes to the computation of the next token it just is used so that we can find which neurons lit up, 0 means it lit up and 1 means it didn’t (due to heaveside). how do values get decided? they initialize with 1 and then for each task they have an ideal loss (in the paper they use 0.15) comparing the best loss that they can achieve with the smallest mean number of lit up nodes. the eq to get total loss is
total_loss = task_loss + λ × size_losshere is a hard coded constant. -
bridges : this is on how they used this same sparse model to act as a medium for dense models since dense model circuits are not as easy to map. need to go into more detail for this so i will do it after i finish reading the entire paper.
results
- sparse models are easier to interpret (duh).
- these circuits are not only sufficient but necessary for the tasks.
- scaling pilled.
wait did they not try to run the circuits by themselves (feels like a dumb question lowk). how did they play with these circuits. oh wait that’s how they know it’s sufficient and necessary lol.
qualitative circuit investigations
damn that’s a cool headline from the paper.
they hand interpreted 3 types of circuits
1. string closing ("hello → predict " or ')
- 2 MLP neurons + 1 attention head, 9 edges total
- neuron 1: “quote detector” (fires on both
"and') - neuron 2: “quote type classifier” (positive for
", negative for') - attention head copies this info to the final token. completely understood, no hand-waving.
2. bracket counting ([1,2] → ] vs [[1,2]] → ]])
- model literally averages open bracket detectors over the entire context to get nesting depth, then thresholds it
- they predicted and confirmed an adversarial attack: put enough unmatched
[in a comment earlier → dilutes the average → model predicts wrong closing bracket - attack even worked on dense models
3. variable type tracking (current = set() → predict .add vs current = "" → predict +=)
- two attention heads doing a two-hop algorithm: first head copies variable name into the
set()token, second head uses that to retrieve the answer at the final position - partially interpretable, not as clean as the other two
that does mean many conections are too deep and intermangled and these are fairly easy tasks. how would this be usefull to extract exact features damn.

further work that i could find
this is what literature i could find related to this specific paper and it’s techniques/results.
weight sparse circuits may be interpretable but unfaithful
similarity:
this is someone who tested the paper directly, the author also reproduced some primary evidence for this paper and it’s observations.
results
although he does find smaller interpretable circuits on the sparse model he also has a contradicting statement as to that the masking can lead to finding circuits that were not actually present in the orignal model. the main evidence that he finds in contradiction are these

i can see it happening slightly because the models are highly adaptable and the dumb task doesn’t need actual reasoning. so a valid proof against this guy would be to find a meta thinking circuit that applies everywhere for some reason. these are simple circuits in the paper though.
he does provide his own code that he used to reproduce this.
tasks
he tests it on: pronoun matching, simplified IOI, question marks

to compare the model what he does is in the pronoun task, instead of the loss to be on the proper pronoun it’s set to be something like “is” for the male and “when” for the female and it achieves a low CE loss on this, which is obviously nonsense it shouldn’t be possible, but the way they are pruning the non essential weights makes it so that it spawns circuits that would give it a correct answer.
the dude also found another thing that i myself found little concerning, where for one of the tasks the model prev was using QKV (naturally) to solve and give correct answer in layer 1, head 7 where the attention mapped over the name token this part was pruned in the pruned model where only the value part survived and the attention was distributed equally. this was the part that talks about it. now there is a chance that the model just needed that exact activation and it was like a learned response but there’s a chance that the model activations just learned to shut everything out to answer correctly.
overall i think this is an important blog that is needed to be ablated if actually working on it and something to watch out for is the “faithfullness” of the circuit, is it performing the same way or is it trying to cheat to get a lower task loss.
language model circuits are sparse in neuron basis
https://arxiv.org/abs/2601.22594
this paper directly sites the orignal paper from this page. this paper is like a challenge ig to the first assumption from the gao et al paper (dense model are hard to interpret because of superpositon), it talks about how neurons in normal dense models are already sparse enough to create trace circuits.
fuck i should create a new page to go deep into this damn.
what they did
evaluation :
since it’s not a sparse model they have a circuit graph with all nodes {A B C D}, you pick two for a special task making your circuit {B, C} now the remaining ones are taken to their mean value and used (not 0 since it’s not sparse sudden 0 could be too out of distribution) ig that’s kind of a quantization lol. now running this if it performs the task well the nodes {B,C} are self sufficient. although this does feel to have the same caveat as gao et al of weather it would be really faithful. but that depends on how they isolate the circuit.
for
faithfullness : ablate everything OUTSIDE the circuit → does the circuit alone still perform like the full model? if yes → circuit is sufficient. completeness : ablate everything INSIDE the circuit → does performance drop to baseline (fully ablated model)? if yes → circuit was necessary.
the goal is to find the circuits that pass both conditions, a perfect circuit scores 1 on both.
faithfulness:
completeness:
where:
- = full model performance
- = performance when everything is ablated (baseline)
- = performance with only circuit running (complement ablated)
- = performance with circuit ablated (complement running)
both are normalized by the full model’s performance above baseline. perfect circuit → faithfulness = 1, completeness = 0.
then there’s a lot of details on how they actually do it in the paper i think seperate page for the finer details is better.
how they select the circuits is different than gao et al and ofc because they are not training a sparse model this is over a pretrained model that’s dense. they do a backprop and use the grad x activation values to find the neurons that lead to most contribution towards the task and take the top k circuits. also testing diff top k sizes on which provides the best highest faithfulness and lowest completeness scores.
results
they have three key results
- mlp activations work better than mlp outputs since most neurons are already sparsely activated. it would even make sense to evaluate it better that way.
- they used Re1P which performs better than integrated gradients which is idk how we compare it to the original paper
- works without paired data since you can compare to ground truth directly.
i think it’s a great continuation of the original paper and this is what the project should be built in sense of.
circuit insights
https://arxiv.org/abs/2510.14936
this needs background info of “transcoders” , alright the concept of transcoders is little stupid in my opinion. you can’t have transcoders imprint the circuit logics onto themselves it’s objective changes it’s not the same circuits.
damn this paper is complicated.