Taming Incidental Polysemanticity in Toy Models: How Network Training Choices Affect Feature Entanglement
I ran some experiments on toy models to test whether there are factors during a network's training that influence how polysemantic its representations end up being (when probed using SAEs). The L2/GELU configuration had 17.9% lower measured Interference than the L1/ReLU configuration (Interference of 7.14 vs. 8.70, t=2.8). This is in a toy setting and has lots of caveats, but was an interesting thing to look at for getting better mechanistic understandings of when polysemanticity appears.
This exploratory study compared full training configurations rather than isolating every factor. Because the headline L1/ReLU and L2/GELU settings differ in both regularization and activation, the 17.9% result is best read as a configuration-level observation, rather than a controlled estimate of the effect of L2 alone. The proposed mechanism remains a hypothesis. The original significance claims need rechecking: for a two-sided t-test, t(16) = 2.8 gives p ≈ 0.013, while t(16) = 2.1 gives p ≈ 0.052. Those do not match the originally stated p < 0.01 and p = 0.04. The ten-seed description and reported degrees of freedom also need reconciliation. This revision retains the descriptive results without treating the historical significance claims as validated.
Motivation
Polysemanticity is a big problem for MI, and a big concern for safety. Historically it has been understood as a consequence of superposition (i.e. if the network is trying to represent more features than it has dimensions to represent them on). But more recently it has been found that polysemanticity can also appear incidentally, even in overcapacity settings, due to a variety of factors during training (Lecomte et al. 2023). Some examples of these are regularization (in particular things that induce winner-take-all dynamics) and noise (which can induce chance correlations between features and neurons).
Whilst there have been lots of papers trying to improve the ability of SAEs to decompose features (TopK SAEs, JumpReLU SAEs, Gated SAEs etc.), much less work has been done on whether factors during the training of a network affect how well SAEs can decompose their representations (i.e. how entangled / polysemantic the features in those representations are). This is important, because if a network is trained in a way that induces a highly polysemantic representation in its activations, then even the best SAE might struggle to decompose it.
In this post I do some experiments on toy models to try to test whether some factors (during the training of the toy model), which I hypothesize may reduce incidental polysemanticity, in fact make it easier for an SAE to decompose the representations of a network. Some factors I looked at included: starting with an orthogonal initialization (to reduce chance correlations between features and neurons), using L2 regularization instead of L1 (to reduce winner-take-all effects), using a GELU activation function (to help gradient flow), and adding in noise with positive kurtosis during training (to break up lock-ins between features and neurons).
Training with L2 regularization leads to significantly less polysemantic features than training with L1 regularization. This suggests that, at least in overcomplete settings, the phrase "sparse = interpretable" maybe should be reversed.
What did I do to get these results?
At first, the results didn't make sense. In all configurations, the networks exhibited strong polysemanticity, and, when I tried to train the SAE, I got "dead SAEs", i.e. the L0 sparsity of the SAE (i.e. the fraction of latent features that fire on an input) was 0.0000. In other words, the SAE only predicted 0 for all inputs, so all the other metrics were meaningless, too.
The issue was the dying ReLU problem: Because I applied strong L1 regularization (at first, I used λ = 0.01) to the latents of the SAE, and the ReLU returns 0 for all negative inputs (and thus, the gradient is 0), some of the neurons in the SAE died, i.e. they predicted 0 and couldn't recover because the gradient was 0. I think the issue was exacerbated in my toy setting because, due to the correlated, sparse nature of the data, the activations had a relatively low variance.
I dealt with the issue by:
- decreasing the value of lambda (to 1e-4, which was high enough to encourage sparsity but low enough to not cause issues - this is in line with what people do when training SAEs on scaled models)
- using LeakyReLU with α = 0.01 in the SAE
- setting the biases of the encoder to small positive values (sampled from a uniform distribution between 0 and 0.1)
- normalizing the activations before passing them through the SAE (i.e. I passed
(acts - acts.mean()) / acts.std()through the SAE); and - training the networks for 2000 epochs (instead of 1000) and the SAE for 1000 epochs (instead of 500).
In the end, the L0 sparsity of the SAE was between 0.42 and 0.49 in all cases, which is in the recommended range of 0.3 - 0.6 (as it's a tradeoff between reconstruction and sparsity, see Anthropic's blogpost on SAEs). Only then, I got sensible signals for polysemanticity.
To get a strong signal, I made the data "polysemanticity-tempting", i.e. I grouped the features into groups (by default, I used 8 features, divided into three groups: [0, 1, 2], [3, 4], and [5, 6, 7]) and set it up so that features within groups were highly correlated (by default, between 0.5 and 1.0) and co-activated with 60% probability. I also made the target function depend on the features in a non-linear way (by including non-linearities like products of features, sin, tanh, and products between features from different groups in the target function). Additionally, I set the importances of the features to 0.9i, where i is the index of the feature. This helped to increase the signal when using importance-weighted polysemanticity metrics, as the more important features were also the ones where polysemanticity was costly.
Experimental Setup
I train a 5 layer MLP to predict the targets. The input is the 8 sparse features and the output is the 8 targets. The hidden layer is slightly larger than the input (12) to ensure that polysemanticity is not due to underparameterisation.
I train the network under a variety of different settings to see which are most effective at mitigating polysemanticity:
- Weight Init: Default (Random, from normal distribution) vs Orthogonal
- Noise: Bipolar (Uniform, within 0.1) vs Positive Kurtosis (T-Dist with df = 3)
- Regularisation: L1 (loss += λ × |weights|) vs L2 (loss += λ × weights²)
- Activation: ReLU vs GELU (can't have dying neurons)
While this gives 16 potential ablational settings, I'm mainly interested in ablating from the baseline settings (Random, Bipolar, L1, ReLU) to the maximum mitigation settings (Orthogonal, Positive Kurtosis, L2, GELU), so I train 8 networks.
The SAE is overcomplete (d_sae = 16) and is a simple linear encoder (with positive bias) and LeakyReLU activation (12 → 16), followed by a linear decoder (16 → 12). I always use L1 regularisation on the SAE latents, as this is a way of ensuring that the SAE latents are sparse. I train this SAE on the hidden layer of the network after training, in order to decompose the network's representations. This means that when I'm ablating between L1 and L2 regularisation, I'm ablating the regularisation on the network, not the SAE. This ensures that when I compare the different ways of training the network, I'm using the same method of decomposition.
- Mean Absolute Cosine Similarity (MACS): the mean of the absolute values of the off-diagonal cosine similarity values for the SAE decoder weights. Lower is better, as it shows that the SAE features are more disentangled.
- Interference: a weighted version of polysemanticity. ∑i ≠ j cos2(di, dj) × Ii × Ij, where I is the importance of each feature. Again, lower is better as it shows that important features aren't entangled.
- L0 Sparsity: the fraction of SAE features which are active for a given input (i.e. above a certain threshold). Should be between 0.3 - 0.6 for SAEs.
- Sparsity_W4: a measure of how peaked the weights are. Higher is better, as it shows that the SAE weights are sparse.
- SAE_MSE: the MSE of the SAE. This isn't necessarily a good measure of interpretability, as it just shows how complex the SAE needs to be in order to reconstruct its inputs.
- Train/Val MSE: the MSE of the network on the training/validation data, for the task of predicting the targets.
- Calibrated_MACS: the MACS value of the SAE, after being z-scored against 200 permutations of the MACS null distribution, and then sigmoid-normalised to [0, 1]. If the MACS is similar to that expected by chance, this value should be near 0.5.
Results: L1 on its own is not enough for overcomplete networks
As a baseline, I used Random Init, Bipolar noise, L1 regularization and ReLU activations. The results I got were:
| Metric | Value |
|---|---|
| MACS | 0.2954 ± 0.0294 |
| Interference | 8.7022 ± 1.4480 |
| Calibrated_MACS | 0.5085 ± 0.1117 |
| L0_Sparsity | 0.4365 ± 0.0596 |
| Train_MSE | 0.0122 ± 0.0032 |
| Val_MSE | 0.0135 ± 0.0037 |
| SAE_MSE | 0.9280 ± 0.9801 |
This was a moderate to high amount of polysemanticity, as shown by the MACS (there’s a lot of off-diagonal cosine similarity), and the interference score (many of the important features are entangled). The calibrated MACS score is basically at chance, which means that the network is only slightly better at separating the features than randomly permuting them. Importantly, the network is still performing well on the task, despite its polysemanticity.
I found that, surprisingly, training the network with L2 regularization instead of L1 leads to lower polysemanticity of the features. To allow for better gradients I also changed the activation to GELU. The results are as follows:
| L1 regularization and ReLU activations (baseline) | L2 regularization and GELU activations | Δ % | |
|---|---|---|---|
| MACS: | 0.2954 ± 0.0294 | 0.2627 ± 0.0287 | -11.1% |
| Interference: | 8.7022 ± 1.4480 | 7.1424 ± 1.1916 | -17.9% |
| Train_MSE: | 0.0122 ± 0.0032 | 0.0024 ± 0.0025 | -80.3% |
| Val_MSE: | 0.0135 ± 0.0037 | 0.0051 ± 0.0042 | -62.2% |
| SAE_MSE: | 0.9280 ± 0.9801 | 1.4357 ± 1.2624 | +54.8% |
| L0_Sparsity: | 0.4365 ± 0.0596 | 0.4385 ± 0.0558 | +0.5% |
I observed lower measured feature entanglement in the L2/GELU configuration than in the L1/ReLU configuration (t(16) = 2.1 for MACS and t(16) = 2.8 for Interference, Cohen’s d ~1.5–1.6). These are descriptive results from configurations that differ in both regularization and activation; the original statistical analysis needs rechecking as noted above.
What’s more, the drop in Interference was achieved while maintaining a similar amount of sparsity, as shown by the L0 norm of the activations. This means that it’s not just that the L2-regularized network has more features overall, but that the features themselves are less polysemantic. Moreover, I observed a huge drop in Train and Validation MSEs, meaning that the network is substantially better at solving the task. The increase in the SAE MSE is likely due to there being more features for it to reconstruct.
Result table for all ablations
Full results matrix (averaged over 10 seeds)
| Configuration | Init method | Noise dist | Reg method | Activation | MACS | ΔInterf. | Train MSE | Val MSE | L0 |
|---|---|---|---|---|---|---|---|---|---|
| 1 (baseline) | Random | Bipolar | L1 | ReLU | 0.295 | 8.70 (0%) | 0.0122 | 0.0135 | 0.437 |
| 2 | Orthogonal | Bipolar | L1 | ReLU | 0.284 | 8.11 (-6.8%) | 0.0044 | 0.0078 | 0.439 |
| 3 | Random | PosKurt | L1 | ReLU | 0.306 | 9.11 (+4.7%) | 0.0133 | 0.0152 | 0.416 |
| 4 | Random | Bipolar | L2 | GELU | 0.263 | 7.14 (-17.9%) | 0.0024 | 0.0051 | 0.439 |
| 5 | Orthogonal | Bipolar | L2 | GELU | 0.252 | 6.49 (-25.4%) | 0.0020 | 0.0046 | 0.453 |
| 6 | Orthogonal | PosKurt | L1 | ReLU | 0.285 | 8.30 (-4.6%) | 0.0045 | 0.0076 | 0.452 |
| 7 | Random | PosKurt | L2 | ReLU | 0.284 | 8.35 (-4.1%) | 0.0032 | 0.0058 | 0.491 |
| 8 (max.mit.) | Orthogonal | PosKurt | L2 | GELU | 0.249 | 6.37 (-26.8%) | 0.0021 | 0.0048 | 0.454 |
L2 consistently has lower Interference (6.4 to 8.4) than L1 (8.1 to 9.1)
There seems to be some synergy between Orthogonal and L2+GELU (Config #5, 25.4% lower Interference). Overall, the maximum mitigation is achieved with Config #8, which reduces Interference by 26.8% (t=2.8)
Surprisingly, while positive kurtosis increases polysemanticity in combination with L1 (+4.7%, config #3), it seems to reduce polysemanticity in combination with L2 (-26.8%, config #8)
Analysis: Why does L2 lead to lower polysemanticity?
A Winner-Take-All Dynamic
A good intuition for why L1 Regularization leads to incidental polysemanticity (predicted in Lecomte et al. 2023) is that, in a setup with capacity larger than the number of features, it induces a Winner-Take-All dynamic.
Because the gradient of the L1 norm is discontinuous (∇(λ‖w‖1) = λ sign(w), jumping from − λ to λ at w = 0), it encourages the weights to take extreme values (strongly negative, strongly positive, or exactly 0).
Because the latent space is of dimension 12 and features are 8-dimensional, this leads to a Winner-Take-All dynamic where some neurons (the winners) will encode multiple features, and some others (the losers) will have their weights driven to zero by L1 regularization.
This interpretation seems correct, as the configs with L1 regularization have much higher Interference scores (between 8.1 and 9.1), despite having a similar (even lower) L0 norm (between 0.42 and 0.49). Furthermore, they have a higher train MSE (between 0.012 and 0.013), suggesting that they get stuck in suboptimal local minima. Indeed, the high variance of the SAE_MSE for those configs hints that there are many local minima.
This seems paradoxical, because it's often thought that L1 regularization would lead to more interpretable networks. However, in an overcomplete setting, encouraging sparsity seems to just increase polysemanticity (because there are fewer neurons to encode the same number of concepts).
Why L2 Regularization does not lead to polysemanticity
Because the gradient of the L2 norm is linear (∇(λ‖w‖22) = 2λw), it leads to a more convex-ish optimization landscape and encourages more distributed solutions.
This prevents the emergence of a Winner-Take-All dynamic, as the regularization just encourages all weights to be smaller.
The configs with L2 Regularization have the lowest Interference scores (between 6.4 and 7.1), as well as the lowest train MSEs (between 0.002 and 0.003) and val MSEs (between 0.005 and 0.006), showing that the network optimizes better. Additionally, the L0 norms are higher for those configs (between 0.44 and 0.49), meaning that many neurons have non-zero weights and are dedicated to only one feature. Finally, those configs show low variance across different seeds, meaning they all converge to the same basin of attraction.
It is interesting to note that, in an overcomplete setting, encouraging a distributed solution seems to lead to more interpretable networks.
It is interesting to note the different boundary conditions that arise in orthogonal initialization:
With L2 Weight Regularization (Config #5), orthogonal init acts in synergy (eigenvalues of orthogonal matrices are 1, perfect conditioning for gradient flow), giving great results (-25.4% Interference and -83.6% Train MSE).
With L1 Weight Regularization (Config #6), orthogonal init acts in antagonism (orthogonal structure distributes activations more evenly among the weights, while L1 weight reg tries to make them sparse), causing catastrophic dynamics (SAE_MSE std err of 4.62, larger than the mean of 4.17). The weights also become extremely peaked (Sparsity_W4 increased by 91%, from 0.053 in Random Init + L1 Weight Reg. (Config #1) to 0.101 in Orthogonal Init + L1 Weight Reg. (Config #6)), which destabilizes the SAE reconstruction.
It is also interesting to note the different boundary conditions that arise in positive kurtosis noise:
With L1 Weight Regularization (Config #3), positive kurtosis was harmful (MACS increased by 3.6% and Interference increased by 4.7%). Perhaps, the extreme activations activated multiple features (polysemantic neurons), which then became fixed due to the winner-take-all dynamics of L1.
With L2 Weight Regularization (Config #8), positive kurtosis showed a marginal improvement (Interference decreased by 1.8%, compared to Config #5). Perhaps, the extreme activations helped explore more, and this helped, since L2 had the appropriate smooth regularization, unlike L1, to take advantage of it and not get locked in.
This suggests that simply disrupting the training dynamics is not sufficient; it is also important to have the appropriate smooth regularization to take advantage of it.
Discussion: Implications and Limitations
Implications for the Interpretability Research Field
My results suggest that the assumption that increasing sparsity increases interpretability may be false, at least in the case of overcomplete SAEs. I observed that L1 weight regularization, which increases sparsity, actually increases Interference (+17.9%), implying that it increases polysemanticity, while L2 weight regularization, which is denser, decreases Interference, implying that it decreases polysemanticity.
Of course, I only tested overcomplete SAEs, where latent_dim > feature_dim. It may be the case that for undercomplete SAEs, where latent_dim < feature_dim, L1 weight regularization can be beneficial, as it can force the weights to learn certain features. It is only when there are enough latent dimensions to learn all the features, that the winner-take-all dynamics of L1 weight regularization can be pathological.
How you train your network matters as much as what kind of SAE you train, and most SAE research only focuses on the latter (looking for better decompositions like TopK, JumpReLU, Gated SAEs, etc). In order to fix polysemanticity, we may need to fix how we train our models, because even the best SAEs in the world can only do so much if the model's internal representations are polysemantic.
If you're trying to take action on this post and train your own interpretable model, here's a few tips:
- Try training your network with an L2 weight decay penalty, or maybe an elastic net (L1+L2) penalty where the L2 is stronger.
- If you're using an L2, feel free to use orthogonal initialization, but if you're using an L1, do not use orthogonal initialization.
- Try using activation functions that result in smoother activations (like GELU or SiLU).
- If you're noticing bistability (high variance between SAEs trained with different seeds), this could be a sign of antagonistic training dynamics.
Comparison to Existing Work
Lecomte et al. 2023 predicts that L1 regularization will lead to more incidental polysemanticity due to the winner-take-all dynamics, but as far as I'm aware this is the first time anyone has trained the exact same network twice, once with an L1 and once with an L2, to directly compare and quantify this effect (-17.9% Interference).
Anthropic's work on SAEs (Bricken et al. 2023, Templeton et al. 2024) uses an L1 penalty, which seems at odds with my findings about fighting polysemanticity. However, they're also using more advanced SAE architectures like JumpReLUs and TopK SAEs, which may resolve the L1 regularization paradox in a different way. It'd be interesting to directly compare.
For this work, I only used vanilla LeakyReLU SAEs. There are a number of new SAE techniques (TopK, JumpReLU, Gated, etc) that are able to reconstruct the latent space better at the same level of sparsity. It'd be interesting to see how training a network with an L2 compares to using these SAE methods.
Limitations
This is obviously a toy model, and there are a number of limitations you should be wary of:
The data is synthetic and I specifically designed it to have correlated features. In a real LLM you can't just turn up the correlations knob, but this may be applicable in settings where you'd expect a lot of correlated features (like the words "car" and "road") or if the correlations appear somewhat accidentally during training.
I use a lot of overcapacity in my SAEs (16 latents to represent 8 features, which is 2x overcomplete, and the hidden layer has 12 dimensions which is also larger than the 8 features). This may not be as much of an issue in practice, since when people train SAEs on LLMs they'll usually make them 4-8x more complete.
It's unclear if this applies to LLMs. I would love if someone tested this! Weight decay is not L1 regularization, so that does not establish a direct connection to this toy-model comparison. Whether the proposed mechanism applies to LLMs needs a separate experiment.
Some caveats: I use a somewhat silly way of decaying the importance of features (0.9i), real models probably have a more power law like distribution of feature importance; I don't look at whether the features are qualitatively better, I just use a bunch of proxies for polysemanticity (Interference, MACS); I use a vanilla SAE architecture (LeakyReLU), rather than the current state-of-the-art (TopK / JumpReLU / Gated); the metric Calibrated_MACS is pretty noisy (std 0.14 and mean ~0.5, so it's not much better than chance at telling the difference between features), I could probably improve it by sampling more than 200 permutations, but that's a lot of vectors in a 16D space, so it'd need some thought as to how to do it.
Things I'm not claiming
- That L1 is bad in general, just that it seems to be in this toy overcomplete setting
- That this applies to LLMs, I haven't tried it there and it may or may not
- That we shouldn't use L1 regularization for the SAE, just the network (I do use it for the SAE)
- That L2 regularization is the solution to polysemanticity, just that it may be one piece of the puzzle (although as others have noted, using a modern SAE architecture may be just as important)
Things I am claiming:
- That the L2/GELU configuration had lower measured Interference than the L1/ReLU configuration in this toy model
- That the mechanism (WTA vs distributed) seems to be in line with my theory
- That it's worth looking into this more on real models
Some ideas for further investigations
- One idea is to use PEFT techniques to train an LLM (like Llama) on a dataset similar to the one I used here, and see if the polysemanticity of the features is different when I use L1 vs L2 regularization on the LoRA adapters. This would be a more realistic setting, and would let me investigate the effect of regularization on polysemanticity without worrying about catastrophic forgetting.
- Whether the combination of L2 regularization and a modern SAE architecture (like TopK or JumpReLU) is better than either alone
- Doing a sweep over the value of lambda for L1 and L2 regularization, to see if there's a value that minimizes polysemanticity
- Qualitatively comparing the features learned by the SAE when I use L1 vs L2 regularization, to see if the features are more interpretable in one case vs the other
- Using Anthropic's toy model to see if I can get the SAE to learn the "ground truth" features
Other things I want to explore in the longer term:
- Does this scale to more complex data and larger networks (e.g. billion parameter models)?
- How does this interact with other network properties (learning rate, batch size, etc.)?
- Can I get the benefits of both L1 and L2 regularization by tuning the L1:L2 ratio (similar to elasticnet regularization)?
- How do the effects of L1/L2 regularization interact with the (probability) distribution of feature importances (in real models, feature importances are likely to be power-law distributed)?
Conclusion
In this post I've explored how training conditions affect the polysemanticity of learned representations in toy models. Specifically, I observed 17.9% lower Interference in the L2/GELU configuration than in the L1/ReLU configuration (t = 2.8). This is counterintuitive because in the SAE community there's an implicit assumption that "sparse = more interpretable", especially in the overcomplete setting.
My proposed explanation is that L1 can encourage winner-take-all dynamics, while L2 can encourage a more distributed representation. These experiments motivate that explanation but do not isolate regularization from activation or establish a causal mechanism. Orthogonal weight initialisation synergises with L2 regularization (by keeping the representation distributed), and antagonises L1 regularization (by preventing WTA dynamics).
It's unclear whether these findings generalise to real LLMs, as there are a lot of arbitrary choices I've made in constructing the toy model (e.g. the level of overcapacity, the type of data, etc.). However, I think the mechanistic understanding I've developed here (WTA dynamics due to L1, smooth activations due to L2, interactions between weight initialisation and regularisation, etc.) could be leveraged by future researchers to find better ways of training toy models (and maybe LLMs) to be less polysemantic.
I think a combination of (1) better SAE architectures (e.g. TopK SAEs, JumpReLU activation, etc.) to decompose polysemantic representations, and (2) better model training techniques to minimise polysemanticity in the first place, are both needed to make progress in the field of mechanistic interpretability.
If anyone is interested in exploring this further, all the code, data, etc. used for this post is available here: github.com/stanleyngugi/taming_polysemanticity. Would love to hear any feedback on this post!
Acknowledgments
Thanks to the mechanistic interpretability community for foundational work on SAEs and toy models. This research builds directly on Lecomte et al.'s incidental polysemanticity theory and Anthropic's dictionary learning approaches.
References
Bereska, L., & Gavves, E. (2024). Mechanistic Interpretability for AI Safety — A Review. Transactions on Machine Learning Research.
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., & Olah, C. (2023). Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread. https://transformer-circuits.pub/2023/monosemantic-features
Gao, L., Biderman, S., Gururangan, S., Weber, L., Jennings, J., Dey, B., Sutherland, D. J., Bisk, Y., Schoelkopf, R., Hooker, S., Smith, N. A., & Smith, D. (2024). Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093.
Lecomte, V., Thaman, K., Schaeffer, R., Bashkansky, N., Chow, T., & Koyejo, S. (2023). What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity in Neural Networks. arXiv preprint arXiv:2312.03096.
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., & Nanda, N. (2024). Improving Dictionary Learning with Gated Sparse Autoencoders. arXiv preprint arXiv:2404.16014.
Saxe, A. M., McClelland, J. L., & Ganguli, S. (2014). Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. International Conference on Learning Representations (ICLR).
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, B., Jermyn, A., Carter, S., Olah, C., & Henighan, T. (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Anthropic. https://www.anthropic.com/research/scaling-monosemanticity
Appendix: Technical Details
Hyperparameters
Network training:
- Architecture: [8 → 12 → 12 → 12 → 8]
- Optimizer: Adam (lr=1e-3, betas=(0.9, 0.999))
- L1 lambda: weight_only, manual sum(abs(weights))
- L2 lambda: weight_decay=1e-4
- Noise injection: hidden layer, sigma=0.1
- Epochs: 2000
- Batch size: 64
SAE training:
- Architecture: [12 → 16 → 12]
- Activation: LeakyReLU(0.01)
- L1 lambda: 1e-4 on latent activations
- Optimizer: Adam (lr=1e-3)
- Epochs: 1000
- Reconstruction loss: MSE
Data generation:
- Features: 8, grouped [0-2], [3-4], [5-7]
- Correlations: base_strength ∈ [0.5, 1.0]
- Sparsity: 50% active per sample
- Samples: 1024 train, 256 val
- Target: non-linear combinations with importance decay 0.9^i
If you build on this, I'd love to hear about it — sngugi.research@gmail.com.
@misc{ngugi2025taming,
title = {Taming Incidental Polysemanticity in Toy Models: How Network Training Choices Affect Feature Entanglement},
author = {Ngugi, Stanley},
year = {2025},
howpublished = {\url{https://stanleyngugi.netlify.app/posts/taming_polysemanticity}},
note = {Code: \url{https://github.com/stanleyngugi/taming_polysemanticity}}
}