Size Doesn't Matter: Cosine-Scored Sparse Autoencoders
We identify a limitation in how sparse autoencoders detect features through inner product scoring — the approach conflates directional alignment with input magnitude. Since sublayer normalization already removes magnitude information that models don't utilize, we propose replacing inner product with a learned combination of cosine similarity and input magnitude. Our method allows each feature to independently determine how much norm-dependence to employ. The cosine-scored approach never recovers full inner product scoring and identifies features aligned with human-recognizable concepts more frequently, despite comparable reconstruction. Cosine scoring should be the default for dictionary learning on normalized representations.
arXiv