Sparsity is recognized as a suitable regularization for deriving interpretable representations from pre-trained embeddings. However, it is unclear how other regularizations, e.g. orthogonality, relate to sparsity, and how they compare wrt. interpretability. We introduce a functionnally-grounded benchmark for measuring dimensional interpretability in speaker representations, comprising unsupervised and supervised metrics, as well as a task-dependent metric as guidance (typicality). By evaluating a wide assortment of variously sparse models trained with three sparse-autoencoder architectures (SPINE, DINE and TopK), we show that sparsity and orthogonality are comparable strategies to train interpretable speaker representation models. Finally, based on its agreement with typicality, we emphasize the relevance of MI-DCI as a supervised disentanglement metric for speaker characterization interpretability, in line with state-of-the-art conclusions on more controlled speech data.