Speaker verification (SV) systems aim to capture speaker-specific traits that are largely independent of linguistic content. However, recent studies have shown that SV performance can still degrade under cross-lingual conditions, suggesting a non-negligible language dependence. Existing analyses mainly focus on data effects while assuming a fixed architecture, leaving open how model design influences cross-lingual behavior.
In this work, we investigate whether architectural choices can mitigate language dependence in SV. Using the Cross-Lingual Transfer Matrix (CLTM), a performance-grounded framework for quantifying pairwise cross-lingual transfer, we systematically compare four representative SV architectures under controlled multilingual conditions and analyze their cross-lingual transfer structures. Our results reveal substantial architectural differences in cross-lingual robustness, providing insights into which modeling strategies are better suited for robust multilingual SV.