ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Large-Kernel 1D CNN for Raw Waveform Spoofing Countermeasures

Guy Perets, Yehuda Ben-Shimol, Itshak Lapidot

This paper proposes a countermeasure (CM) for spoofing-robust automatic speaker verification (SASV) that applies a large context filter to raw audio waveforms. The CM uses a 1D ResNet34 architecture with a first convolution layer with a large kernel. The model eschews fixed time-frequency representations and enables direct learning of discriminatory features from waveform signals. An evaluation of several kernel lengths determined that a 160-sample (10ms) kernel size is the most effective choice for the first convolution layer of the proposed architecture. This notable kernel size captures a broader temporal context, which is useful for distinguishing bonafide speech from spoofed speech. When tested with the ASVspoof2019 LA evaluation set, the proposed CM attained an EER of 1.37% and a min t-DCF of 0.0429, whereas the integrated ECAPA-TDNN ASV achieved a min t-DCF of 0.0478. These results underscore the effectiveness of large-kernel temporal context modeling for anti-spoofing.