ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

J-SPAW2: A Japanese Corpus for Speaker Verification and Anti-Spoofing with Challenging Replay and Speech Synthesis Attacks

Sayaka Shiota, Suzuka Horie, Sawato Furubayashi, Shinnosuke Takamichi

Spoofing attacks against automatic speaker verification (ASV) are increasingly serious, as deep learning enables high-fidelity speech synthesis and realistic replay attacks. We construct J-SPAW2, a Japanese corpus for evaluating ASV and antispoofing countermeasures (CM) under challenging conditions. J-SPAW2 extends J-SPAW in two directions. For physical access (PA) attacks, spoofed speech was re-recorded under varied conditions, including playback device, distance, and volume. CM models degrade in low-volume, far-field settings, whereas close-range attacks are more likely to bypass ASV. t-DCF analysis highlights this critical vulnerability. For logical access (LA) attacks, deepfake speech is generated from noisy, non-consensually recorded speech using zero-shot TTS. CM and ASV evaluations confirm that the synthesized speech bypasses existing security systems. J-SPAW2 is designed to broaden benchmark diversity by covering realistic and varied PA and LA threat conditions, thereby supporting research on robust speaker verification and anti-spoofing.