In recent years, Deep Neural Network (DNN) architectures have successfully been applied to Text-Independent Speaker Verification (TI-SV), bringing impressive performance improvements. Despite its broad field of application and strong commercial interest, Text-Dependent Speaker Verification (TD-SV) has seen limited research effort. This paper presents our attempt to take advantage of DNN architectures to improve TD-SV systems, through linguistic verification (LV) followed by speaker verification (SV). For LV, we introduce an original attention mechanism based on activation sequence alignment. For SV, a standard ResNet trained on speaker classification task is used. Evaluation of the Common Voice and DeepMine datasets demonstrates the effectiveness of the proposed approach.