The practice of evaluating TTS models by conducting listening tests to rate the ‘naturalness’ of synthesized speech using Mean Opinion Score (MOS) is under increased scrutiny. In the standard implementation of a MOS-based listening test, the results are decontextualized and do not provide detailed information that researchers can use to determine the source of errors, improve their models or evaluate them with respect to a particular use case. In this work we aim to address these challenges by presenting an approach to listening test evaluation that produces time-aligned error annotations based on the type of error, as well as the severity of the error with respect to the use case. Our approach makes specific considerations to the low-resource TTS context but we also argue for the wider adoption of contextual evaluation in speech synthesis research.