In voice anonymization methods, the preservation of the emotional content is often overlooked, yet it is a key element of vocal communication. The VoicePrivacy Challenge 2024 (VPC2024) addresses this by proposing an anonymization benchmark that uses the unweighted average recall (UAR) to assess anonymization systems' ability to preserve the emotional content. The UAR compares the output labels of a speech emotion recognition (SER) model after anonymization with the ground-truth labels. However, this metric relies on hard label matching and does not account for the classification models' uncertainty. In this paper, we propose three measures to expand the application of the metric to non-labeled datasets by comparing the SER models' soft labels or logits. Validation through a MUSHRA listening experiment confirms that our proposed metric, based on soft-output probabilities, is more robust in real-world settings and is well correlated with subjective test results.