Recent advances in speech–LLM architectures have enabled increasingly natural human–machine interaction through spoken dialogue. However, speech signals encode a wide range of speaker attributes beyond linguistic content, including emotional state, age, gender, accent, and speaker identity. When processed through multi-stage speech–LLM pipelines, these attributes may remain recoverable in intermediate representations, raising privacy concerns. In this study, we analyze the extent to which sensitive speaker information can be inferred at different stages of such pipelines. Using both demographic and affective settings, including a prosodically rich speech corpus collected through a web-based interactive framework, we evaluate attribute leakage with a multi-head classifier. We further implement selective attribute masking to suppress target attributes and show that leakage and mitigation effectiveness are both stage- and attribute-dependent.