The "I'm Fine" Effect
When the words say one thing and the voice says another.
I designed and evaluated the Alignment Discrepancy Score (ADS), an interpretable measure of the distance between text-derived sentiment and audio-derived prosody in a shared valence-arousal space.
What the result showed: ADS did not sharply separate PHQ groups. The PHQ-elevated group instead showed a tighter distribution with fewer high-ADS outliers, consistent with reduced emotional range. A threshold selected for case detection achieved 80% sensitivity; a balanced threshold yielded 74% sensitivity and 51% specificity.
Scope: Sentiment-ADS outperformed the audio-only baseline but trailed the text-only model. Its value is interpretability: ADS exposes cross-modal variability that text-only prediction cannot show. It is best understood as an exploratory cue, not a replacement for text features or a diagnostic tool.
text-prosody-evidence ↗ turns lessons from the research into a dataset-neutral reference implementation for leakage-aware evaluation, calibration, reliability, and auditable evidence. It is a stricter engineering direction, not a reproduction or validation of the original findings.
Positive words.
Flattened delivery.