DOI : 10.5281/zenodo.22482097
- Open Access

- Authors : Vanetta P. Silveira, Reva L. Gaonkar, Riya N. Padwalkar, Sanisha Faria, Louella M. Colaco, Supriya Patil, Meghana Pai Kane
- Paper ID : IJERTV15IS080661
- Volume & Issue : Volume 15, Issue 08 , August – 2026
- Published (First Online): 06-09-2026
- ISSN (Online) : 2278-0181
- Publisher Name : IJERT
- License:
This work is licensed under a Creative Commons Attribution 4.0 International License
Multimodal Sentiment Analysis for Mental Health Detection
Vanetta P. Silveira
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Riya N. Padwalkar
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Louella M. Colaco
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Reva L. Gaonkar
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Sanisha Faria
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Supriya Patil
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Meghana Pai Kane
Department of Computer Engineering Padre Conceição College of Engineering Verna, Goa, India
Abstract – Mental health is a critical component of overall well- being, yet conditions such as depression are frequently underdiag- nosed due to the subjective and time-intensive nature of traditional assessment. This paper proposes a multimodal approach for auto- mated depression detection that combines textual and acoustic rep- resentations extracted from patient interviews. A BERT-based en- coder, enriched with auxiliary sentiment features, is used to model linguistic patterns, while a WavLM-based encoder captures acous- tic and prosodic cues from speech. Text and audio embeddings are combined using a learned gated fusion mechanism at the ut- terance level, and a Bidirectional Gated Recurrent Unit further models temporal dependencies at the session level. The system was evaluated on the E-DAIC WOZ dataset, where the text-only model achieved 73.21% accuracy for depression detection, and the fu- sion model obtained the highest precision (72.73%) for identifying individuals with depression. The auxiliary sentiment classication head further attained an accuracy of 87.34% at the utterance level, demonstrating the benet of jointly modelling sentiment and de- pression cues through complementary textual and acoustic modal- ities.
Keywords – multimodal sentiment analysis; depression detection; mental health; BERT; WavLM; multimodal fusion
-
INTRODUCTION
Mental health is a state of well-being that allows one to deal with lifes stresses, to realize their potential, to perform duties normally and to be an asset to the society. Chisholm et al. [1] stated that without mental well-being there is no true physical health, highlighting the strong interdependence between physi- cal and mental well-being. Mental health is often overlooked in comparison to physical health. An imbalance between the two contributes to many physical and mental disorders.
According to a survey conducted by the World Health Orga- nization (WHO) in 2021, 1 in 7 individuals globally are affected by mental disorders, with anxiety and depression being the most prevalent. Anxiety and depression alone affect over 359 and 280 million people respectively, with prolonged untreated depres- sion linked to increased suicide risk [2 Despite the severity, these disorders are often overlooked due to various factors, par- ticularly stigma and lack of awareness [1].
Mental disorders are diagnosed through various approaches, primarily relying on clinical interviews and self-reporting ques- tionnaires. Some of the detection methods include the Patient Health Questionnaire (PHQ-8/9), the Beck Depression Inven- tory (BDI-II), the Hamilton Depression Scale, and the Cen- ter for Epidemiological Studies Depression Scale (CES-D) [5]. Nevertheless, there is a lack of focus on early detection of men-
tal disorders; hence, depression and suicidal tendencies often remain undetected at the initial stages. Affected individuals often show unclear clinical characteristics, making traditional manual diagnosis complex, time-consuming, and potentially bi- ased [5]. The shortage of professionals in the eld of psychiatry makes the situation even more critical [68].
Over the years, researchers have identied w ays to miti- gate mental health issues using Articial Intelligence (AI) tech- niques [911]. Natural Language Processing (NLP) plays a crucial role in sentiment analysis (SA), gaining popularity as a powerful diagnostic tool with the capability of early detection of emotions and disorder symptoms.
Despite signicant progress in SA, most existing systems re- main unimodal, predominantly using text. Consequently, they miss important characteristics that can be obtained using a mul- timodal approach. Recent studies explored systems that uti- lize social media data, consisting of tweets, texts, and images [12, 13]. These approaches proved effective in monitoring men- tal health, particularly among the youth. However, these meth- ods may not be practical, since not all users are active on social media or have access to it. Hence, there is a need for an au- tomated system that integrates multiple modalities to detect the sentiments of individuals.
This paper proposes a multimodal system that integrates tex- tual and audio data to address the mental health issue of de- pression. The proposed system is trained on the Extended Dis- tress Analysis Interview Corpus-Wizard of Oz (E-DAIC WOZ) dataset, using the audio data and their corresponding text. The system aims to detect depression based on audio and text and classies individuals as depressed or non-depressed.
The remainder of this paper is organized as follows. Sec- tion II reviews recent works that motivate this study. Section
-
presents the proposed system and methodology. Sections
-
and V present and discuss the experiments conducted along with the results. Finally, Section VI concludes the paper.
-
-
RELATED WORKS
SA is an NLP application that has been extensively researched and widely applied in various domains. It has signicantly con- tributed to the eld of mental health, where sentiments are used in diagnosing an individuals mental state. However, existing SA systems are largely unimodal, with text being the most com- mon input.
Initially, traditional machine learning (ML) methods such as Logistic Regression (LR), k-Nearest Neighbors (k-NN), and Naive Bayes (NB) formed the foundation of most existing ap- proaches. Alvarez et al. [14] employed six ML methods on Electronic Health Records (EHRs), where Random Forest (RF) classifier gave the best performance (Table I). In parallel, so- cial media emerged as a valuable source for mental health de- tection and analysis [12]. Detection of depression and Post- Traumatic Stress Disorder (PTSD) among Twitter users was ex- plored through sentiment analysis of tweets. Multiple ML algo- rithms were implemented, where Decision Tree (DT) achieved the best performance, while RF performed least effectively [15]. Selva et al. [13] proposed a deep learning framework integrat- ing word2vec and BERT embeddings with a Bi-LSTM classifier and knowledge distillation for depression and anxiety detection
from social media data.
Advanced methods such as Deep Learning (DL), Large Lan- guage Models (LLMs), and transformer-based architectures were widely adopted [1618]. Social media data sourced from Reddit and Twitter was used to predict multiple emotional states such as happiness, anger, surprise, sadness, and fear. A range of techniques, including LR, NB, RF, and Support Vector Ma- chine (SVM), along with DL approaches such as Long Short- Term Memory (LSTM) and Bidirectional Encoder Representa- tion from Transformers (BERT), were employed, achieving a maximum accuracy of 93.28% [19]. Using 30-second speech segments, Zhang et al. [20] trained a deep learning classier to distinguish depressed from non-depressed individuals, with evaluation spanning three corpora: Multimodal Open Dataset for Mental Disorder Analysis (MODMA), DAIC-WOZ, and the Chinese Outpatient Audio Depression Corpus (COAD).
Hassan et al. [21] identied depression using assessment scales such as the Hamilton Depression Rating Scale (HAM- D) and Personal Health Questionnaire (PHQ-8). ML meth- ods, including SVM, Gaussian Mixture Model (GMM), and LR, along with deep learning models such as CNN, RNN, and LSTM, were applied. DL methods excelled in speech emo- tion recognition, and CNNs effectively captured acoustic fea- tures for speech classication. Jadhav et al. [22] utilized six audio and two video features, adopting a voting-based ensem- ble classier comprising LR, Support Vector Classier (SVC), RF, and Gradient Boosting (GB). Textual data was handled via a Mamdani fuzzy approach for textual analysis , achieving up to 99.54% accuracy through late fusion (Table I).
Consequently, researchers recognized that unimodal systems failed to capture important information available in modali- ties such as speech and visual data, leading to the develop- ment of multimodal systems that integrate multiple data sources to improve analysis and performance. Gandhi et al. [23] ex- plored Multimodal Sentiment Analysis (MSA) using speech, visual, and textual data through fusion techniques, and pre- sented various MSA architectures along with interdisciplinary applications and future research directions. Wang et al. [24] conducted sentiment analysis on the Weibo platform, classi- fying the data into positive and negative categories, using the SMP2020 Weibo Emotion Classication Evaluation dataset and weibo_senti_100k, comprising both textual and visual data.
Ikeuchi et al. [17] focused on depression detection from con- versational data through ne-tuning. After converting speech to text via Whisper-large-v3, the authors benchmarked several LLMs, including Llama, Gemma, Mistral, and Qwen for the de- tection task. QLoRA-based ne-tuning was employed to reduce model parameters and optimize memory usage. Among these, Llama 3.1-8B Instruct achieved the strongest results. Liu et al. [25] proposed an AI-based mobile application, MoodEcho, for automatic speech-based depression detection. The system relied on Mel-spectrogram representations to train a CNN aug- mented with Squeeze-and-Excitation (SE) blocks. Iyortsuun et al. [26] introduced an Additive Cross-Modal Attention Network (ACMA), a Bi-LSTM architecture with cross-modal attention that fuses text and audio cues for depression classication, eval- uated on the DAIC-WOZ dataset. It should be noted that accu- racies reported on social-media-derived datasets [13, 19, 22] are
not directly comparable to those obtained on clinical-interview datasets such as DAIC-WOZ/E-DAIC, since social media la- bels are typically self-reported or keyword-derived, whereas E- DAIC WOZ relies on clinically validated PHQ-8 assessments [27], representing a more challenging and realistic detection task. Table I summarises the datasets, methods, and reported performance of the works discussed above.
-
METHODOLOGY
This work presents a multimodal approach for detecting depres- sion through audio signals and their corresponding transcripts, along with extracted sentiments, as depicted in Fig. 1. The au- dio input is decomposed into two modalities, speech signals and textual transcripts. Both modalities undergo preprocessing to remove noise, eliminate irrelevant information, and improve data quality. Subsequently, feature extraction is performed to derive acoustic features from speech signals and contextual em- beddings from transcripts using a pretrained BERT model.
The extracted features from both modalities are combined and fused to leverage complementary information, enabling a more effective understanding of the individuals mental and emotional state. Additionally, sentiment features are extracted from the textual data to capture emotional information, which further helps improve the overall performance of detecting de- pression.
-
Dataset
This work uses the E-DAIC WOZ dataset, a semi-clinical mul- timodal dataset introduced by DeVault et al. [28, 29], designed for automated mental health assessment. It comprises 275 semi-structured human-computer interviews conducted with a virtual assistant, operating in both Wizard-of-Oz (human con- trolled) and fully autonomous interaction modes. The dataset is split into training, development, and test sets while preserv- ing speaker diversity in terms of age, gender, and depression severity based on the eight-item PHQ-8 [27]. The training and development sets contain a mix of human-controlled and AI- driven sessions, whereas the test set consists exclusively of fully automated interactions [28]. Each interview includes synchro- nized audio recordings, textual transcripts, and pre-trained mul- timodal features, with clinically validated labels such as PHQ-8 scores and PTSD severity measures.
-
Data Pre-processing
As the proposed system utilises multiple modalities as input, preprocessing was carried out separately for each modality.
-
Text Pre-processing: Textual transcripts were preprocessed to improve data quality while preserving semantically impor- tant linguistic information. The preprocessing pipeline involved merging the PHQ transcripts with train, development, and test splits using a common participant identier to ensure proper alignment with labels.
-
Sentiment Feature Extraction: To enrich textual represen- tations, sentiment-based features were extracted using a lo- cally stored RoBERTa-based model. Long transcripts were seg- mented into overlapping chunks of up to 510 tokens with a stride of 256 tokens using a sliding window approach. Each chunk was processed independently to obtain sentiment proba-
bilities such as negative, neutral, and positive.
-
Audio Pre-processing: Audio preprocessing aims to en- hance signal quality, remove noise, and isolate the target speaker for effective feature extraction and model training. All audio recordings were resampled to 16 kHz for consistency, and non-relevant introductory and ending segments were removed. Speech regions were identied using WebRTC Voice Activity Detection (VAD) [30]. Speaker separation was performed by extracting embeddings using Resemblyzer [31], followed by K-means clustering [22]. Noise reduction included high-pass ltering, declipping, spike removal, and soft compression, fol- lowed by peak normalization.
-
-
Training
The proposed multimodal framework was trained separately on textual and audio modalities to effectively learn contextual, acoustic, and emotional patterns associated with depression.
-
Text-based Training: The text-based framework was trained using a multi-task learning strategy involving depression classi- cation and auxiliary sentiment prediction. Cross-entropy loss was employed for both tasks, with lower weight assigned to the auxiliary sentiment loss. To address class imbalance, weighted random sampling with softened class weighting was adopted. The model was optimized using the AdamW optimizer with learning rate scheduling, warm-up, gradient clipping, dropout, and layer normalization. Since predictions were generated at the utterance level, session-level depression predictions were obtained using aggregation strategies such as mean probability, condence-weighted mean, and median probability.
-
Audio-based Training: The audio-based framework was trained using acoustic and prosodic representations extracted using a WavLM-based encoder fom the preprocessed speech signals. The extracted audio representations were used to train a depression classication model. Optimization was performed using the AdamW optimizer along with learning rate scheduling and gradient clipping. Utterance-level or segment-level predic- tions were aggregated to obtain session-level predictions.
-
-
Fusion
The proposed system adopts a two-stage fusion strategy that combines textual, sentiment, and acoustic representations to ob- tain a unied session-level depression prediction.
-
Utterance-Level Gated Fusion: At the utterance level, text and audio embeddings are combined using a learned gated fu- sion mechanism. The text branch produces a 128-dimensional embedding by concatenating four BERT pooling strategies, forming a 4 × 768-dimensional vector projected to 128 di- mensions using a linear layer with LayerNorm and GELU. The audio branch produces a 128-dimensional embedding from WavLM frame-level hidden states pooled using an attention pooling layer. The two embeddings are concatenated and passed through a gating network:
g = (Wg[et; ea]+ bg) (1)
where et and ea denote the text and audio embeddings respec- tively, Wg is a learned weight matrix, and is the sigmoid ac- tivation. The gate g is split into text-specic (gt) and audio-
TABLE I. Comparison of existing approaches for mental health detection.
Sr.
No.
Author
Dataset
Method
Accuracy /
F1-Score
Classication
SA
Present
1
Alvarez et al. [14]
EHRs
RF, LR, k-NN, NB, SVM
0.72 /
Binary
Yes
2
Tiwari et al. [15]
Twitter
DT, RF, ML methods
/
Binary
Yes
3
Selva et al. [13]
Reddit, Twitter
word2vec + BERT + Bi-LSTM
98.5% /
Binary
Yes
4
Borah et al. [19]
Reddit, Twitter
LR, NB, RF, SVM, LSTM, BERT
93.28% /
Multi-class
Yes
5
Hassan et al. [21]
DAIC-WOZ
SVM, GMM, CNN, RNN, LSTM
/
Binary
No
6
Zhang et al. [20]
MODMA, DAIC-WOZ, COAD
DL (speech, 30s samples)
/ 0.94, 0.67, 0.93
Binary
No
7
Jadhav et al. [22]
E-DAIC, PHQ-8
Ensemble (LR, SVC, RF, GB) + Mamdani
99.54% /
Binary
No
8
Gandhi et al. [23]
Multiple
Multimodal Fusion (Speech + Visual + Text)
/
Multi-class
Yes
9
Wang et al. [24]
SMP2020,
weibo_senti_100k
Multimodal (Text + Visual)
99.10% /
Binary
Yes
10
Ikeuchi et al. [17]
DAIC-WOZ, PROMPT
Llama, Gemma, Mistral (QLoRA)
/ 0.84
Binary
No
11
Liu et al. [25]
DAIC-WOZ
CNN + SE blocks (Mel-spectrogram)
/ 0.86
Binary
No
12
Iyortsuun et al. [26]
DAIC-WOZ
ACMA + Bi-LSTM (Multimodal)
75.8% /
Binary
No
specic (ga) components. An element-wise interaction term is also computed, and all three are concatenated to form the fused representation:
f = [gt 0 et ; ga 0 ea ; et 0 ea] (2)
where 0 denotes element-wise multiplication. The resulting vector f is linearly projected to 256 dimensions and then passed through a residual block, wherein the MLP consists of two blocks of LayerNorm, GELU activation, and dropout, with an intermediate 256-to-256 linear transformation.
-
Attention Pooling: Since audio features are extracted frame- by-frame from WavLM, an attention pooling mechanism is em- ployed to aggregate frame-level representations into a single ut- terance embedding:
ea = softmax(si) · hi (3)
i
where hi are the WavLM hidden states and si are the learned attention scores.
-
Session-Level Fusion via BiGRU: To capture temporal de- pendency across a clinical interview, a Bidirectional Gated Re- current Unit (BiGRU) session encoder is applied on top of the utterance-level embeddings. For each utterance i, the text em- bedding, audio embedding, and their element-wise interaction are concatenated and projected:
x(i) = proj([e(i); e(i); e(i) 0 e(i)]), proj : R384 R256 (4)
-
Calibration and Threshold Tuning: Temperature scaling is applied as a post-hoc calibration step. A single learnable tem- perature parameter T is optimized on the development set by minimizing cross-entropy loss on the scaled logits z/T . The decision threshold is tuned on the development set by maximis- ing balanced recall, constrained to the range [0.30, 0.70].
-
-
-
RESULTS
The proposed system was evaluated on the E-DAIC WOZ test split (56 sessions, 4,614 utterances) across four depression- detection congurations and two sentiment-classication gran- ularities.
-
Depression Detection
Table II outlines the test-set performance across all four cong- urations. The text-only BERT model achieved the highest over- all accuracy (73.21%) and macro F1-score (0.7214). The audio- only WavLM model reached 67.86% accuracy but showed a pronounced recall imbalance (91.43% Non-Depressed vs. 28.57% Depressed). The Fusion MLP achieved the highest precision for the Depressed class (72.73%), reducing false- positive predictions relative to the text-only model (62.50%). The Session BiGRU obtained the highest Depressed-class re- call (71.43%), at the cost of lower overall accuracy (58.93%). Fig. 2 depicts the precisionrecall trade-off.
TABLE II. Depression Detection Performance on the E-DAIC
t a t a
The projected sequence is fed into the BiGRU:
H = BiGRU(x(1),…, x(n)) (5)
The BiGRU produces hidden states H Rn×128 (64 units per direction). Attention pooling is applied to obtain a session-level
vector, passed through a linear classification head.
WOZ Test Set
Model
Acc.
Prec.
Rec.
Macro F1
Text-only (BERT)
0.7321
0.7188
0.7286
0.7214
Audio-only (WavLM)
0.6786
0.6738
0.6000
0.5902
Fusion MLP (BERT+WavLM)
0.7143
0.7192
0.6476
0.6500
Session BiGRU
0.5893
0.6094
0.6143
0.5881
Input
Raw Audio
Audio
Corresponding Transcripts
Text Pre-processing
-
Space normalization
-
Filler removal (uh, um, hmm)
-
Interviewer artifact removal
-
Sentence boundary x/p>
-
Short utterance removal (<3 words)
Acoustic Feature Extraction
-
MFCC
-
Prosodic features
-
Pitch variance
Text Feature Extraction
-
BERT chunking (512 tok)
-
50% overlap stride
-
CLS embeddings
-
Weighted by chunk count
Sentiment Extraction
-
RoBERTa-based model
-
Sliding window (510 tok)
-
p_neg, p_pos scores
-
Sentiment strength score
-
BertWithSentiment model
-
Linear(768+3, 2) classier
-
Focal Loss + class weighting
-
Dev threshold tuning
Depression Detection + Sentiment Prediction
Feature Fusion
(BERT CLS + p_neg + p_pos + Sentiment Strength)
Audio Pre-processing
-
Trim intro/outro (70s / 40s)
-
VAD segmentation (WebRTC)
-
Speaker separation (K-means)
-
High-pass lter (60 Hz)
-
Declip + spike removal
-
Soft compression
-
Peak normalization
Fig. 1. Proposed multimodal pipeline for depression detection using the E-DAIC-WOZ dataset.
Score
1
0.5
0
0.71
0.63
0.67
0.73
0.71
0.47
0.38
0.29
Text-only Audio-only Fusion MLP Session BiGRU
BERT text encoder captures linguistic depression markers ef- fectively but on its own is prone to false positives, while the WavLM audio encoder, despite lower standalone accuracy, supplies acoustic cues that sharpen precision when combined through the gated fusion mechanism. This is reected in the Fusion MLPs improved Depressed-class precision (72.73% vs. 62.50% for text-only), which is clinically relevant since it re- duces unnecessary referrals. Compared to prior DAIC-WOZ fusion models such as the ACMA framework by Iyortsuun et al. [26] (75.8% accuracy), the proposed Fusion MLP achieves a comparable 71.43% accuracy while additionally performing
Precision (Dep.)
Recall (Dep.)
Fig. 2. Precisionrecall trade-off for the Depressed class across the four model congurations.
-
-
Sentiment Classication
The auxiliary sentiment head, trained jointly with the text branch, achieved 87.34% accuracy (macro F1 = 0.8734) at the utterance level across 4,614 test utterances. Aggregating pre- dictions to the session level using a tuned threshold of 0.5206 yielded 85.71% accuracy (macro F1 = 0.8564) across the 56 test sessions.
joint sentiment classication, a capability absent from existing DAIC-WOZ depression-detection systems.
The gap between utterance-level and session-level perfor- mance suggests that temporal dynamics are harder to learn from the small labelled set (21 Depressed, 35 Non-Depressed ses- sions in the test split), though the BiGRUs higher Depression- class recall makes it a useful complementary high-sensitivity signal alongside the higher-precision Fusion MLP.
-
-
DISCUSSION
The results indicate that text and audio modalities contribute complementary information rather than redundant signal. The
The consistently strong sentiment classication performance at both utterance (87.34%) and session (85.71%) levels con- rms that the shared BERT encoder generalises well across the two related tasks, supporting the multi-task design choice.
-
CONCLUSION
This paper presented a multimodal approach for depression de- tection that combines textual and acoustic representations ex- tracted from patient interviews in the E-DAIC WOZ dataset. Four congurations were evaluated: a text-only model, an audio-only model, an utterance-level gated fusion model, and a session-level BiGRU model. The text-only model achieved the highest overall accuracy (73.21%), while the gated fu- sion model achieved the highest precision for the Depressed class (72.73%). The session-level BiGRU achieved the high- est Depressed-class recall (71.43%), making it a useful high- sensitivity signal for agging at-risk patients. The auxiliary sentiment classication head achieved strong performance at both the utterance level (87.34% accuracy) and the session level (85.71% accuracy). These ndings show that the proposed sys- tem achieves performance comparable to existing DAIC-WOZ- based approaches while uniquely combining depression detec- tion with sentiment classication, offering added clinical in- terpretability beyond single-task depression screening systems. Future work will explore incorporating additional modalities such as facial expressions, expanding the model to detect re- lated conditions such as anxiety and stress, and validating the system on larger and more diverse clinical datasets.
REFERENCES
-
K. Srivastava, K. Chatterjee, and P. S. Bhat, Mental health awareness: The indian scenario, Industrial Psychi- atry Journal, vol. 25, no. 2, pp. 131134, 2016, editorial.
-
World Health Organization, Mental disorders, Sep. 2025, accessed: 2026-03-26. [Online]. Avail- able: https://www.who.int/news-room/fact-sheets/detail/ mental-disorders
-
, Depressive disorder (depression), Aug. 2025, accessed: 2026-03-26. [Online]. Available: https:
//www.who.int/news-room/fact-sheets/detail/depression
-
Institute for Health Metrics and Evaluation (IHME), Global burden of disease study 2021 results, 2024, accessed: 2026-03-26. [Online]. Available: https:
//vizhub.healthdata.org/gbd-results/
-
N. Cummins, S. Scherer, J. Krajewski, S. Schnieder,
J. Epps, and T. F. Quatieri, A review of depression and suicide risk assessment using speech analysis, Speech communication, vol. 71, pp. 1049, 2015.
-
Y. Li, S. Kumbale, Y. Chen, T. Surana, E. S. Chng, and
-
Guan, Automated depression detection from text and audio: A systematic review, IEEE Journal of Biomedical and Health Informatics, 2025.
-
-
M. Kowalewski, M. Stroinski, K. Kwarciak, V. Laptiev, and D. Hemmerling, End-to-end multimodal system for depression detection from online recordings, in 2023 45th Annual International Conference of the IEEE Engi- neering in Medicine & Biology Society (EMBC). IEEE, 2023, pp. 14.
-
S. Teng, S. Chai, J. Liu, T. Tateyama, L. Lin, and Y.-W. Chen, Multi-modal and multi-task depression detection with sentiment assistance, in Proceedings of the 2024 IEEE International Conference on Consumer Electronics (ICCE), 2024, pp. 15.
-
C. Madhumitha, Prediction of mental health (depression) using data science and machine learning techniques, B. Eng. thesis, Dept. Comput. Sci. Eng., 2022.
-
M. Maulyanda and S. A. Nazhifah, Sentiment analysis of mental health using support vector machine (svm) with fastapi implementation, Brilliance: Research of Articial Intelligence, vol. 5, no. 1, pp. 568575, 2025.
-
P. Nedungadi, G. Veena, K.-Y. Tang, R. R. Menon, and
R. Raman, Ai techniques and applications for online so- cial networks and media: Insights from bertopic model- ing, Ieee Access, 025.
-
F. Arias, M. Z. Nunez, A. Guerra-Adames, N. Tejedor- Flores, and M. Vargas-Lombardo, Sentiment analysis of public social media as a tool for health-related topics, Ieee Access, vol. 10, pp. 74 85074 872, 2022.
-
G. Selva Mary, J. Blesswin, M. Venkatesan, S. Vairagar,
S. Adagale, C. Shravage, and J. Barpute, Original re- search article enhancing conversational sentimental analy- sis for psychological depression prediction with bi-lstm, Journal of Autonomous Intelligence, vol. 7, no. 1, 2023.
-
E. Alvarez-Mellado, E. Holderness, N. Miller, F. Dhang,
P. Cawkwell, K. Bolton, J. Pustejovsky, and M.-H. Hall, Assessing the efcacy of clinical sentiment analysis and topic extraction in psychiatric readmission risk predic- tion, in Proceedings of the Tenth International Workshop on Health Text Mining and Information Analysis (LOUHI 2019), 2019, pp. 8186.
-
P. K. Tiwari, M. Sharma, P. Garg, T. Jain, V. K. Verma, and A. Hussain, A study on sentiment analysis of mental illness using machine learning techniques, in IOP confer- ence series: materials science and engineering, vol. 1099, no. 1. IOP Publishing, 2021, p. 012043.
-
R. V. Sharan, C. Mascolo, and B. W. Schuller, Emotion recognition from speech signals by mel-spectrogram and a cnn-rnn, in 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, 2024, pp. 14.
-
K. Ikeuchi, T. Kishimoto, F. Nakai, T. Horigome, M. Ki- tazawa, and T. Ohtsuki, Efcient and effective ne- tuning method for depression detection from conversa- tion, in 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, 2025, pp. 15.
-
J. He and H. Hu, Mf-bert: Multimodal fusion in pre- trained bert for sentiment analysis, IEEE Signal Process- ing Letters, vol. 29, pp. 454458, 2021.
-
T. Borah and S. Ganesh Kumar, Application of nlp and machine learning for mental health improvement, in International Conference on Innovative Computing and Communications: Proceedings of ICICC 2022, Volume 3. Springer, 2022, pp. 219228.
-
S. Zhang, Z. Zheng, D. He, H. Zhu, Y. Gu, Y. Hu, X. Lv,
T. Zhu, and W. Zhang, Speech-based depression de- tection: Enhancing emotional support via clinical data, IEEE Transactions on Affective Computing, 2025.
-
A. Hassan and S. Bernadin, A comprehensive analysis of speech depression recognition systems, in SoutheastCon 2024. IEEE, 2024, pp. 15091518.
-
G. D. Jadhav, S. D. Babar, and P. N. Mahalle, Ensemble- based multimodal analysis for depression detection, En- gineered Science, vol. 35, p. 1604, 2025.
-
A. Gandhi, K. Adhvaryu, S. Poria, E. Cambria, and
A. Hussain, Multimodal sentiment analysis: A system- atic review of history, datasets, multimodal fusion meth- ods, applications, challenges and future directions, Infor- mation Fusion, vol. 91, pp. 424444, 2023.
-
C. Wang, J. Konpang, A. Sirikham, and S. Tian, Enhanc- ing weibo sentiment analysis with multi-modal learning: Integrating text and synthesized images with contrastive learning, IEEE Access, 2025.
-
L. Liu, F. Tydeman, W. Xie, and Y. Wang, Develop- ment of an ai-based mobile app for automatic depression screening using speech in english and chinese, in 2025 47th Annual International Conference of the IEEE Engi- neering in Medicine and Biology Society (EMBC). IEEE, 2025, pp. 15.
-
N. K. Iyortsuun, S.-H. Kim, H.-J. Yang, S.-W. Kim, and
M. Jhon, Additive cross-modal attention network (acma) for depression detection based on audio and textual fea- tures, IEEE Access, vol. 12, pp. 20 47920 489, 2024.
-
S. S. Dhingra, K. Kroenke, M. M. Zack, T. W. Strine, and
L. S. Balluz, Phq-8 days: a measurement option for dsm- 5 major depressive disorder (mdd) severity, Population health metrics, vol. 9, no. 1, p. 11, 2011.
-
J. Gratch, R. Artstein, G. M. Lucas, G. Stratou, S. Scherer,
A. Nazarian, R. Wood, J. Boberg, D. DeVault, S. Marsella et al., The distress analysis interview corpus of human and computer interviews. in Lrec, vol. 14. Reykjavik, 2014, pp. 31233128.
-
F. Ringeval, B. Schuller, M. Valstar, N. Cummins,
R. Cowie, L. Tavabi, M. Schmitt, S. Alisamir, S. Amiripar- ian, E.-M. Messner et al., Avec 2019 workshop and chal- lenge: state-of-mind, detecting depression with ai, and cross-cultural affect recognition, in Proceedings of the 9th International on Audio/visual Emotion Challenge and Workshop, 2019, pp. 312.
-
J. Wiseman and I. Y. Bondarenko, Python interface to the webrtc voice activity detector, URL: https://github. com/wiseman/py-webrtcvad ( : 15.04. 2025), 2016.
-
C. Jemine, Resemblyzer: Voice encoder for speaker ver- ication, https://github.com/resemble-ai/Resemblyzer, 2019, accessed: 2026.
