融合多尺度特征金字塔与注意力机制的脑电情绪识别模型

EEG-Based Emotion Recognition via Multi-Scale Feature Pyramid and Attention Mechanism

  • 摘要: 现有基于脑电图(Electroencephalogram,EEG)的情绪识别方法在局部特征表征能力和特征提取效率方面仍存在不足,限制了识别性能的进一步提升。针对上述问题,文章提出一种融合多尺度特征金字塔与注意力机制的情绪识别模型(Multi-Scale Pyramid Attention Transformer,MSPAT)。首先,设计了一种基于多尺度卷积的特征金字塔模块,用于从EEG拓扑图中提取多尺度空间与频域特征;其次,引入多尺度注意力聚合模块,以增强关键通道及空间区域中与情绪相关的判别性特征;最后,利用时序Transformer模块对情绪变化的时序动态信息进行建模。为验证模型的有效性与稳定性,在DEAP与SEED两个公开数据集上,将MSPAT模型与4D-CRNN、STFCGAT、Conformer等具有代表性的基线模型进行对比实验,并进行了混淆矩阵分析和消融实验。对比实验结果表明:在DEAP数据集上,MSPAT模型在Valence、Arousal维度上的分类准确率分别达到97.61%、97.12%,相较于所有的基线模型,分别提升了0.8%~7.8%、0.5%~6.5%;在SEED数据集上的分类准确率达到96.17%,与所有的基线模型相比,提升了0.9%~12.4%。同时,消融实验结果进一步验证了所提出的特征金字塔模块与多尺度注意力模块在增强局部特征表征能力和提升特征提取效率方面的有效性。结果表明,MSPAT模型能够充分挖掘脑电信号的深层关联信息,有效改善特征表达不充分、信息挖掘不全面的现状,具备优异的识别稳定性与跨数据集泛化性能。

     

    Abstract: Existing EEG-based emotion recognition methods still suffer from limited local feature representation capability and inefficient feature extraction, which hinder further improvements in recognition performance. To address these limitations, an emotion recognition model integrating a multi-scale feature pyramid and attention mechanism, termed the Multi-Scale Pyramid Attention Transformer (MSPAT), is proposed. Firstly, a multi-scale convolution-based feature pyramid module is designed to extract multi-scale spatial and frequency-domain features from EEG topological maps. Subsequently, a multi-scale attention aggregation module is introduced to enhance emotion-rela-ted discriminative features in key channels and spatial regions. Finally, a temporal Transformer module is employed to model the temporal dynamics of emotional variations. To evaluate the effectiveness and stability of the proposed model, comparative experiments were conducted on the DEAP and SEED datasets against representative mainstream baseline models, including 4D-CRNN, STFCGAT, and Conformer, together with confusion matrix analysis and ablation studies. Experimental results demonstrate that MSPAT achieves classification accuracy of 97.61% and 97.12% on the Valence and Arousal dimensions of the DEAP dataset, respectively, with improvements of 0.8%-7.8% and 0.5%-6.5% over all baseline models in the two dimensions; it obtains an classification accuracy of 96.17% on the SEED dataset, representing an improvement of 0.9%-12.4% compared with all baseline models. Meanwhile, the results of ablation studies further verify the effectiveness of the proposed multi-scale feature pyramid module and multi-scale attention module in enhancing local feature representation capability and improving feature extraction efficiency. The results indicate that MSPAT can fully mine the deep correlation information of EEG signals, effectively alleviate the problems of insufficient feature expression and incomplete information mining, and exhibit exce-llent recognition stability and cross-dataset generalization performance.

     

/

返回文章
返回