现代制造工程 ›› 2026, Vol. 551 ›› Issue (8): 122-131.doi: 10.16731/j.cnki.1671-3133.2026.08.017

• 仪器仪表/检测/监控 • 上一篇    下一篇

融合影像与时域特征的多维注意力机制增强的跨模态sEMG手势识别方法*

黄斌浩1, 吕健1, 强利刚2   

  1. 1 贵州大学现代制造技术教育部重点实验室,贵阳 550025;
    2 贵州航天控制技术有限公司,贵阳 116026
  • 收稿日期:2025-04-08 出版日期:2026-08-18 发布日期:2026-09-01
  • 通讯作者: 吕健,副教授,硕士生导师,博士,主要研究方向为人机交互、智能设计。 E-mail:jlv@gzu.edu.cn
  • 作者简介:黄斌浩,硕士研究生,主要研究方向为外骨骼机器人设计、智能交互。
  • 基金资助:
    *贵州省科技厅科学支撑计划项目(黔科合支撑[2022]一般197)

Cross-modal sEMG gesture recognition method enhanced by multi-dimensional attention mechanism integrating image and temporal features

HUANG Binhao1, LÜ Jian1, QIANG Ligang2   

  1. 1 Key Laboratory of Advanced Manufacturing Technology of the Ministry of Education,Guizhou University,Guiyang 550025,China;
    2 Guizhou Aerospace Control Technology Co.,Ltd.,Guiyang 116026,China
  • Received:2025-04-08 Online:2026-08-18 Published:2026-09-01

摘要: 现有表面肌电(surface Electromyography,sEMG)信号手势识别方法在时空特征提取和多模态信息融合方面存在局限,传统二维卷积神经网络(2D Convolutional Neural Network,2D CNN)和长短期记忆网络(Long Short-Term Memory,LSTM)难以精准分类复杂手势。为解决这一问题,提出了一种基于sEMG灰度图和多维注意力机制的跨模态手势识别方法。通过将sEMG信号转化为灰度特征图和时域特征表格,并结合3D ResNet-18网络进行时空特征提取,同时引入TabAttention模块和CSRA残差注意力机制,以增强特征融合和空间细节捕捉。实验结果表明,所提方法在6种手势识别中达到98.13 %的准确率,优于2D CNN、LSTM、AlexNet和VGGNet模型,为sEMG手势识别提供了更高效的解决方案。

关键词: 表面肌电, 3D卷积神经网络, 时空特征提取, 跨模态信息融合, 手势分类

Abstract: Existing surface Electromyography (sEMG) gesture recognition methods exhibit certain limitations in spatiotemporal feature extraction and multimodal information fusion,with traditional 2D Convolutional Neural Network (3D CNN) and Long Short-Term Memory (LSTM) models struggling to accurately distinguish complex gestures. To address this issue,it proposes a cross-modal sEMG gesture recognition method based on sEMG grayscale images and a multi-dimensional attention mechanism. The proposed method transforms sEMG signals into grayscale feature maps and tabular time-domain features,utilizing a 3D ResNet-18 model for spatiotemporal feature extraction. Additionally,TabAttention and CSRA residual attention mechanisms are introduced to enhance feature fusion and optimize spatial detail extraction. Experimental results demonstrate that the proposed approach achieves an accuracy of 98.13 % in recognizing six types of gestures,significantly outperforming 2D CNN,LSTM,AlexNet,and VGGNet,thus providing a more superior solution for sEMG gesture recognition.

Key words: surface Electromyography (sEMG), 3D Convolutional Neural Network (3D CNN), spatiotemporal feature extraction, cross-modal information fusion, gesture classification

中图分类号: 

版权所有 © 《现代制造工程》编辑部 
地址:北京市东城区东四块玉南街28号 邮编:100061 电话:010-67126028 电子信箱:2645173083@qq.com
本系统由北京玛格泰克科技发展有限公司设计开发 技术支持:support@magtech.com.cn