现代制造工程 ›› 2026, Vol. 551 ›› Issue (8): 59-67.doi: 10.16731/j.cnki.1671-3133.2026.08.008

• 机器人技术 • 上一篇    下一篇

基于积分强化学习的轮腿机器人姿态平衡控制*

鲁金亮1,2, 皮明1,2, 杨涛1,2, 张良2,3, 曾国鑫1,2   

  1. 1 西南科技大学信息与控制工程学院,绵阳 621000;
    2 特殊环境机器人技术四川重点实验室,绵阳 621000;
    3 绵阳师范学院机电工程学院,绵阳 621000
  • 收稿日期:2025-06-25 出版日期:2026-08-18 发布日期:2026-09-01
  • 通讯作者: 皮明,博士,特聘副教授,主要研究方向为智能机器人控制。 E-mail:756408918@qq.com
  • 作者简介:鲁金亮,硕士研究生,主要研究方向为机器人控制。杨涛,博士,教授,主要研究方向为机电系统仿真。张良,硕士,副教授,主要研究方向为机电系统建模与仿真。曾国鑫,硕士研究生,主要研究方向为机器人控制。E-mail:1912693105@qq.com
  • 基金资助:
    *四川省重点研发计划项目(2024YFFK0039);西南科技大学博士基金项目(21zx7142)

Posture balance control of wheel-legged robots based on integral reinforcement learning

LU Jinliang1,2, PI Ming1,2, YANG Tao1,2, ZHANG Liang2,3, ZENG Guoxin1,2   

  1. 1 School of Information and Control Engineering,Southwest University of Science and Technology,Mianyang 621000,China;
    2 Key Laboratory of Sichuan Province for Robot Technology Used for Special Environment,Mianyang 621000,China;
    3 School of Mechanical and Electrical Engineering,Mianyang Normal University,Mianyang 621000,China
  • Received:2025-06-25 Online:2026-08-18 Published:2026-09-01

摘要: 针对轮腿机器人运动过程中的姿态平衡,对运动学解耦后线性化的状态方程引入零力矩点(Zero Moment Point,ZMP)稳定性判据结合积分强化学习(Integral Reinforcement Learning,IRL)的在线最优控制实现机器人直立姿态平衡与行走稳定性。每隔固定步长对控制结果积分,并通过最小二乘法拟合代价函数的Riccati矩阵P,优化在线策略。实验结果表明,在收敛速度和稳态误差等方面IRL控制均优于固定最优控制。由于线性化后的模型并未涵盖非线性因素,提出一种串级模糊PID控制,ADAMS与MATLAB联合仿真表明,串级模糊PID控制姿态角以及速度(稳态误差减少38.46 %)方面显著优于PID。

关键词: 轮腿机器人, 零力矩点, 最优控制, 积分强化学习, 串级模糊PID控制

Abstract: To address posture balance control of wheel-legged robots in motion, the Zero Moment Point (ZMP) stability criterion and online optimal control based on Integral Reinforcement Learning (IRL) were integrated with the linearized state equation obtained after kinematic decoupling. This integrated framework achieved upright posture balance and stable walking. Control inputs were integrated at fixed time intervals,and the Riccati matrix P of the cost function was fitted via least squares to optimize the online control policy. Experimental results demonstrate that IRL control outperforms fixed-gain optimal control in terms of convergence speed and steady-state error. To address nonlinear factors unaccounted for in the linearized model,a cascaded fuzzy PID controller was proposed. ADAMS and MATLAB co-simulation results confirm that the cascaded fuzzy PID controller significantly outperforms conventional PID control in regulating posture angle and velocity tracking, achieving a 38.46 % reduction in steady-state error.

Key words: wheel-legged robot, Zero-Moment Point (ZMP), optimal control, Integral Reinforcement Learning (IRL), cascaded fuzzy PID control

中图分类号: 

版权所有 © 《现代制造工程》编辑部 
地址:北京市东城区东四块玉南街28号 邮编:100061 电话:010-67126028 电子信箱:2645173083@qq.com
本系统由北京玛格泰克科技发展有限公司设计开发 技术支持:support@magtech.com.cn