现代制造工程 ›› 2026, Vol. 549 ›› Issue (6): 51-59.doi: 10.16731/j.cnki.1671-3133.2026.06.006

• 机器人技术 • 上一篇    下一篇

面向识别抓取任务的轻量化LLM人机交互模型*

安延伟1,2, 秦建军1,2,3, 王陆旸1,2,3, 李国通1,2,3, 李昕1,2   

  1. 1 北京建筑大学机电与车辆工程学院,北京 100044;
    2 北京建筑大学“人工智能+”研究院具身智能与机器人技术研究中心,北京 100044;
    3 城市建筑施工智能建造装备北京市重点实验室,北京 100044
  • 收稿日期:2025-12-21 出版日期:2026-06-18 发布日期:2026-07-02
  • 通讯作者: 王陆旸,博士,主要研究方向为车路协同自动驾驶、机器人具身智能和多模态融合感知。E-mail:wangluyang@bucea.edu.cn
  • 作者简介:安延伟,硕士研究生,主要研究方向为机器人控制、具身智能。E-mail:2108020023013@stu.bucea.edu.cn
  • 基金资助:
    *国家地方共建人形机器人创新中心【2025】年度开放基金项目(HRCPUR20251106007);运载工具先进制造与测控技术教育部重点实验室开放课题项目(Z26008);北京建筑大学“人工智能+”研究院“揭榜挂帅”项目(06080925012);北京建筑大学培育项目-种子计划(X25014)

A lightweight LLM-based human-robot interaction model for robotic manipulator recognition and grasping tasks

AN Yanwei1,2, QIN Jianjun1,2,3, WANG Luyang1,2,3, LI Guotong1,2,3, LI Xin1,2   

  1. 1 School of Mechanical-Electronic and Vehicle Engineering,Beijing University of Civil Engineering and Architecture,Beijing 100044,China;
    2 The Embodied Intelligence and Robotics Technology Research Center of the ″Artificial Intelligence+″ Institute, Beijing University of Civil Engineering and Architecture,Beijing 100044,China;
    3 Beijing key Laboratory of Intelligent Construction Equipment for Urban Building Construction, Beijing 100044,China
  • Received:2025-12-21 Online:2026-06-18 Published:2026-07-02

摘要: 面向开放桌面场景下由自然语言驱动的识别与抓取任务,现有方法多依赖超大参数量模型进行决策,虽然能力强,但本地部署常面临显存占用高、推理时延,以及算力与能耗成本高等问题,难以满足机械臂在线闭环控制的实时性与稳定性。同时,语言决策、环境感知与动作控制之间缺乏统一的结构化接口约束,导致闭环执行过程中易出现不可执行动作、交互冗余与失败恢复不稳定。以中等规模大语言模型(Large Language Model,LLM)为决策中枢的函数调用式机械臂人机交互与决策框架,将视觉检测、目标筛选、抓取点估计、运动控制与执行反馈等能力统一封装为具备明确输入/输出协议的机器人统一的结构化接口,并将机械臂结构化状态编码为状态文本,约束模型输出可解析、可校验的动作调用,从而实现在线闭环执行的稳定性与实时性。基于4 B参数量的Qwen3(4 B-Qwen3)模型,构建状态到函数调用的轨迹数据并进行监督微调,提升工具调用行为与任务级决策能力。在Kinova Gen3真实机械臂平台上,以整理桌面杂物为代表性评测任务进行实验,结果表明,微调后的模型在任务成功率、动作效率及函数调用合法性与准确性等指标上显著优于未微调的Qwen3-4 B模型,并在关键指标上接近DeepSeek-R1-671 B模型,验证了中等规模大语言模型在结构化约束与任务特化训练下具备兼顾性能与部署成本的可行性。

关键词: 大语言模型, 机械臂抓取, 人机交互, 结构化机器人接口, 监督微调

Abstract: For natural language driven object recognition and grasping in open tabletop environments,many existing methods rely on extremely large models as a central decision maker. While powerful, but they incur high memory, compute cost,large latency,and expensive deployment,making them unsuitable for real time closed loop robotic control. In addition,the lack of a unified structured interface between language level decisions,perception,and control often leads to non-executable outputs,redundant interaction,and unstable failure recovery. It proposes a function calling human robot interaction framework that uses a medium scale language model as the decision core. Visual detection,target selection,grasp point estimation,motion control,and execution feedback are wrapped as robot functions with explicit input output protocols,and the structured robot state is encoded as a text prefix. This design constrains the model to produce parsable and verifiable action calls,enabling stable closed loop execution with feedback-based recovery. Using the 4 B parameter Qwen3 model, it builds state to function call trajectory data and perform supervised fine tuning to improve tool use and task level decision quality while keeping the system lightweight for deployment. Experiments on a real Kinova Gen3 robotic arm with tabletop clutter rearrangement show that the fine-tuned model achieves higher task success,fewer action steps,and better function call validity and accuracy than the non-fine-tuned baseline,and it approaches a much larger 671 B model on key metrics. These results indicate that structured interfaces and task specific fine tuning can make medium scale models practical for real world interactive grasping tasks.

Key words: Large Language Model (LLM), robotic arm grasping, Human-Robot Interaction (HRI), structured robot inter-faces, Supervised Fine-Tuning (SFT)

中图分类号: 

版权所有 © 《现代制造工程》编辑部 
地址:北京市东城区东四块玉南街28号 邮编:100061 电话:010-67126028 电子信箱:2645173083@qq.com
本系统由北京玛格泰克科技发展有限公司设计开发 技术支持:support@magtech.com.cn