QwenGrasp: A Usage of Large Vision-Language Model for Target-Oriented Grasping

Chen, Xinyu; Yang, Jian; He, Zonghan; Yang, Haobin; Zhao, Qi; Shi, Yuhui

Computer Science > Robotics

arXiv:2309.16426 (cs)

[Submitted on 28 Sep 2023 (v1), last revised 25 Dec 2023 (this version, v3)]

Title:QwenGrasp: A Usage of Large Vision-Language Model for Target-Oriented Grasping

Authors:Xinyu Chen, Jian Yang, Zonghan He, Haobin Yang, Qi Zhao, Yuhui Shi

View PDF HTML (experimental)

Abstract:Target-oriented grasping in unstructured scenes with language control is essential for intelligent robot arm grasping. The ability for the robot arm to understand the human language and execute corresponding grasping actions is a pivotal challenge. In this paper, we propose a combination model called QwenGrasp which combines a large vision-language model with a 6-DoF grasp neural network. QwenGrasp is able to conduct a 6-DoF grasping task on the target object with textual language instruction. We design a complete experiment with six-dimension instructions to test the QwenGrasp when facing with different cases. The results show that QwenGrasp has a superior ability to comprehend the human intention. Even in the face of vague instructions with descriptive words or instructions with direction information, the target object can be grasped accurately. When QwenGrasp accepts the instruction which is not feasible or not relevant to the grasping task, our approach has the ability to suspend the task execution and provide a proper feedback to humans, improving the safety. In conclusion, with the great power of large vision-language model, QwenGrasp can be applied in the open language environment to conduct the target-oriented grasping task with freely input instructions.

Subjects:	Robotics (cs.RO)
Cite as:	arXiv:2309.16426 [cs.RO]
	(or arXiv:2309.16426v3 [cs.RO] for this version)
	https://doi.org/10.48550/arXiv.2309.16426

Submission history

From: Xinyu Chen [view email]
[v1] Thu, 28 Sep 2023 13:23:23 UTC (12,072 KB)
[v2] Sun, 8 Oct 2023 05:54:47 UTC (12,056 KB)
[v3] Mon, 25 Dec 2023 08:59:37 UTC (12,733 KB)

Computer Science > Robotics

Title:QwenGrasp: A Usage of Large Vision-Language Model for Target-Oriented Grasping

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Robotics

Title:QwenGrasp: A Usage of Large Vision-Language Model for Target-Oriented Grasping

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators