Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution

IROS 2026

Hanyi Zhang, Khang Nguyen, Charith Munasinghe, Basu Hela, Tianyu Li, Zihong Luo, Hoan Nguyen, Hans Wernher van de Venn, Yalin Zheng, Ravi Prakash, Tung D. Ta, Anh Nguyen, Baoru Huang
GCA-Bench overview showing instruction levels, scenario understanding, grasping actions, dataset sources, and complex trajectory stages

Abstract

Robust robotic grasping remains a fundamental challenge for complex real-world applications. Existing benchmarks primarily focus on isolated visual grasp pose detection, leaving out tasks that require multi-step reasoning, semantic understanding, and reliable execution. GCA-Bench addresses this gap with challenging complex action scenarios that combine scene-level reasoning and semantic constraints. The benchmark evaluates recent foundation-model and grasping baselines under shared settings, with results showing success rates below 70% on complex grasping scenarios. GCA-Bench also introduces execution-aware metrics and failure analysis to guide more robust, generalizable grasping systems.

GCA-Bench Benchmark

GCA-Bench evaluates complex robotic grasping as a full pipeline rather than a single grasp pose prediction problem. It combines language instructions, semantic and scene understanding, trajectory-level execution, and task-specific constraints across 102 grasping tasks in singulated, cluttered, constrained, and semantic scenarios.

102Grasping tasks
4Scenario groups
3Instruction levels
Sim + RealData sources
GCA-Bench scenario design with singulated, cluttered, constrained, and semantic scenes
Scenario design covers singulated objects, cluttered layouts, constrained spaces, and semantic tasks that require following language instructions.
GCA-Bench task design and evaluation metrics
Tasks evaluate semantic understanding, scene understanding, and staged task execution. Metrics include Detection Success Rate, Grasp Success Rate, Task Success Rate, SPL, and Execution Time.

Data and Environment

The benchmark uses NVIDIA Isaac Lab with a Franka Emika Panda setup and four RGB-D viewpoints: wrist, front, side, and over-the-shoulder. For fine-tuning and analysis, the dataset includes manually demonstrated trajectories with 2000 simulation trajectories and 800 trajectories from real robots.

Complex trajectory collectionDemonstrations include successful and failed trials, multi-stage actions, object state changes, and gripper commands.
Execution-aware graspingTasks require preparatory pushes, safe approach paths, collision avoidance, and post-grasp stabilization.
Collected trajectories showing complex cluttered grasping actions
Complex cluttered scenes require trajectory planning beyond visual grasp detection, including clearing surrounding objects and preparing better grasping poses.
Real-world validation setup across singulated, clutter, constrained, and semantic tasks
Real-world Setup pairs simulation scenarios with UR5 robot validation tasks across the four benchmark categories.
Real-world data collection setup with Franka Panda, SO-ARM 100, UR5, wrist camera, and third-view camera
Real-world Collection uses Franka Panda, SO-ARM 100, and UR5 robots with wrist-mounted and third-view cameras.
Isaac Lab simulation setup with multiple cameras and diverse assets
Simulation environments combine multi-camera observations with diverse objects, shelves, baskets, and other constrained-space assets.

Results

GCA-Bench compares grasp detection pipelines, motion-planning variants, and VLA policies. Detection-based methods remain brittle in complex scenes, while fine-tuned VLA models improve results but still degrade as task complexity increases, especially under semantic constraints.

Comparison of baseline methods on GCA-Bench
Overall baseline comparison across singulated, cluttered, constrained-space, and semantic task categories.
Evaluation results across different GCA-Bench scenarios
Scenario-level results show sharp drops in flat objects, stacked clutter, packed clutter, baskets, and shelves.
Instruction following results across basic, simple constraint, and high-level instructions
Instruction-following results separate basic grasping from simple constraints and high-level semantic goals.
Real-world performance compared with simulation baseline
Real-world validation broadly follows simulation trends, with semantic tasks reaching the highest success rate.

Takeaways

GCA-Bench exposes a persistent gap between perception and reliable execution. Even when targets are detected, stable grasping can fail because of occlusion, thin or slippery objects, constrained motion, missing feedback, and semantic constraints such as preserving balance or avoiding unsafe contact. The benchmark is designed to support future evaluation of closed-loop, language-conditioned, and execution-aware robotic grasping systems.

BibTeX

@inproceedings{zhang2026gcabench,
      title={Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution},
      author={Zhang, Hanyi and Nguyen, Khang and Munasinghe, Charith and Hela, Basu and Li, Tianyu and Luo, Zihong and Nguyen, Hoan and van de Venn, Hans Wernher and Zheng, Yalin and Prakash, Ravi and Ta, Tung D. and Nguyen, Anh and Huang, Baoru},
      year={2026},
      booktitle={IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)}
    }