StableGrasp: Reconstructing Physically Stable Human Hand Grasps from Single Images
Quick Answer
StableGrasp introduces a differentiable simulation-based optimization framework for reconstructing stable human hand grasps from single RGB images.
Quick Take
By separating visual hand pose from control targets, it significantly enhances the stability of reconstructed grasps, outperforming existing methods in physical simulation consistency and geometric plausibility.
Key Points
- StableGrasp optimizes hand geometry and control to minimize grasp kinetic energy.
- The method ensures visual consistency while enhancing physical stability of grasps.
- Experiments demonstrate superior stability compared to traditional hand-control strategies.
- Joint optimization leads to more plausible and stable grasp reconstructions.
- The framework benefits visual-only grasp reconstruction pipelines significantly.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reconstructing a physically stable human grasp from a single RGB image is challenging because physically modeling grasps is itself difficult, and the problem requires estimating not only a visually constrained hand pose but also a control target that stabilizes the grasp. Existing methods either model only visual hand geometry without considering physics, or rely on less plausible physical modeling, which limits the physical validity of the resulting grasps. In this paper, we present StableGrasp, a differentiable simulation-based optimization framework that explicitly separates the visual hand pose from the control target that determines the grasping forces. Our method jointly optimizes hand geometry and control by minimizing the kinetic energy of the grasp in a differentiable simulator, while regularizing the hand geometry to preserve visual consistency and geometric plausibility. The reconstructed grasps are substantially more stable under rigorous physical simulation, while remaining visually consistent with the input images and geometrically plausible. Experiments show that our approach produces far more stable grasps than alternative hand-control strategies, benefiting visual-only grasp reconstruction pipelines by turning their outputs into physically stable grasps.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO) |
| Cite as: | arXiv:2610.09195 [cs.CV] |
| (or arXiv:2610.09195v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09195 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Han Jiang [view email]
[v1]
Tue, 6 Oct 2026 22:45:51 UTC (42,340 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.