Recovering Physically Plausible Human-Object Interactions from Monocular Videos

arXiv cs.CV·Dingbang Huang, Etienne Vouga, Qixing Huang, Georgios Pavlakos

2d ago

·~1 min·6/5/2026·en·0

Quick Answer

This paper shows that The RePHO method reconstructs physically plausible human-object interactions from monocular videos, overcoming common kinematic artifacts.

Quick Take

The RePHO method reconstructs physically plausible human-object interactions from monocular videos, overcoming common kinematic artifacts. By employing a physics-guided framework and reinforcement learning, it significantly improves interaction quality on standard benchmarks, outperforming state-of-the-art methods in physical plausibility metrics.

Key Points

RePHO uses a physics-guided reconstruction framework to enhance human-object interactions.
It employs reinforcement learning to refine kinematic estimates for better accuracy.
An adaptive sampling strategy identifies the most reliable kinematic frames.
The method shows significant improvements in physical plausibility over existing techniques.
Demonstrated effectiveness on two standard human-object interaction benchmarks.

Article Excerpt

From source RSS / original summary

arXiv:2606. 05359v1 Announce Type: new Abstract: In this paper, we propose RePHO, a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often result in physically implausible artifacts such as interpenetration and object floating. To overcome these issues, we introduce a physics-guided reconstruction framework.

We begin with a kinematic estimate and then refine it by training a policy with reinforcement learning (RL). This policy is optimized to reproduce the interaction in a physics simulator. Because kinematic estimates are typically noisy, naive RL training can fail. Therefore, we propose an adaptive sampling strategy with a dual self-updating mechanism that can identify the frames with the most informative and reliable kinematic reconstruction.

Our process progressively improves reconstruction quality and yields physically consistent HOI sequences. We demonstrate our approach on two standard HOI benchmarks and achieve clear improvements in physical plausibility metrics over state-of-the-art methods. Project Page: https://dingbang777. github. io/RePHO/

Reader Mode unavailable (could not extract clean content).

Read on arxiv.org

Want this in your inbox every morning?

Daily brief at your local 8am — bilingual EN/中文, free.

Subscribe — it's free

More from arXiv cs.CV

See more →

arXiv cs.CV·Shahrzad Esmat, Chaunte W. Lacewell, Sameh Gobriel, Nilesh Jain, Ali Jannesari

2d ago

FeaturedOriginal

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

AI Summary

A phase-aware LLM agent optimizes human-object interaction retrieval, outperforming Optuna TPE by 33.3% and VDTuner by 34.2% on the HICO-DET benchmark. This method enhances throughput by 15.3x over UniIR and demonstrates strong transferability across vector database management systems.

#LLM #Agent #Inference #AI Startup

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

Quick Answer

Quick Take

Key Points

Article Excerpt

Want this in your inbox every morning?

More from arXiv cs.CV

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

Biomazon: A Multimodal Dataset for 3D Forest Structure and Biomass Modeling in the Amazon Basin

Optimal Transport Flow Matching by Design

Related in this space

The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane

Aptiv to Deliver Production-Ready Edge AI with Long-Term Support with NVIDIA

TorqueAGI Announces Collaborations with NVIDIA, John Deere, and Dexterity to Advance Physical AI for Enterprise-Grade Robots