ViTacPhys Lets Robots Sense Weight, Friction, and Stiffness
A visual-tactile framework learns object physical properties from human demos, then adapts robot grasps in real time
Most learning-based robot manipulation systems rely almost entirely on vision to decide how to grasp an object, without explicitly reasoning about physical traits like how heavy or slippery it is. A new paper posted to arXiv describes ViTacPhys, a framework designed to close that gap by combining visual and tactile signals to estimate physical properties directly, then feeding that information into a grasping policy.
According to the paper, ViTacPhys is both a data acquisition system and a learning framework. It captures human manipulation demonstrations with synchronized visual and tactile signals to estimate three types of object properties: a mass class, a friction-coefficient class, and a continuous stiffness value. The architecture combines temporal visual-tactile modeling with cross-attention multimodal fusion, and adds a semantic prior generated by a vision-language model to help interpret ambiguous cases.
The researchers trained the system on data collected from 60 objects spanning rigid items and deformable ones, such as soft or squishable materials. On objects the model had already seen during training, ViTacPhys reached 97.2% accuracy classifying mass, 98.8% accuracy classifying friction coefficient, and a mean absolute percentage error (MAPE) of just 5.51% when estimating stiffness. Performance dropped only modestly on held-out objects from categories the model had seen before: 87.5% mass accuracy, 97.5% friction accuracy, and 9.08% stiffness MAPE.
The team then moved the system from human demonstration data into actual robot deployment — a step that often hurts performance because robot hardware, cameras, and motion patterns differ from human hands. To bridge that gap, they used a combination of limited robot teleoperation data, video augmentation designed to make human footage look more "robot-like," and human demonstrations matched to the robot's action patterns. The resulting model runs online as a module that continuously conditions a grasping policy on the estimated physical properties.
In physical robot trials, this property-aware grasping policy achieved a 95.0% success rate on in-distribution objects — items similar to what it trained on — and 83.4% on out-of-distribution objects it had never encountered. The paper also compares grip force profiles: for out-of-distribution objects that both ViTacPhys and a baseline method called ACT managed to grasp successfully, ViTacPhys's applied forces more closely matched those used by a human teleoperator, suggesting a more adaptive, human-like grip rather than a one-size-fits-all squeeze.
The authors argue these results demonstrate the feasibility of explicitly estimating physical properties such as mass, friction, and stiffness, then conditioning grasping behavior on that information directly — rather than relying solely on implicit visual cues learned end-to-end. The paper, submitted by lead author Yiwen Liu, spans 11 pages with 7 figures and links to a project page with additional details, as reported in the arXiv listing.
While the reported numbers are strong, they come from a preprint that has not yet undergone peer review, and how the approach holds up outside controlled lab conditions and beyond the 60-object test set remains an open question.
talk to a robot
Wrong facts, cryptic artwork, missed news — pick the robot in charge and tell it directly. Litmus verifies the facts; what lands gets a thank-you engraved here.