VT-MUSE Fuses Vision and Touch to Sharpen Robot Manipulation
A new two-stage learning framework blends visual and tactile signals over time, boosting task success by 11 points
A team of researchers has proposed VT-MUSE, short for Multimodal Unified SEquential representation learning, a new framework designed to help robots better combine what they see with what they feel during manipulation tasks, as reported by arXiv (cs.RO).
Most existing visuotactile systems process visual and tactile data separately before merging them, which can miss subtle links between the two senses. These systems also tend to focus only on the current moment, ignoring how contact forces and textures change over time as a robot grasps or manipulates an object. VT-MUSE tackles both issues with a two-stage design.
In the first stage, separate encoders for vision and touch are trained together using cross-modal temporal alignment and masked-view consistency, helping the two modalities learn compatible representations. In the second stage, a conditional variational latent model takes in partially masked sequences of visual frames alongside complete histories of tactile readings. Auxiliary decoders then try to reconstruct the hidden recent visual frames and predict changes in tactile depth, pushing the learned representation to capture both broad visual context and fine-grained contact dynamics.
This combined representation is fed into a lightweight Transformer-based control policy through a gated cross-attention mechanism, allowing the robot's decision-making to draw on both senses as needed.
According to the paper, VT-MUSE outperformed the strongest existing baseline across all tasks in a simulation benchmark by 11 percentage points, and also delivered substantial improvements in real-world manipulation experiments, though specific task names, benchmark details, and exact real-world figures were not included in the available excerpt.
The work, submitted by lead author Congsheng Xu, was posted to arXiv on August 21, 2026, under the categories Robotics (cs.RO) and Computer Vision and Pattern Recognition (cs.CV). It adds to a growing body of research aiming to give robots more human-like integration of sight and touch, a capability seen as important for delicate or contact-rich manipulation tasks such as assembly, grasping deformable objects, and dexterous handling.
talk to a robot
Wrong facts, cryptic artwork, missed news — pick the robot in charge and tell it directly. Litmus verifies the facts; what lands gets a thank-you engraved here.
Sources
This story was written by Robopedia based on the sources below.
Learn more