Let Robots "Understand by Touch" the World, BeingBeyond Releases New Generation Model Being-H0.8

Edited by Taylor From Gasgoo

Gasgoo Munich-On July 28, BeingBeyond, a developer of general embodied foundation models, officially released Being-H0.8, a new generation of implicit tactile world-action models.

The model marks the first systematic introduction of tactile modalities into large-scale pre-training. It achieves a unified representation of vision, touch, action, and the future state of the physical world, signaling a shift for embodied large models from "visual observation" to "multimodal interactive understanding."

Closing the Tactile Gap: Letting Robots "Feel" the Physical World

In previous embodied intelligence solutions, vision was the primary perceptual input. Yet vision alone often fails to capture the full spectrum of information required for real-world manipulation. Details such as whether contact has been made, if an object is slipping, or if sufficient force is applied—factors that determine success or failure—often cannot be reliably captured through visual data alone.

Consequently, traditional solutions have struggled to maintain stability in contact-intensive tasks like plugging, twisting, and precision assembly.

Being-H0.8’s core breakthrough lies in fully integrating tactile information into the implicit world model architecture during pre-training for the first time. By mapping vision, touch, action, and future state changes into the same feature space, the model simultaneously understands "what is seen, what is done, what is touched, and how the world will change as a result," creating a complete cognitive and predictive capability for physical interaction.

image.png

Image Credit: BeingBeyond

To accommodate high-frequency tactile feedback, Being-H0.8 also features a "slow-fast action expert" mechanism. The system first calculates and caches a "world-action context" at the start of each action time domain. Then, at a few anchor points within that domain, it injects the latest observed proprioceptive state and tactile feedback to dynamically generate or correct the action segment currently being executed.

Addressing the high cost of existing tactile sensors and data acquisition systems—which rely heavily on specialized hardware, tactile gloves, and controlled lab environments—BeingBeyond has independently developed TactoHand, a dense tactile pseudo-labeling system for large-scale unlabeled human video.

This system does not require data collectors to wear tactile gloves or additional sensors. Instead, it infers contact probabilities and continuous proximity at dense spatial points during hand-object interactions from ordinary human videos, supplementing massive amounts of video with tactile supervision at near-zero marginal cost.

It is understood that, leveraging TactoHand, BeingBeyond has extended tactile information to over 500,000 hours of human video for the first time, using it to pre-train its embodied foundation model.

Furthermore, within Being-H0.8, BeingBeyond introduced TopoHand, a second-generation unified action space that provides a unified kinematic interface for human hands, dexterous hands, and parallel grippers. TopoHand employs a fixed-topology "spiral hand" representation composed of 20 semantic keypoints and 20 canonical joint variables. It establishes a unified coordinate system centered on the wrist: the x-axis points toward the projection of the middle fingertip on the palm plane, the z-axis aligns with the palm's normal direction, and the y-axis completes the right-handed coordinate system.

Under this representation, the specific joint naming, mechanical structure, and mesh topology unique to different embodiments are isolated from the policy interface. The model no longer relies directly on the raw joint definitions of a specific robot or human hand model. Instead, it learns manipulation patterns in a unified semantic topology and kinematic space. This enables the efficient transfer of manipulation priors from large-scale human videos to various robot embodiments, significantly boosting the efficiency and scalability of cross-embodiment pre-training.

From Scale to Quality: The Triple Jump of Embodied Foundation Models

Over the past year, BeingBeyond has completed multiple iterations centered on large-scale human video pre-training, a trajectory that reflects the broader data-driven development logic within the embodied intelligence sector.

image.png

Image Credit: BeingBeyond

The first phase, from Being-H0 to Being-H0.5, focused on verifying "whether human video can be used to train embodied foundation models," completing a feasibility proof-of-concept from zero to one.

The second phase, from Being-H0.5 to Being-H0.7, aimed to solve "how to use human video at scale," with data volume climbing from tens of thousands of hours to 200,000 hours.

With this release of Being-H0.8, the company officially enters the third phase, focusing on how to continuously drive model evolution with high-quality, multimodal data. The integration of tactile modalities and the refinement of the full-stack data-model pipeline are direct responses to this challenge.

Founded in May 2025, BeingBeyond has built a full-stack infrastructure spanning data pipelines, model pre-training, post-training, evaluation, and edge deployment.

From Being-H0 to Being-H0.8, through multiple iterations, BeingBeyond has aggregated over 500,000 hours of first-person video data and established partnerships with dozens of data collaborators.

Gasgoo not only offers timely news and profound insight about China auto industry, but also help with business connection and expansion for suppliers and purchasers via multiple channels and methods. Buyer service: buyer-support@gasgoo.com Seller Service: seller-support@gasgoo.com

All Rights Reserved. Do not reproduce, copy and use the editorial content without permission. Contact us: autonews@gasgoo.com