AgiBot reports that scaling GE Act 2.0’s embodied training data 100×, from 300 to 30,000 hours, raised zero shot real robot success from 17.1% to 44.1% on G1 OP and from 13.4% to 31.1% on G2 90D. The model is trained from random initialization on manipulation data rather than adapted from a pretrained web video mode...
Δημοσιεύτηκε απόΕπεξεργασία με GPT-5.6 TerraΕικόνες δημιουργήθηκαν με GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is AgiBot’s GE-Act 2.0 native world-action model, and how do its random-initialization training approach, 100-fold embodied-data scalin. Article summary: GE-Act 2.0 is AgiBot’s “native” world–action model: its visual representation, future-state generation, and action-prediction components are trained from random initialization on manipulation data, rather than adapting a. Topic tags: general, general web, user generated, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
GE-Act 2.0 is AgiBot’s attempt to build a robot-focused world–action model from the ground up. Rather than adapting a general video model trained on web data, its trainable visual, future-state, and action components are initialized randomly and pretrained on embodied manipulation data. The central claim is not simply that the model becomes more reliable with more data: AgiBot reports that previously unsuccessful fine-manipulation tasks begin to work as training grows from 300 to 30,000 hours. 17
That is a meaningful research result, but it needs a careful reading. The evidence comes from AgiBot’s reported evaluation protocol on two robot embodiments and 100 atomic tasks. It does not show that GE-Act 2.0 can autonomously handle arbitrary homes, factories, or robot hardware.
A world–action model links perception, prediction, and control. For a language instruction such as placing an object somewhere, the system must identify the relevant scene state, predict what a successful near-future state should look like, and issue motor commands that can produce it.
GE-Act 2.0 divides that job into three reported components:
AgiBot calls the training method Knowledge-Aligned Selective Optimization (KASO). The paper presents it as a way to coordinate the visual-planning and action-learning pieces while allowing them to be pretrained on complementary data. 17
The team also reports action-chunk inference of 104 milliseconds on a single RTX 5090. That figure matters because a manipulation policy must repeatedly observe and react; it is not enough to make a visually plausible plan if producing the next control command is too slow. 5
Many robot-learning systems benefit from a model first trained on broad image or web-video corpora. GE-Act 2.0 instead initializes its trainable generative and action components from scratch and trains them on manipulation data. 17
That makes the scaling experiment more focused. The researchers are testing whether additional embodied data—data tied to physical interactions and actions—produces systematically better robot manipulation, rather than measuring the value of a pre-existing web-video foundation model.
The reported recipe can use multiple kinds of data, including robot teleoperation trajectories, first-person human video, unlabeled rollout data, and failure data. This is important because robot demonstrations are costly; a training approach that can learn from imperfect or partially labeled interaction data could make data collection more practical. 6
20
AgiBot trained versions of the model using 300, 1,200, 5,000, and 30,000 hours of manipulation data. The final comparison therefore spans a 100-fold increase in data. 4
17
The primary evaluation was designed as a real-robot zero-shot test:
Under that protocol, reported aggregate success improved as the data scale increased:
| Robot embodiment | Success at 300 hours | Success at 30,000 hours |
|---|---|---|
| G1-OP | 17.1% | 44.1% |
| G2-90D | 13.4% | 31.1% |
These are reported results from the GE-Act 2.0 study. 3
17
The most direct conclusion is that, for this model and task distribution, more embodied data was associated with substantially higher zero-shot task success. The reported curves between 5,000 and 30,000 hours did not show an obvious plateau, which the authors interpret as evidence that further scaling may still help. 5
A higher average success rate can mean a robot is simply becoming more reliable at tasks it could already complete. AgiBot argues that its larger model-data regime showed something stronger: a growing number of tasks achieved non-zero success.
For G1-OP, the reported count of atomic tasks with successful completion increased from 39 to 76 out of 100 as training data rose from 300 to 30,000 hours. 4
The examples cited include:
This is the basis for the paper’s “light-up” framing. In the reported trials, data scaling was associated not only with better execution of established abilities, but with fine skills becoming feasible in zero-shot settings after failing at smaller scales.
Still, “light-up” should not be read as proof of sudden general intelligence. It is a description of observed task coverage in a finite benchmark. Whether the same pattern holds for more varied objects, longer-horizon work, safety-critical settings, or new robot designs remains an open question.
The G2-90D provides the study’s evidence for cross-embodiment transfer. AgiBot reports that this wheeled mobile manipulator accounted for less than 2% of the co-training data, yet its aggregate success improved by 17.7 percentage points between the small- and large-data conditions. 3
17
That result suggests that at least some learned visual and action-relevant structure can be shared across different robot bodies. It is a useful sign for a future in which policies do not have to be rebuilt entirely for every machine.
But it does not establish hardware-agnostic robot intelligence. The evidence is limited to two evaluated embodiments and the study’s task distribution. The G2-90D also had representation in co-training, so the findings do not show zero-data transfer to an arbitrary new robot.
GE-Act belongs to AgiBot’s broader Genie Envisioner (GE) platform, which is described as combining robot policy learning, evaluation, and simulation in a shared video-generative framework. 18
Within that broader framing, GE-Act is the action-oriented component, while GE-Sim is a world-simulation component intended to generate action-conditioned rollouts for evaluation and development. 18
19
The distinction matters: the headline result for GE-Act 2.0 is based on zero-shot tests on physical robots. GE-Sim may be related infrastructure in the wider platform, but it is not the basis for the central claim that the model executed the reported real-world manipulation tasks without task-specific fine-tuning or evaluation demonstrations.
AgiBot also reports fine-tuned results on simulation-oriented benchmarks. In the provided reporting, GE-Act 2.0 achieved 60.52% success in a RoboTwin unseen-environment test and 0.770 on GenieSim instruction following; the latter was reported as above the compared π0.5 and GR00T N1.7 results. 20
Those results can be useful measures of adaptation and benchmark performance, but they answer a different question from the zero-shot physical-robot experiment. Fine-tuned benchmark scores should not be used as evidence that the same capability was achieved with no task-specific adaptation.
GE-Act 2.0 is notable because it tests a simple but important proposition for embodied AI: can a world–action model trained directly on manipulation data improve predictably as that data grows?
In AgiBot’s reported experiment, the answer is yes. Expanding the training set from 300 to 30,000 hours coincided with large gains in zero-shot success across G1-OP and G2-90D, broader task coverage, and fine skills such as towel folding and cup nesting appearing in the larger-scale condition. 3
4
17
The appropriate conclusion is narrower than the marketing promise of a general-purpose robot: GE-Act 2.0 offers evidence that scaling embodied manipulation data can produce new, measurable capabilities in a controlled real-robot evaluation. Independent replication, broader hardware tests, and stronger long-horizon and safety evaluations are still needed before treating it as a general solution to robot autonomy.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
AgiBot reports that scaling GE Act 2.0’s embodied training data 100×, from 300 to 30,000 hours, raised zero shot real robot success from 17.1% to 44.1% on G1 OP and from 13.4% to 31.1% on G2 90D.
AgiBot reports that scaling GE Act 2.0’s embodied training data 100×, from 300 to 30,000 hours, raised zero shot real robot success from 17.1% to 44.1% on G1 OP and from 13.4% to 31.1% on G2 90D. The model is trained from random initialization on manipulation data rather than adapted from a pretrained web video model, and pairs visual planning with inverse dynamics to produce robot actions.
Its strongest reported finding is that larger datasets expanded the set of skills that worked at all—including towel folding and cup nesting—while also improving cross embodiment performance.