AgiBot reports that scaling GE Act 2.0’s embodied training data from 300 to 30,000 hours raised zero shot real robot success from 17.1% to 44.1% on G1 OP and from 13.4% to 31.1% on G2 90D. The system is trained from random initialization on manipulation data rather than adapted from a pretrained web video model, com...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is AgiBot’s GE-Act 2.0 native world-action model, and how do its random-initialization training approach, 100-fold embodied-data scalin. Article summary: GE-Act 2.0 is AgiBot’s “native” world–action model: its visual representation, future-state generation, and action-prediction components are trained from random initialization on manipulation data, rather than adapting a. Topic tags: general, general web, user generated, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
GE-Act 2.0 is AgiBot’s attempt to build a robot-focused world–action model from the ground up. Rather than adapting a general video model trained on web data, its trainable visual, future-state and action components begin from random initialization and are pretrained on embodied manipulation data. The central claim is not merely that the system becomes more reliable with more data: AgiBot reports that previously unsuccessful fine-manipulation tasks started working as training grew from 300 to 30,000 hours. 17
That is a meaningful research result—but it needs to be read carefully. The evidence comes from AgiBot’s reported protocol: two robot embodiments, 100 atomic tasks and zero-shot physical-robot testing. It does not demonstrate that GE-Act 2.0 can independently manage arbitrary homes, factories or unfamiliar robot hardware.
A world–action model connects perception, prediction and control. Given a language instruction such as placing an object in a particular location, a robot must identify the relevant state of the scene, anticipate what a successful next state should look like, and issue motor commands that can create it.
GE-Act 2.0 divides that process into three reported components:
AgiBot calls its training approach Knowledge-Aligned Selective Optimization (KASO). The paper presents KASO as a way to coordinate visual planning and action learning while letting the components be pretrained on complementary types of data. 17
The team also reports action-chunk inference in 104 milliseconds on a single RTX 5090 GPU. That matters for manipulation because a robot needs to repeatedly observe and respond; a visually convincing plan is not enough if generating the next control command takes too long. 5
Many robot-learning systems benefit from models first trained on huge image or web-video datasets. GE-Act 2.0 instead initializes its trainable generative and action components from scratch, then trains them on manipulation data. 17
That makes its scaling experiment more focused. The researchers are testing whether more embodied data—data connected to physical interaction and actions—can systematically improve robot manipulation, rather than measuring the benefit of an existing web-video foundation model.
The reported recipe can incorporate several kinds of data, including robot teleoperation trajectories, first-person human video, unlabeled rollout data and failure data. This matters because high-quality robot demonstrations are expensive to collect; learning from imperfect or partially labeled interaction data could make training data more practical to assemble. 6
20
AgiBot trained versions of the model with 300, 1,200, 5,000 and 30,000 hours of manipulation data. The smallest-to-largest comparison is therefore a 100-fold increase in training data. 4
17
The main evaluation was designed as a real-robot zero-shot test, featuring:
Under that protocol, aggregate success rose with training scale:
| Robot embodiment | Success at 300 hours | Success at 30,000 hours |
|---|---|---|
| G1-OP | 17.1% | 44.1% |
| G2-90D | 13.4% | 31.1% |
These are results reported by the GE-Act 2.0 study. 3
17
The clearest conclusion is that, for this model and task distribution, more embodied data was associated with substantially better zero-shot task success. The reported curves between 5,000 and 30,000 hours showed no obvious plateau, which the authors interpret as a sign that further scaling may still help. 5
A higher average success rate can simply mean a robot is becoming more dependable at tasks it could already perform. AgiBot argues that its larger model-and-data setting showed something stronger: more tasks reached non-zero success.
For G1-OP, the reported number of atomic tasks with successful completion rose from 39 to 76 out of 100 as training data increased from 300 to 30,000 hours. 4
The examples cited include:
That is the basis for the paper’s “light-up” framing. In the reported trials, data scaling was linked not just to stronger performance on established abilities, but to fine skills becoming feasible in zero-shot settings after failing at smaller scales.
Still, “light-up” should not be taken as proof of sudden general intelligence. It describes observed task coverage in a finite benchmark. Whether the pattern extends to more varied objects, longer multi-stage work, safety-critical tasks or entirely new robot designs remains unresolved.
The G2-90D supplies the study’s evidence for cross-embodiment transfer. AgiBot reports that this wheeled mobile manipulator made up less than 2% of co-training data, yet its aggregate success increased by 17.7 percentage points between the small- and large-data conditions. 3
17
The result suggests that at least some learned visual and action-relevant structure can be shared across different robot bodies. That would be important if future robot policies can avoid being rebuilt from scratch for every machine.
But it does not establish hardware-agnostic robot intelligence. The evidence covers only two evaluated embodiments and the study’s task distribution. G2-90D was also represented in co-training, so the result is not a demonstration of zero-data transfer to an arbitrary new robot.
GE-Act is part of AgiBot’s broader Genie Envisioner (GE) platform, described as combining robot policy learning, evaluation and simulation in a shared video-generative framework. 18
Within that broader platform, GE-Act is the action-oriented component, while GE-Sim is a world-simulation component intended to generate action-conditioned rollouts for evaluation and development. 18
19
The distinction matters. GE-Act 2.0’s headline finding comes from zero-shot tests on physical robots. GE-Sim may be related infrastructure within the wider platform, but it is not the basis for the central claim that GE-Act 2.0 executed the reported manipulation tasks on real robots without task-specific fine-tuning or evaluation demonstrations.
AgiBot also reports fine-tuned results on simulation-oriented benchmarks. In the provided reporting, GE-Act 2.0 achieved 60.52% success in a RoboTwin unseen-environment test and scored 0.770 on GenieSim instruction following; the latter was reported as higher than the compared π0.5 and GR00T N1.7 results. 20
Those numbers may be useful measures of adaptation and benchmark performance, but they answer a different question than the zero-shot physical-robot experiment. Fine-tuned benchmark scores should not be treated as proof that the same capability was achieved with no task-specific adaptation.
GE-Act 2.0 is notable because it tests a straightforward but important proposition in embodied AI: can a world–action model trained directly on manipulation data improve predictably as its data grows?
In AgiBot’s reported experiment, the answer is yes. Scaling from 300 to 30,000 hours coincided with large zero-shot gains across G1-OP and G2-90D, broader task coverage and fine skills such as towel folding and cup nesting appearing in the larger-scale condition. 3
4
17
The appropriately narrow conclusion is not that general-purpose robot autonomy has arrived. Rather, GE-Act 2.0 provides evidence that scaling embodied manipulation data can yield new, measurable capabilities in a controlled real-robot evaluation. Independent replication, broader hardware testing, and stronger evaluations of long-horizon reliability and safety are still needed before it can be considered a general solution to robot autonomy.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
AgiBot reports that scaling GE Act 2.0’s embodied training data from 300 to 30,000 hours raised zero shot real robot success from 17.1% to 44.1% on G1 OP and from 13.4% to 31.1% on G2 90D.
AgiBot reports that scaling GE Act 2.0’s embodied training data from 300 to 30,000 hours raised zero shot real robot success from 17.1% to 44.1% on G1 OP and from 13.4% to 31.1% on G2 90D. The system is trained from random initialization on manipulation data rather than adapted from a pretrained web video model, combining visual planning with inverse dynamics to generate robot actions.
Its most notable reported result is broader task coverage: skills including towel folding and cup nesting began to succeed only in the larger data condition, though the findings remain limited to the study’s two robot...