Both efforts move robot learning from task-specific retraining toward inference-time imitation: a pretrained system treats a brief demonstration as context and immediately controls the robot accordingly. HOST offers the stronger academic evidence; GEN-1.5 is a promising company-r Comparison
Research answer

Create a landscape editorial hero image for this Studio Global article: How do the two research efforts, HOST from Beijing Institute of Technology, X Square Robot, and Tsinghua University and GEN 1.5 from General. Article summary: Both efforts move robot learning from task specific retraining toward inference time imitation: a pretrained system treats a brief demonstration as context and immediately controls the robot accordingly.. Topic tags: general web, ai safety, prompt engineering, ai, workflow. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with
Both efforts move robot learning from task-specific retraining toward inference-time imitation: a pretrained system treats a brief demonstration as context and immediately controls the robot accordingly. HOST offers the stronger academic evidence; GEN-1.5 is a promising company-reported result but lacks comparable independent validation.
HOST shows a concrete version of human-to-robot observational learning: the robot does not merely replay a recovered hand trajectory. It uses the video as a progress-conditioned reference while grounding its own observations and body state in execution.
Its 50-task, 20-trials-per-task evaluation provides evidence that this approach can transfer across new objects, tools, and manipulation primitives, although the 62% average success rate also shows that it is not yet reliably deployable for every task.
HOST’s reported robustness tests retained most of its 62% baseline under changes in lighting, objects, scene, and human disturbance, with reported losses of 1%, 4%, 6%, and 9%, respectively.
GEN-1.5 frames the same idea as emergent in-context learning in a large embodied model: the demonstration is placed in a context window rather than converted into a newly trained policy. The company also claims human-to-robot transfer in some cases, composition of two prompted behaviors, and simulation-video prompting for physical execution.
GEN-1.5’s alternative adaptation path is important: it is not strictly “no training” in every use case. One-shot physical prompting requires no update, but the company also offers few-shot gradient-based adaptation for harder tasks.
These systems support a shift toward a more human-like teaching interface: show rather than program, reward-design, collect hundreds of robot demonstrations, or perform a full training run. A broad pretrained model or policy supplies prior physical knowledge; the observed demonstration specifies the new goal, procedure, and relevant interactions.
The distinction is consequential: the reported efficiency is not that the robots learn from no prior data. Rather, costly large-scale prior training is amortized, so a new end user may teach a task with one short observation instead of running task-specific optimization. HOST makes that trade-off quantitatively explicit through its 29-second acquisition and claimed 507× speedup.
HOST is an academic preprint, not a guarantee of open-world reliability. Its evaluation is controlled laboratory experimentation on 50 selected novel tasks, with an average success rate of 62%; broader household, industrial, long-horizon, safety-critical, and independently replicated performance remains unproven.
GEN-1.5’s evidence is substantially less independently assessable. The model, its task claims, training duration, no-simulation assertion, prompt duration, and adaptation claims are published by Generalist AI or relayed from its announcements.
No comparable public benchmark is established for GEN-1.5. There is insufficient public, independent evidence of a standardized task count, trial protocol, aggregate one-shot success rate, baseline comparison, model/data release, or third-party reproduction. Thus its demonstrations are evidence of potential capability, not yet a verified performance frontier.
Taken together, HOST is evidence that single-video, parameter-free skill acquisition can work across a defined manipulation test set, while GEN-1.5 suggests that sufficiently broad embodied pretraining may make similar behavior an in-context capability of a generalist model. The compelling direction is real, but the practical conclusion should remain measured: observational learning is emerging, not yet a demonstrated replacement for robust task validation, safety engineering, and adaptation in unconstrained environments.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Both efforts move robot learning from task-specific retraining toward inference-time imitation: a pretrained system treats a brief demonstration as context and immediately controls the robot accordingly. HOST offers the stronger academic evidence; GEN-1.5 is a promising company-r
Both efforts move robot learning from task-specific retraining toward inference-time imitation: a pretrained system treats a brief demonstration as context and immediately controls the robot accordingly. HOST offers the stronger academic evidence; GEN-1.5 is a promising company-r ## Comparison