Skild AI says S1 can watch one video of an unseen task lasting up to 10 minutes and execute it without changing its model weights or collecting task specific data. S1’s demonstrations include pour over coffee, plant potting, kit assembly, and pancake flipping, with the model reportedly adapting to changed objects an...
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Skild AI’s S1 robotics foundation model, unveiled by co-founder Abhinav Gupta at SMARTCLOUD SHOW 2026 in Seoul on August 25, 2026, a. Article summary: S1 is Skild AI’s flagship robotics foundation model for in-context learning: a robot watches one video demonstration—including an unseen task lasting up to 10 minutes—and then executes it without collecting task-specific. Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
Skild AI’s S1 is a robotics foundation model built around in-context learning: instead of being retrained for every new job, a robot watches a video demonstration and uses it as the specification for what to do. Skild says S1 can handle previously unseen tasks lasting up to 10 minutes, without task-specific fine-tuning or changing the model’s weights. 1
5
That makes S1 an intriguing step toward more adaptable physical AI. It does not, however, prove that general-purpose robots are ready for unrestricted household or factory work. The strongest results currently come from Skild’s own demonstrations and benchmarks, so independent testing remains important.
Many robot-learning systems are trained around a particular task, robot body, or set of operating conditions. S1 instead treats a demonstration video as an executable prompt. The model must infer the goal, sequence of actions, object relationships, and physical adjustments needed to reproduce the task on a robot. 1
3
That is different from simply telling a robot to “make coffee” or “move the plant.” Language can identify a goal, but a video can show the order of operations, the way an object is grasped, how tools are used, and how the task unfolds over many steps. Skild’s central claim is that this richer prompt lets S1 generalize to long-horizon tasks that were not present in pretraining. 1
5
The company describes the approach as a robotics analogue of prompting a language model: the model’s underlying capabilities stay fixed while the demonstration supplies the task-specific context.
Skild has shown S1 performing several manipulation workflows, including:
Some demonstrations run for as long as 10 minutes and involve dozens of sequential actions. 1
3
5
Skild also reports qualitative tests designed to show that S1 is not merely replaying a memorized motion. In those tests, the model reportedly responded to changes in the scene, retried after some errors, substituted a cup when the demonstrated watering can was unavailable, and corrected a demonstrator’s awkward handling of an egg rather than copying it exactly. These are useful signs of behavioral flexibility, although video demonstrations alone do not establish reliability across a broad range of environments. 5
The most specific comparison supplied by Skild is an internal evaluation of unseen tasks. At 100,000 hours of training data, the company reported a 66% average per-step success rate for its video-prompted in-context policy, compared with 9% for a matched language-prompted vision-language-action, or VLA, baseline. 1
4
Skild also says the language-prompted VLA degraded by as much as three times more than S1 when deployment conditions changed substantially—for example, when objects differed from those in the demonstration or when the execution plan required the opposite arm. 1
Another company-reported comparison estimated that one video prompt produced performance comparable to roughly 380 task-specific post-training episodes on a pretrained policy. 1
These numbers suggest why video prompting could matter: a demonstration communicates not only the desired outcome, but also the intended sequence, timing, object arrangement, tool use, and physical strategy. But the results should be read as Skild’s internal evidence, not as an independently reproduced head-to-head benchmark. The supplied material does not include a public paper, model weights, or a standardized external evaluation. 4
15
Skild describes S1 as operating in real time after watching the demonstration, and reporting around the launch said the robot could execute the task immediately after viewing it. 2
5
There is not enough evidence, however, to give a verified figure such as “within seconds” or “within 11 minutes” for the general system. A specific transition time from the end of the video to the start of execution is not established by the supplied sources. That distinction matters in industrial settings, where setup time, inference latency, safety checks, and recovery behavior can be as important as the model’s raw task success.
Skild credits a broad data pipeline with helping it build a general-purpose robotics model. The company has described using large-scale simulation, internet videos of human actions, robot data, teleoperation, and information from real deployments. Its Series C materials also describe a data flywheel in which deployment generates more experience that can improve the model across robots and environments. 37
At SMARTCLOUD SHOW, co-founder Abhinav Gupta reportedly said Skild uses “every kind” of robot-data collection and has accumulated 100 to 1,000 times more data than competitors. That multiplier is a company assertion, not a publicly audited industry comparison. 6
The practical strategy is to combine data sources that teach different parts of the problem: human videos provide examples of objects and actions, simulation expands the range of possible robot bodies and environments, and real robot data connects those abstractions to physical control.
S1 is part of Skild’s broader strategy for an “omni-bodied” robot brain. The term does not mean that a biped, quadruped, wheeled robot, and robotic arm all move in the same way. Their joints, sensors, limbs, wheels, and physically possible actions are different. 45
Instead, the intended architecture learns physical regularities that transfer across platforms and then maps those expectations onto the capabilities of a particular robot. A single model can therefore aim to share knowledge across embodiments while adapting its outputs to each robot’s available actuators and sensors. 45
This approach is designed to reduce the need to build and maintain a completely separate model for every robot configuration. Skild has also described training across large numbers of simulated robot bodies as part of its effort to learn more general physical behavior. 42
45
Skild has announced collaborations with ABB Robotics and Universal Robots to integrate its software into industrial systems. Reuters also reported that Skild and Nvidia were deploying the robot brain on assembly lines associated with Nvidia Blackwell systems. 34
44
Those announcements show a path toward commercial testing, but they do not provide a complete public performance record. The supplied evidence does not establish independently audited figures for factory throughput, uptime, error rates, safety incidents, or return on investment. Deployment is evidence that the technology is being tested or integrated—not proof that it has already replaced conventional automation at scale.
Investor interest in Skild has been substantial. Bloomberg reported a $1.4 billion Series C led by SoftBank that valued the company at more than $14 billion, more than tripling its valuation in roughly seven months. 17
That financing reflects confidence in the possibility that robotics software could become a scalable platform layer across many types of hardware. Other reports put Skild’s cumulative funding above $1.8 billion, but the available figures vary by source and should not be treated as a definitive audited total. 18
21
22
The “ChatGPT moment for robots” comparison is best understood as an analogy, not a settled conclusion. S1 points toward a possible change in the way robots acquire new skills—from task-by-task retraining toward learning from demonstrations. Investor Ravi Mhatre described the launch as a “GPT-3 moment” for long-horizon human labor tasks, but that is an investor judgment rather than independent scientific validation. 10
S1’s most important idea is not simply that a robot can imitate a video. It is that a foundation model might use one demonstration to compose skills, adapt to a changed scene, and execute a long sequence without updating its weights for that specific task.
Skild’s reported 66% versus 9% benchmark result and its long-horizon demonstrations make the claim technically significant. 1
4 The caveat is equally important: the public evidence is still dominated by company materials, selected demonstrations, and early deployment announcements. Until outside researchers can reproduce the results under standardized conditions—and industrial customers publish reliability and safety data—S1 should be viewed as a promising direction for physical AI rather than definitive proof that a general-purpose robot brain has arrived.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Skild AI says S1 can watch one video of an unseen task lasting up to 10 minutes and execute it without changing its model weights or collecting task specific data.
Skild AI says S1 can watch one video of an unseen task lasting up to 10 minutes and execute it without changing its model weights or collecting task specific data. S1’s demonstrations include pour over coffee, plant potting, kit assembly, and pancake flipping, with the model reportedly adapting to changed objects and recovering from some mistakes.
The technology is being positioned for industrial robots, but public evidence does not yet establish production grade metrics such as uptime, throughput, or return on investment.