Skild AI says S1 can execute previously unseen, multi step tasks lasting up to 10 minutes after watching one human video, without fine tuning; its reported 66% step success rate on unseen tasks exceeded a language pro... S1 treats the video as an instruction at inference time, combining skills learned during pretrai...
Published byEdited with GPT-5.6 LunaImages generated with GPT Image 1.5
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Skild AI’s S1 robotics foundation model, how does it enable robots to learn and perform previously unseen long-horizon manipulation. Article summary: Skild AI’s S1 is a robotics foundation model designed for visual in-context learning: rather than retraining a policy for each new job, a robot conditions on one video demonstration and then attempts the task in real tim. Topic tags: general, general web, news, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts wi
Skild AI’s S1 is designed to make robot programming look more like teaching by example. A person demonstrates a task in a video, and the model uses that demonstration as a visual prompt rather than starting a new task-specific training run. Skild says S1 can handle previously unseen, multi-step manipulation tasks lasting up to 10 minutes, with no fine-tuning or post-observation training cycle. 1
That distinction matters because robotics systems have traditionally required substantial task-specific data, teleoperation, engineering, or retraining before a robot can perform a new job. S1’s promise is not that a robot memorizes a video frame by frame, but that it can interpret the goal and compose capabilities it learned earlier into a new sequence of actions.
S1 uses visual in-context learning. The human demonstration functions as a task prompt at inference time. The model observes what happens, infers the intended sequence and objective, and translates that information into robot actions in the current environment. The model’s weights are not changed for each new demonstration, according to Skild’s description of S1. 1
This is closer to showing an AI system what to do than to conventional robot training. The model must connect visual observations with previously learned skills, such as grasping, moving, placing, and manipulating objects, then order those skills into a longer procedure. Skild’s public examples include pour-over coffee, plant potting, kit assembly, and pancake flipping. 1
35
The difficult part is the long horizon. A short action can often be learned or scripted in isolation; a 10-minute task may contain dozens of decisions, changing object positions, and opportunities for error. S1 is intended to maintain the task objective across that extended sequence and adapt its actions to the scene rather than simply replaying the demonstrator’s exact movements.
Skild reports a substantial difference between video prompting and language prompting on its unseen-task evaluation. At a stated 100,000-hour training scale, the video-prompted policy reached a 66% average per-step success rate, compared with 9% for a comparable language-prompted vision-language-action policy. 32
The result suggests that a visual demonstration can communicate physical details that words may leave ambiguous: which object to grasp, how it should be oriented, the order of actions, and how the demonstrator responds to the environment. But the number needs careful interpretation. It comes from Skild’s reported evaluation and is not presented here as an independently replicated industry benchmark. A per-step success figure also should not be read as a 66% guarantee that an entire 10-minute task will be completed from beginning to end.
The public demonstrations cover manipulation tasks such as:
These examples are useful because they combine perception, object handling, ordering, and sustained control rather than testing a single isolated movement. Skild describes the tasks as unseen during pretraining, although the available evidence comes primarily from the company and reports summarizing its claims. 1
33
The demonstrations show the intended interface: a person performs the task once, and the robot attempts to generalize the procedure. They do not establish that S1 can reliably perform arbitrary household work, operate without supervision, or handle every physical variation found in a production environment.
S1’s advertised advantage is that it can begin operating after observing the demonstration, without a separate post-observation training phase. Skild and coverage of the launch describe this as real-time in-context execution. 1
30
However, the available sources do not provide a reliable, precise latency measurement—such as the number of seconds between the end of the video and the robot’s first action. “No fine-tuning” means the model does not undergo a new task-specific weight-update cycle; it does not necessarily mean that every video is processed instantaneously or that the robot needs no setup, calibration, or safety checks.
S1 is part of Skild’s larger Skild Brain strategy: train a general-purpose model across tasks, robot morphologies, simulation, and visual human data instead of building a separate policy for every robot and job. Skild says it created a simulated universe containing 100,000 robot variants and tested generalization to bodies excluded from the training set. 6
Training across different embodiments is intended to expose the model to general control principles. Skild’s materials describe a system meant to work across humanoids, quadrupeds, tabletop arms, and mobile manipulators. 3
10
Claims that Skild has 100 to 1,000 times more data than competitors should be treated cautiously. The supplied sources do not provide a standardized, independently auditable comparison of dataset size, quality, or useful robot experience across companies. The stronger defensible point is that Skild is pursuing scale through cross-embodiment training and simulation rather than relying only on bespoke demonstrations for one hardware platform.
“Omni-bodied” is Skild’s term for a model intended to transfer across different robot bodies. In principle, a shared model can learn task-level concepts separately from the precise geometry of one machine, allowing it to adapt to different limbs, wheels, manipulators, and actuator layouts. Skild says its simulated tests included held-out robot forms controlled zero-shot. 6
That does not make hardware irrelevant. Robots still differ in sensors, grippers, joint limits, control interfaces, latency, payloads, and safety constraints. A model that generalizes across embodiments may reduce adaptation work, but deployment still requires compatible hardware, system integration, and validation in the target environment.
Public reporting says Skild’s software is operating on hundreds of robots across factories, data centers, and logistics settings. Reported examples include NVIDIA’s Houston factory, a LaGuardia trial, and work with wiring and construction companies. The company has also announced relationships involving ABB Robotics, Universal Robots, and NVIDIA. 17
4
Those reports indicate commercial interest and real-world testing, but they do not provide enough public evidence to calculate production completion rates, uptime, cost savings, or return on investment. A laboratory or controlled evaluation result should not be treated as a production metric. One report also describes a planned deployment with Foxconn on assembly lines associated with NVIDIA Blackwell GPU server systems, which is evidence of a test or deployment effort rather than proof of broad autonomous factory operation. 43
Skild’s financial backing has helped make its approach one of the most closely watched efforts in physical AI. Bloomberg reported that the company raised about $1.4 billion in a SoftBank-led Series C in January 2026 at a valuation above $14 billion. 13 Skild’s earlier Series A, announced in July 2024, was $300 million at a reported $1.5 billion valuation.
9
The “ChatGPT moment” analogy captures the hoped-for change in user experience: instead of programming or retraining a robot for every task, a worker could demonstrate the desired behavior and let a general model handle the translation into actions. S1 is a notable step toward that interface.
It is not yet proof that general-purpose robots have arrived. The strongest public results are company-reported, the task set remains limited, and independently verified industrial metrics are not available in the supplied evidence. The more measured conclusion is that S1 demonstrates a promising way to reduce task-specific robot training: use a video as an instruction, rely on broad pretraining to supply reusable skills, and test whether those skills can be composed into a new long-horizon task.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Skild AI says S1 can execute previously unseen, multi step tasks lasting up to 10 minutes after watching one human video, without fine tuning; its reported 66% step success rate on unseen tasks exceeded a language pro...
Skild AI says S1 can execute previously unseen, multi step tasks lasting up to 10 minutes after watching one human video, without fine tuning; its reported 66% step success rate on unseen tasks exceeded a language pro... S1 treats the video as an instruction at inference time, combining skills learned during pretraining into actions for the robot’s current environment.
The broader Skild Brain strategy trains across many robot bodies and simulated variants, but public evidence still does not establish production level completion rates, uptime, or return on investment.
Skild AI says S1 can execute previously unseen, multi step tasks lasting up to 10 minutes after watching one human video, without fine tuning; its reported 66% step success rate on unseen tasks exceeded a language pro... S1 treats the video as an instruction at inference time, combining skills learned during pretrai...
Published byEdited with GPT-5.6 LunaImages generated with GPT Image 1.5
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Skild AI’s S1 robotics foundation model, how does it enable robots to learn and perform previously unseen long-horizon manipulation. Article summary: Skild AI’s S1 is a robotics foundation model designed for visual in-context learning: rather than retraining a policy for each new job, a robot conditions on one video demonstration and then attempts the task in real tim. Topic tags: general, general web, news, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts wi
Skild AI’s S1 is designed to make robot programming look more like teaching by example. A person demonstrates a task in a video, and the model uses that demonstration as a visual prompt rather than starting a new task-specific training run. Skild says S1 can handle previously unseen, multi-step manipulation tasks lasting up to 10 minutes, with no fine-tuning or post-observation training cycle. 1
That distinction matters because robotics systems have traditionally required substantial task-specific data, teleoperation, engineering, or retraining before a robot can perform a new job. S1’s promise is not that a robot memorizes a video frame by frame, but that it can interpret the goal and compose capabilities it learned earlier into a new sequence of actions.
S1 uses visual in-context learning. The human demonstration functions as a task prompt at inference time. The model observes what happens, infers the intended sequence and objective, and translates that information into robot actions in the current environment. The model’s weights are not changed for each new demonstration, according to Skild’s description of S1. 1
This is closer to showing an AI system what to do than to conventional robot training. The model must connect visual observations with previously learned skills, such as grasping, moving, placing, and manipulating objects, then order those skills into a longer procedure. Skild’s public examples include pour-over coffee, plant potting, kit assembly, and pancake flipping. 1
35
The difficult part is the long horizon. A short action can often be learned or scripted in isolation; a 10-minute task may contain dozens of decisions, changing object positions, and opportunities for error. S1 is intended to maintain the task objective across that extended sequence and adapt its actions to the scene rather than simply replaying the demonstrator’s exact movements.
Skild reports a substantial difference between video prompting and language prompting on its unseen-task evaluation. At a stated 100,000-hour training scale, the video-prompted policy reached a 66% average per-step success rate, compared with 9% for a comparable language-prompted vision-language-action policy. 32
The result suggests that a visual demonstration can communicate physical details that words may leave ambiguous: which object to grasp, how it should be oriented, the order of actions, and how the demonstrator responds to the environment. But the number needs careful interpretation. It comes from Skild’s reported evaluation and is not presented here as an independently replicated industry benchmark. A per-step success figure also should not be read as a 66% guarantee that an entire 10-minute task will be completed from beginning to end.
The public demonstrations cover manipulation tasks such as:
These examples are useful because they combine perception, object handling, ordering, and sustained control rather than testing a single isolated movement. Skild describes the tasks as unseen during pretraining, although the available evidence comes primarily from the company and reports summarizing its claims. 1
33
The demonstrations show the intended interface: a person performs the task once, and the robot attempts to generalize the procedure. They do not establish that S1 can reliably perform arbitrary household work, operate without supervision, or handle every physical variation found in a production environment.
S1’s advertised advantage is that it can begin operating after observing the demonstration, without a separate post-observation training phase. Skild and coverage of the launch describe this as real-time in-context execution. 1
30
However, the available sources do not provide a reliable, precise latency measurement—such as the number of seconds between the end of the video and the robot’s first action. “No fine-tuning” means the model does not undergo a new task-specific weight-update cycle; it does not necessarily mean that every video is processed instantaneously or that the robot needs no setup, calibration, or safety checks.
S1 is part of Skild’s larger Skild Brain strategy: train a general-purpose model across tasks, robot morphologies, simulation, and visual human data instead of building a separate policy for every robot and job. Skild says it created a simulated universe containing 100,000 robot variants and tested generalization to bodies excluded from the training set. 6
Training across different embodiments is intended to expose the model to general control principles. Skild’s materials describe a system meant to work across humanoids, quadrupeds, tabletop arms, and mobile manipulators. 3
10
Claims that Skild has 100 to 1,000 times more data than competitors should be treated cautiously. The supplied sources do not provide a standardized, independently auditable comparison of dataset size, quality, or useful robot experience across companies. The stronger defensible point is that Skild is pursuing scale through cross-embodiment training and simulation rather than relying only on bespoke demonstrations for one hardware platform.
“Omni-bodied” is Skild’s term for a model intended to transfer across different robot bodies. In principle, a shared model can learn task-level concepts separately from the precise geometry of one machine, allowing it to adapt to different limbs, wheels, manipulators, and actuator layouts. Skild says its simulated tests included held-out robot forms controlled zero-shot. 6
That does not make hardware irrelevant. Robots still differ in sensors, grippers, joint limits, control interfaces, latency, payloads, and safety constraints. A model that generalizes across embodiments may reduce adaptation work, but deployment still requires compatible hardware, system integration, and validation in the target environment.
Public reporting says Skild’s software is operating on hundreds of robots across factories, data centers, and logistics settings. Reported examples include NVIDIA’s Houston factory, a LaGuardia trial, and work with wiring and construction companies. The company has also announced relationships involving ABB Robotics, Universal Robots, and NVIDIA. 17
4
Those reports indicate commercial interest and real-world testing, but they do not provide enough public evidence to calculate production completion rates, uptime, cost savings, or return on investment. A laboratory or controlled evaluation result should not be treated as a production metric. One report also describes a planned deployment with Foxconn on assembly lines associated with NVIDIA Blackwell GPU server systems, which is evidence of a test or deployment effort rather than proof of broad autonomous factory operation. 43
Skild’s financial backing has helped make its approach one of the most closely watched efforts in physical AI. Bloomberg reported that the company raised about $1.4 billion in a SoftBank-led Series C in January 2026 at a valuation above $14 billion. 13 Skild’s earlier Series A, announced in July 2024, was $300 million at a reported $1.5 billion valuation.
9
The “ChatGPT moment” analogy captures the hoped-for change in user experience: instead of programming or retraining a robot for every task, a worker could demonstrate the desired behavior and let a general model handle the translation into actions. S1 is a notable step toward that interface.
It is not yet proof that general-purpose robots have arrived. The strongest public results are company-reported, the task set remains limited, and independently verified industrial metrics are not available in the supplied evidence. The more measured conclusion is that S1 demonstrates a promising way to reduce task-specific robot training: use a video as an instruction, rely on broad pretraining to supply reusable skills, and test whether those skills can be composed into a new long-horizon task.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Skild AI says S1 can execute previously unseen, multi step tasks lasting up to 10 minutes after watching one human video, without fine tuning; its reported 66% step success rate on unseen tasks exceeded a language pro...
Skild AI says S1 can execute previously unseen, multi step tasks lasting up to 10 minutes after watching one human video, without fine tuning; its reported 66% step success rate on unseen tasks exceeded a language pro... S1 treats the video as an instruction at inference time, combining skills learned during pretraining into actions for the robot’s current environment.
The broader Skild Brain strategy trains across many robot bodies and simulated variants, but public evidence still does not establish production level completion rates, uptime, or return on investment.