Skip to content
Source
Skild AI — Blog·· 48 days agoSignalEditorial score85

Introducing S1: In-Context Learning for Robotics

Introducing S1: In-Context Learning for Robotics

Summary

Skild AI introduces S1, a robotic foundation model that leverages in-context learning to execute complex, long-horizon tasks without post-training. This marks a significant shift from traditional fine-tuning approaches, enabling rapid deployment and reducing data requirements. The model demonstrates strong performance on unseen tasks, including plant potting, pancake cooking, and kit assembly, and shows robustness to perturbations and common-sense reasoning.

Editorial context

Skild AI introduces S1, a robotic foundation model that leverages in-context learning to execute complex, long-horizon tasks without post-training. This marks a significant shift from traditional fine-tuning approaches, enabling rapid deployment and reducing data requirements. The model demonstrates strong performance on unseen tasks, including plant potting, pancake cooking, and kit assembly, and

Evidence and limits

Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.

Environment:
Controlled physical settingSource E1Source E2Source E3
Control:
Reported autonomousSource E1Source E2Source E3
Data origin:
Human egocentric dataSource E1Source E2Source E3
Reported quantities Scroll across to read all columns.
MetricValue / unitBasis / contextEvidence
hours50 hoursUnique elapsed hours

teleoperation

Source wording: “To collect these 380 long-horizon demos (4-10 minutes long), it took 50-100 hours of teleoperation.”

Source E4
percent66 percentSummed sensor-hours

success rate

Source wording: “The in-context learning policy is flat because it never trains — it is given one demonstration in its prompt. For long-horizon tasks (over four minutes), collecting 380 demonstrations takes 50–100 hours of teleoperation. A single demonstration in context is worth roughly 380 post-training examples (the exact crossing was estimated by interpolating between measured points).”

Source E4
percent96 percentSummed sensor-hours

success rate

Source wording: “S1's ICL outperforms conventional VLAs as we scale pretraining to achieve 96% accuracy, which is very high for long-horizon tasks.”

Source E4
trajectories380 trajectoriesTrajectory count

teleoperation

Source wording: “For long-horizon tasks (over four minutes), collecting 380 demonstrations takes 50–100 hours of teleoperation.”

Source E4
percent86 percentSummed sensor-hours

success rate

Source wording: “Post-training does pass it eventually — achieving an 86% success rate with 2,000 demonstrations — but we expect this gap to shrink as we continue to scale ICL pre-training and/or perform ICL post-training.”

Source E4
  • dataset: Not reported
Source excerpts and review record

No manual editorial approval recorded.

Original source quotation: “The conventional robot learning pipeline. Every new task requires data collection and fine-tuning.”

Source E1

Original source quotation: “S1 is built on NVIDIA AI infrastructure, which provides the accelerated computing foundation needed to train at scale across our diverse mix of robotics data.”

Source E2

Original source quotation: “S1 performs tasks that run up to ten minutes and do not appear in the training data.”

Source E3

Original source quotation: “For long-horizon tasks (over four minutes), collecting 380 demonstrations takes 50–100 hours of teleoperation.”

Source E4

Source:Skild AI — Blog · skild.ai