Introducing S1: In-Context Learning for Robotics
Introducing S1: In-Context Learning for Robotics
Skild AI introduces S1, a robotic foundation model that leverages in-context learning to execute complex, long-horizon tasks without post-training. This marks a significant shift from traditional fine-tuning approaches, enabling rapid deployment and reducing data requirements. The model demonstrates strong performance on unseen tasks, including plant potting, pancake cooking, and kit assembly, and shows robustness to perturbations and common-sense reasoning.
Skild AI introduces S1, a robotic foundation model that leverages in-context learning to execute complex, long-horizon tasks without post-training. This marks a significant shift from traditional fine-tuning approaches, enabling rapid deployment and reducing data requirements. The model demonstrates strong performance on unseen tasks, including plant potting, pancake cooking, and kit assembly, and
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
| Metric | Value / unit | Basis / context | Evidence |
|---|---|---|---|
| hours | 50 hours | Unique elapsed hours teleoperation Source wording: “To collect these 380 long-horizon demos (4-10 minutes long), it took 50-100 hours of teleoperation.” | Source E4 |
| percent | 66 percent | Summed sensor-hours success rate Source wording: “The in-context learning policy is flat because it never trains — it is given one demonstration in its prompt. For long-horizon tasks (over four minutes), collecting 380 demonstrations takes 50–100 hours of teleoperation. A single demonstration in context is worth roughly 380 post-training examples (the exact crossing was estimated by interpolating between measured points).” | Source E4 |
| percent | 96 percent | Summed sensor-hours success rate Source wording: “S1's ICL outperforms conventional VLAs as we scale pretraining to achieve 96% accuracy, which is very high for long-horizon tasks.” | Source E4 |
| trajectories | 380 trajectories | Trajectory count teleoperation Source wording: “For long-horizon tasks (over four minutes), collecting 380 demonstrations takes 50–100 hours of teleoperation.” | Source E4 |
| percent | 86 percent | Summed sensor-hours success rate Source wording: “Post-training does pass it eventually — achieving an 86% success rate with 2,000 demonstrations — but we expect this gap to shrink as we continue to scale ICL pre-training and/or perform ICL post-training.” | Source E4 |
- dataset: Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “The conventional robot learning pipeline. Every new task requires data collection and fine-tuning.”
Source E1
Original source quotation: “S1 is built on NVIDIA AI infrastructure, which provides the accelerated computing foundation needed to train at scale across our diverse mix of robotics data.”
Source E2
Original source quotation: “S1 performs tasks that run up to ten minutes and do not appear in the training data.”
Source E3
Original source quotation: “For long-horizon tasks (over four minutes), collecting 380 demonstrations takes 50–100 hours of teleoperation.”
Source E4
Source:Skild AI — Blog · skild.ai