StepFun positions Step 5 Preview for work that spans large amounts of information and multiple agent actions, particularly software engineering and professional knowledge tasks. It accepts text, images and video within a one-million-token context window; StepFun recommends its smaller Flash models when speed or a pure-text workload matters more.
1
3
Release status and open weights
Step 5 Preview is listed for use through StepFun’s API. Reports say StepFun plans to release open weights on October 15, 2026—a future date as of September 25, not a completed official release.
3
32 Some third-party Hugging Face repositories already claim to offer BF16 weights, but those listings alone do not establish that StepFun has published an official checkpoint.
5
11
Architecture, context and intended uses
Step 5 Preview is described as a sparse mixture-of-experts model with 600 billion parameters in total and approximately 27 billion active per token. It supports text, image and video input and a 1M-token context window. That capacity makes it a candidate for workflows with substantial source material, though the context limit alone does not demonstrate reliable recall or task completion across a full-length run.
1
5
StepFun recommends the model for software engineering, professional knowledge work and multistep agents. Its announcement also emphasizes sustained work across software environments and finance-related tasks; these are intended strengths, not guarantees for every deployment.
1
16
Attention, training and efficiency claims
A third-party model listing attributes the long context to Sparse GQA. A separate report describes on-policy, long-horizon reinforcement learning and lists train–inference alignment for MoE routing, MTP-3 speculative decoding, FP8 MoE and KV-cache offload. It attributes a more than 3× end-to-end speedup for long-horizon RL to StepFun. These details and the speedup are reported claims, not independently established results in the available primary documentation.
5
37
The provided material does not substantiate a specific role for Step 5 Preview in generating verifiable tasks. Nor does it isolate how much any individual attention or training technique improves reasoning or tool use. Those should not be treated as confirmed capabilities or measured gains.
1
37
Price and measured performance
Artificial Analysis gives Step 5 Preview an Intelligence Index score of 44. It lists prices of $1.00 per million input tokens and $2.70 per million output tokens, and measures approximately 83 output tokens per second through StepFun’s API. That speed is a measurement under its test conditions, not a fixed rate for every prompt or provider.
21
Step 5 Preview vs. Step 3.5 Flash and Step 3.7 Flash
StepFun recommends Step 3.5 Flash for pure-text workloads and Step 3.7 Flash for high-speed multimodal reasoning. Step 3.7 Flash supports image and video understanding, offers three reasoning-effort levels and has a 256K-token context window. Its sparse MoE architecture has 198B total parameters and about 11B active per token—smaller figures than Step 5 Preview’s 600B/27B and 1M-token context.
1
2
5
The practical choice is therefore about workload, not parameter count alone: Step 5 Preview is positioned for larger-context, sustained agent tasks, while the Flash models are StepFun’s recommendations for faster multimodal work or text-only use.
1