StepFun AI is a Chinese artificial intelligence company developing large language models, multimodal systems, agentic AI, speech models, and developer tools. Its current lineup includes Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash, and the StepAudio 3 family. StepFun also provides an AI Studio, subscription plans, and an OpenAI-compatible API for developers.
Step 5 Preview is the company’s newest flagship model. It is designed for complex agentic work, software engineering, professional knowledge work, finance, and multimodal tasks. StepFun says Step 5 Preview has 600 billion total parameters, 27 billion active parameters per token, a 1-million-token context window, and vision input. The company currently offers it through its products and API, with an open-weight release planned for October 15, 2026.
Quick Summary
- StepFun AI is a Chinese AI company developing language, multimodal, agentic, and speech models.
- Step 5 Preview is its latest flagship model, focused on reasoning, coding, vision, research, finance, and long-horizon agentic tasks.
- Step 3.7 Flash targets efficient multimodal AI agents with tool use, coding, web search, and visual understanding.
- Step 3.5 Flash focuses on fast reasoning, coding, long-context processing, and agent workflows.
- StepAudio 3 brings real-time voice interaction, speech recognition, voice generation, and music capabilities.
- Developers can access StepFun through AI Studio and an OpenAI-compatible API.
- StepFun supports use cases including AI coding, research, automation, financial analysis, document processing, multimodal applications, and voice AI.
- The company is increasingly moving from standalone chatbots toward AI agents capable of using tools and completing multi-step tasks.
- Step 5 Preview’s open-weight release is planned for October 15, 2026, according to StepFun.
What is StepFun AI?
StepFun AI, operated by Shanghai-based StepFun, is an AI research and product company focused on foundation models and AI applications. The company’s model development has progressed from large language models to multimodal systems, image and video generation, speech, reasoning, and agentic AI.
StepFun’s current strategy is increasingly centered on AI agents that can perform tasks rather than simply generate text. Its latest models are designed to reason, use tools, interact with software environments, work with documents and visual information, and complete longer workflows.
The company’s public platform now brings together:
- Frontier language models
- Agentic and coding models
- Speech recognition and generation
- Real-time voice interaction
- AI Studio
- API access
- Subscription plans
- Developer and startup programs
The official StepFun platform currently lists Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash, and several StepAudio 3 models.
StepFun was founded in 2023 by Jiang Daxin, a former Microsoft executive. Reuters reported in April 2026 that the company was restructuring its offshore corporate structure as it prepared for a potential Hong Kong listing, while also noting backing from investors including Tencent and Shanghai government-linked capital.
What Are the Main StepFun AI Models?
The StepFun model family has expanded considerably since the original Step-1 and Step-2 releases.
The most important current models are:
| Model | Main focus | Key capabilities | Current status |
|---|---|---|---|
| Step 5 Preview | Complex work and agents | Reasoning, coding, vision, finance, long-horizon tasks | Current flagship preview |
| Step 3.7 Flash | Efficient agents | Multimodal understanding, tool use, coding, web/visual search | Current |
| Step 3.5 Flash | Fast reasoning and agents | Coding, reasoning, tool use, long context | Current |
| StepAudio 3 Realtime | Voice agents | Full-duplex conversation, reasoning while speaking, tool use | Current |
| StepAudio 3 ASR | Speech recognition | Speech-to-text and audio understanding | Current |
| StepAudio 3 Gen | Voice generation | Speech and voice generation | Current |
| StepAudio 3 Music | Music generation | AI music creation | Current |
| Step-2 | Large language model | General language, reasoning and coding | Earlier generation |
| Step-1.5V | Multimodal understanding | Image and video understanding | Earlier generation |
| Step-1X | Image generation | Text-to-image generation | Earlier generation |
| Step R-mini | Reasoning | Planning, reasoning and reflection | Earlier generation |
Step 5 Preview
Step 5 Preview is currently StepFun’s flagship model for complex real-world work.
It is built using a sparse Mixture-of-Experts architecture with 600B total parameters and 27B active parameters per token. It also supports a 1M-token context window and vision input.
StepFun positions Step 5 Preview around several areas:
- Software engineering
- Coding agents
- Long-running development tasks
- Professional knowledge work
- Financial analysis
- Research
- Document analysis
- Multimodal workflows
- Tool use
- Computer and software interaction
One of the more important differences from traditional chat-oriented models is the emphasis on long-horizon execution.
For example, StepFun describes experiments in which Step 5 Preview was given extended periods to optimize an inference kernel, improve a post-training pipeline, and operate a game environment over thousands of interactions. These examples are intended to demonstrate sustained planning, execution, feedback, and correction rather than one-shot question answering.
Step 5 Preview also supports vision. StepFun demonstrates workflows involving screenshots, documents, spreadsheets, software interfaces, and visual development environments.
Step 3.7 Flash
Step 3.7 Flash is designed around efficient agentic execution.
It adds capabilities including:
- Native multimodal understanding
- Visual input
- Web search
- Visual search
- Tool orchestration
- Coding
- GUI interaction
- Agent frameworks
- Long-running workflows
StepFun says Step 3.7 Flash can understand images such as product interfaces, documents, charts, and natural scenes and then use code or tools to act on that information. It is also designed to work with agent frameworks including Claude Code, KiloCode, Hermes Agent, and OpenClaw.
The model uses a sparse architecture and is positioned as an efficient option for real-world agents.
StepFun reports 196B+ total parameters and approximately 11B active parameters, with multimodal capabilities and support for local, cloud, and data-center deployment.
This makes Step 3.7 Flash particularly relevant when an application needs a model to repeatedly
"observe → reason → use a tool → inspect the result → act again."
Step 3.5 Flash
Step 3.5 Flash preceded Step 3.7 Flash and established StepFun’s current emphasis on efficient agentic reasoning.
It uses a sparse MoE architecture with approximately 196B total parameters and 11B active parameters per token. StepFun reports a 256K context window and generation speeds of approximately 100–300 tokens per second in typical usage, with higher throughput reported for some coding tasks.
The model was specifically optimized for:
- Coding
- Reasoning
- Tool use
- Agent workflows
- Long-context processing
- Local deployment
StepFun reports 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0 for Step 3.5 Flash. These are vendor-reported benchmark results and should be interpreted in the context of each benchmark’s methodology rather than treated as a universal measure of model quality.
One notable technical feature is Multi-Token Prediction (MTP), which StepFun uses to increase decoding efficiency.
The model can also be deployed locally. StepFun documents INT4 GGUF deployment on high-memory systems including Mac Studio M4 Max, NVIDIA DGX Spark, and AMD Ryzen AI Max+ 395 systems.
What Is StepAudio 3?
StepAudio 3 is StepFun’s newer family of speech-focused models.
Rather than treating speech as simply another input or output format, the family is designed for applications where AI needs to listen, understand, speak, and interact in real time.
The current StepAudio 3 family includes:
- StepAudio 3 ASR — speech recognition
- StepAudio 3 Realtime — real-time voice interaction
- StepAudio 3 Gen — voice generation
- StepAudio 3 TTS — text-to-speech
- StepAudio 3 Music — music generation
The official platform describes StepAudio 3 as a family covering speech recognition, real-time interaction, voice generation and music generation.
StepAudio 3 Realtime
StepAudio 3 Realtime is particularly focused on conversational voice agents.
StepFun describes it as supporting:
- Full-duplex interaction
- Natural turn-taking
- Interruptions
- Audio understanding
- Emotion and conversational cues
- Reasoning during speech
- Tool use
The full-duplex design means a user does not necessarily need to wait for the model to completely finish speaking before interacting again.
This architecture is relevant to customer-service agents, voice assistants, call-center applications, conversational interfaces, and other systems where low-latency interaction matters.
How Has StepFun AI Evolved?
StepFun’s development can be understood in several stages.
Step-1 and Step-1V
The early Step series established StepFun’s work in large language and multimodal models.
Step-1 focused on language capabilities, while Step-1V expanded into multimodal understanding.
Step-2
Step-2 represented a major increase in scale and adopted a Mixture-of-Experts architecture.
StepFun introduced a preview of Step-2 in 2024 and later released the formal version alongside other models at the 2024 World Artificial Intelligence Conference.
Step-1.5V and Step-1X
Step-1.5V focused on multimodal understanding, including images and video.
Step-1X moved into image generation and represented StepFun’s expansion beyond language models.
Step R-mini
Step R-mini marked a shift toward dedicated reasoning models.
It was designed around the ability to plan, reason, attempt solutions, and reflect before producing an answer.
Step-Video, Step-Audio and Other Multimodal Models
StepFun subsequently expanded into:
- Video generation
- Speech recognition
- Voice generation
- Audio interaction
- 3D generation
- Multimodal understanding
The company’s earlier public model portfolio therefore covered considerably more than conventional chatbots.
Step 3.5 Flash and Step 3.7 Flash
The 3.x generation represents another strategic shift.
Instead of optimizing primarily for chatbot responses, StepFun increasingly emphasizes agents.
That means the model is expected to work with:
- Tools
- Browsers
- Terminals
- APIs
- Search
- Code execution
- GUI environments
- Other agents
Step 3.5 Flash introduced this direction strongly, while Step 3.7 Flash expanded multimodal and tool-oriented capabilities.
Step 5 Preview
Step 5 Preview moves further toward professional and autonomous knowledge work.
The model combines:
- Reasoning
- Coding
- Vision
- Tool use
- Long context
- Research
- Financial analysis
- Professional workflows
StepFun currently describes it as its flagship model for agentic work.
What Technology Does StepFun Use?
Mixture of Experts
Several StepFun models use Mixture-of-Experts (MoE) architectures.
Instead of activating the entire model for every token, an MoE system routes each token through a subset of specialized experts.
The practical objective is to achieve a large model’s representational capacity without paying the full computational cost of activating every parameter on every token.
Step 3.5 Flash, for example, has approximately 196B total parameters but activates around 11B per token.
Step 5 Preview similarly uses sparse MoE architecture, with 600B total parameters and 27B active parameters per token.
Long Context
Long context is increasingly important for agents because real tasks may involve:
- Large codebases
- Multiple documents
- Research materials
- Spreadsheets
- Conversation history
- Tool outputs
- Search results
Step 3.5 Flash supports a 256K context window, while Step 5 Preview expands that to 1 million tokens.
Multimodal Processing
Newer StepFun models increasingly combine text with visual information.
Step 3.7 Flash, for example, is designed to interpret screenshots, documents, charts and interfaces and then act through tools.
Step 5 Preview also supports vision input.
Agentic Tool Use
A major part of StepFun’s current technology direction is tool orchestration.
An agent can potentially:
- Understand a user’s objective.
- Break the objective into smaller tasks.
- Search for information.
- Call APIs or other tools.
- Execute code.
- Inspect results.
- Correct mistakes.
- Continue until the task is complete.
This is fundamentally different from a model that only generates a text response.
What Can StepFun AI Be Used For?
StepFun’s current models have applications across both consumer and professional workflows.
1. AI Coding
Step 5 Preview and Step 3.7 Flash can be used for:
- Code generation
- Debugging
- Refactoring
- Documentation
- Feature development
- Testing
- Software maintenance
- Coding agents
StepFun’s StepCodeBench evaluation specifically covers tasks such as feature modification, bug repair, refactoring, documentation, performance tuning, code generation, CI/CD operations and environment setup.
2. AI Agents
Agentic applications are one of StepFun’s clearest areas of focus.
A developer can use a StepFun model as the reasoning layer behind an agent that interacts with tools, websites, software, databases or APIs.
Potential applications include:
- Research agents
- Coding agents
- Productivity agents
- Data-analysis agents
- Business workflow automation
- Customer-support agents
- Browser agents
- Computer-use systems
StepFun also operates an Agent Builder Program aimed at developers building agents, workflows, tools and integrations.
3. Research and Deep Research
Long-context reasoning and tool use make StepFun models suitable for research workflows.
A research agent could:
- Search multiple sources
- Extract information
- Compare documents
- Organize evidence
- Perform calculations
- Produce a structured report
Step 5 Preview has been tested by StepFun on large-scale research workflows, including a climate study involving 1,000 locations and hundreds of thousands of records.
4. Financial Analysis
Finance is an explicit focus of Step 5 Preview.
StepFun describes internal evaluations covering:
- Live financial search
- Corporate valuation
- Financial deep research
- Financial data analysis
- Evidence reconciliation
The company also reports results on the external FrontierFinance benchmark.
For real financial decisions, however, model-generated analysis should still be independently checked against primary financial documents and current market data.
5. Document and Spreadsheet Analysis
StepFun’s newer models can be applied to knowledge-work tasks involving documents, tables and structured information.
Potential workflows include:
- Summarizing reports
- Extracting structured data
- Comparing documents
- Building spreadsheets
- Financial analysis
- Creating presentations and reports
- Research synthesis
StepFun demonstrates a Step 5 Preview workflow producing a multi-sheet analytical workbook containing source data, formulas, charts and regional analysis.
6. Multimodal Applications
With vision-enabled models, developers can build applications that understand both text and images.
Examples include:
- Screenshot analysis
- UI understanding
- Chart interpretation
- Document processing
- Product analysis
- Visual search
- Image-based research
- Visual coding workflows
7. Voice AI
StepAudio 3 opens another category of applications.
Potential use cases include:
- Voice assistants
- Customer-service agents
- Call-center automation
- Conversational AI
- Speech transcription
- Voice interfaces
- Voice generation
- Music applications
The real-time model is especially suited to applications requiring natural conversational turn-taking and interruption handling.
How Can Developers Use StepFun AI?
Developers can access StepFun through its official Open Platform.
The platform provides an API with an OpenAI-compatible interface. StepFun’s documentation shows developers using the OpenAI SDK while changing the API base URL to StepFun’s endpoint.
A basic workflow is:
- Create a StepFun account.
- Generate an API key.
- Select a supported model.
- Use the StepFun API endpoint.
- Send prompts or structured requests.
- Process the model’s response inside your application.
The international API endpoint is:
https://api.stepfun.ai/v1
StepFun also operates a China-specific platform and endpoint.
Because the API is designed to be compatible with the OpenAI SDK, developers familiar with OpenAI-style integrations can generally adapt existing application architecture rather than learning an entirely different interface.
Does StepFun AI Have a No-Code Option?
Yes.
StepFun currently provides AI Studio, allowing users to interact with its models through a browser without building an API integration first.
This is useful for:
- Testing prompts
- Comparing model behavior
- Experimenting with workflows
- Prototyping AI applications
- Evaluating a model before development
StepFun describes Studio as a no-code environment where users can test models before connecting them to the API.
The company also says that capabilities previously associated with its Experience Center have been migrated to AI Studio.
Is StepFun AI Free?
Some StepFun access is available through free or promotional options, but StepFun is not simply a permanently free AI service.
The current platform uses several access models:
- AI Studio access
- Step Plan subscriptions
- Usage-based API billing
- Developer programs
- Startup programs
The API platform says users can sign up and call public models using usage-based billing.
StepFun’s Step Plan currently includes several subscription tiers, with different monthly credit allocations and benefits such as flagship-model access, smart routing, MCP tool coverage and priority API access on higher tiers.
Because subscription pricing and promotional offers can change, users should check the current official StepFun plan page before purchasing rather than relying on an older article.
What Are the StepFun Step Plan Options?
The current Step Plan structure includes:
| Plan | Positioning | Listed monthly credits | Additional benefits |
|---|---|---|---|
| Flash Mini | Entry-level | 400M | Flagship models, smart routing, MCP tools |
| Flash Plus | Regular use | 1,600M | Priority API access, priority support |
| Flash Pro | Professional use | 8,000M | Priority API access, priority support |
| Flash Max | Heavy use | 40,000M | Priority API access, priority support |
StepFun’s current plan page also says Studio receives additional creation credits equivalent to 40% of the plan quota.
The actual monetary prices should be checked directly on the current subscription page because plan pricing and promotions can change.
How Does StepFun Compare With Traditional Chatbots?
The biggest distinction is where the product is headed.
A conventional chatbot primarily focuses on answering a user’s message.
StepFun’s newer models are increasingly designed around:
Reason → Search → Use tools → Execute → Inspect → Correct → Continue
This makes them more relevant to agentic workflows.
For example, a traditional chatbot might explain how to build a website.
An agentic StepFun workflow could potentially:
- Understand the website requirements.
- Generate the code.
- Run the application.
- Inspect the result.
- Identify errors.
- Modify the code.
- Test it again.
- Produce the final project.
Step 5 Preview’s published evaluations specifically emphasize these longer workflows and sustained execution.
Step 3.5 Flash vs Step 3.7 Flash vs Step 5 Preview
| Feature | Step 3.5 Flash | Step 3.7 Flash | Step 5 Preview |
|---|---|---|---|
| Primary focus | Efficient reasoning and agents | Multimodal agents | Complex professional work |
| Total parameters | ~196B | ~196B | +600B |
| Active parameters | ~11B | ~11B | 27B |
| Context | 256K | Not highlighted as 1M | 1M |
| Vision | No | Yes | Yes |
| Coding | Yes | Yes | Yes |
| Tool use | Yes | Yes | Yes |
| Web/visual search | Limited compared with 3.7 | Yes | Yes |
| Long-horizon work | Yes | Yes | Strong focus |
| Finance | General | General | Explicit focus |
| Local deployment | Yes | Yes | Current deployment details differ |
| Status | Current | Current | Preview flagship |
The progression is important: Step 3.5 Flash emphasizes efficient reasoning and agent execution; Step 3.7 Flash adds stronger multimodal and visual-agent capabilities; Step 5 Preview expands toward high-end professional and long-horizon work.
The specifications above are based primarily on StepFun’s published model information and should not be interpreted as a standardized third-party benchmark comparison.
What Are the Limitations of StepFun AI?
StepFun’s models are capable, but there are several reasons not to treat benchmark results as proof that every task will work reliably.
1. Model performance varies by task
A model can perform extremely well on one benchmark while struggling with another.
Step 5 Preview’s own evaluation table demonstrates this variation across reasoning, coding, computer-use, multimodal and finance benchmarks.
2. Agent reliability is different from chat quality
An agent must repeatedly make correct decisions.
A single incorrect tool call can cause a long workflow to fail even if the underlying model is strong at ordinary question answering.
3. Current information still requires verification
Models should not automatically be treated as authoritative sources for:
- Current financial information
- Legal advice
- Medical decisions
- Business-critical data
- Live pricing
- Current regulations
Applications should connect models to reliable data sources and implement verification where accuracy matters.
4. Some capabilities are still evolving
Step 5 Preview is explicitly a preview model, and StepFun plans to release its open weights on October 15, 2026.
That means the capabilities, availability and deployment options can change.
Who Should Use StepFun AI?
StepFun is particularly relevant to:
i) Developers
Developers can use the API to integrate reasoning, coding, vision and speech capabilities into their applications.
ii) AI Agent Builders
The platform’s current direction strongly emphasizes agents, tool use and autonomous workflows.
iii) Startups
StepFun has a Startup Program offering selected early-stage AI teams potential API credits, early model access, ecosystem support and promotional opportunities.
iv) Researchers
Open models such as Step 3.5 Flash and Step 3.7 Flash can be relevant for researchers experimenting with local inference, agent architectures and model behavior.
v) Businesses
Organizations can explore StepFun for document processing, research, coding, finance, customer support, automation and other knowledge-work applications.
What Is the Future Direction of StepFun AI?
StepFun’s recent releases indicate a clear movement from large language models toward multimodal, agentic and professional AI systems.
The progression can be summarized as:
Language → Multimodal → Reasoning → Agents → Multimodal Agents → Professional Autonomous Work
Step 5 Preview represents the latest stage of this progression.
The model is not presented merely as a chatbot. StepFun is positioning it as an AI system capable of handling complex tasks across software engineering, research, finance and professional knowledge work.
The planned open-weight release of Step 5 Preview on October 15, 2026 could also be significant for developers and researchers interested in self-hosting and customization. Until that release occurs, however, it should be described as a planned release rather than an already available open-weight model.
Conclusion
StepFun AI has evolved from a Chinese large-language-model developer into a broader AI platform covering reasoning, coding, multimodal understanding, agents, professional knowledge work and voice AI.
Its current model lineup is led by Step 5 Preview, while Step 3.7 Flash and Step 3.5 Flash provide more efficient options for agentic and coding workloads. The StepAudio 3 family extends the ecosystem into speech recognition, real-time conversation, voice generation and music.
For developers, StepFun is also more than a collection of models. Its ecosystem now includes AI Studio, an OpenAI-compatible API, Step Plan subscriptions, developer programs and startup support.
The most important development to watch next is Step 5 Preview’s planned October 15, 2026 open-weight release, which could make its latest flagship technology available to a much wider developer and research community.
Frequently Asked Questions
1. What is StepFun AI?
StepFun AI is an AI company developing large language models, multimodal models, agentic systems and speech technologies. Its current lineup includes Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash and StepAudio 3.
2. What is the latest StepFun AI model?
As of September 21, 2026, Step 5 Preview is StepFun’s newest flagship model. It has 600B total parameters, 27B active parameters per token, a 1M-token context window and vision input.
3. Is StepFun AI open source?
Some StepFun models are open source or available with open weights, including Step 3.5 Flash and Step 3.7 Flash. StepFun says Step 5 Preview is currently available through its products and API, with an open-weight release planned for October 15, 2026.
4. Does StepFun AI have an API?
Yes. StepFun provides an OpenAI-compatible API through its Open Platform. Developers can use the StepFun API to integrate its public models into applications and workflows.
5. Is StepFun AI free?
StepFun provides different access methods, including Studio, subscription plans and usage-based API access. Some free or promotional access may be available, but pricing and promotions can change. The current Step Plan page should be checked for the latest subscription details.
6. What is StepFun mainly used for?
StepFun’s newer models are designed for coding, AI agents, research, document analysis, financial workflows, multimodal applications, automation and voice AI. StepAudio 3 extends the platform into real-time voice interaction and speech applications.
Also Read –
Step 5 Preview: StepFun’s 600B Agentic AI Model


