StepFun AI Guide: Models, Products, Uses & API

StepFun AI models and agentic AI technology platform with multimodal and developer workflows.

StepFun AI is a Chinese artificial intelligence company developing large language models, multimodal systems, agentic AI, speech models, and developer tools. Its current lineup includes Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash, and the StepAudio 3 family. StepFun also provides an AI Studio, subscription plans, and an OpenAI-compatible API for developers.

Step 5 Preview is the company’s newest flagship model. It is designed for complex agentic work, software engineering, professional knowledge work, finance, and multimodal tasks. StepFun says Step 5 Preview has 600 billion total parameters, 27 billion active parameters per token, a 1-million-token context window, and vision input. The company currently offers it through its products and API, with an open-weight release planned for October 15, 2026.

Quick Summary

  • StepFun AI is a Chinese AI company developing language, multimodal, agentic, and speech models.
  • Step 5 Preview is its latest flagship model, focused on reasoning, coding, vision, research, finance, and long-horizon agentic tasks.
  • Step 3.7 Flash targets efficient multimodal AI agents with tool use, coding, web search, and visual understanding.
  • Step 3.5 Flash focuses on fast reasoning, coding, long-context processing, and agent workflows.
  • StepAudio 3 brings real-time voice interaction, speech recognition, voice generation, and music capabilities.
  • Developers can access StepFun through AI Studio and an OpenAI-compatible API.
  • StepFun supports use cases including AI coding, research, automation, financial analysis, document processing, multimodal applications, and voice AI.
  • The company is increasingly moving from standalone chatbots toward AI agents capable of using tools and completing multi-step tasks.
  • Step 5 Preview’s open-weight release is planned for October 15, 2026, according to StepFun.

What is StepFun AI?

StepFun AI, operated by Shanghai-based StepFun, is an AI research and product company focused on foundation models and AI applications. The company’s model development has progressed from large language models to multimodal systems, image and video generation, speech, reasoning, and agentic AI.

StepFun’s current strategy is increasingly centered on AI agents that can perform tasks rather than simply generate text. Its latest models are designed to reason, use tools, interact with software environments, work with documents and visual information, and complete longer workflows.

The company’s public platform now brings together:

  • Frontier language models
  • Agentic and coding models
  • Speech recognition and generation
  • Real-time voice interaction
  • AI Studio
  • API access
  • Subscription plans
  • Developer and startup programs

The official StepFun platform currently lists Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash, and several StepAudio 3 models.

StepFun was founded in 2023 by Jiang Daxin, a former Microsoft executive. Reuters reported in April 2026 that the company was restructuring its offshore corporate structure as it prepared for a potential Hong Kong listing, while also noting backing from investors including Tencent and Shanghai government-linked capital.

What Are the Main StepFun AI Models?

The StepFun model family has expanded considerably since the original Step-1 and Step-2 releases.

The most important current models are:

Model Main focus Key capabilities Current status
Step 5 Preview Complex work and agents Reasoning, coding, vision, finance, long-horizon tasks Current flagship preview
Step 3.7 Flash Efficient agents Multimodal understanding, tool use, coding, web/visual search Current
Step 3.5 Flash Fast reasoning and agents Coding, reasoning, tool use, long context Current
StepAudio 3 Realtime Voice agents Full-duplex conversation, reasoning while speaking, tool use Current
StepAudio 3 ASR Speech recognition Speech-to-text and audio understanding Current
StepAudio 3 Gen Voice generation Speech and voice generation Current
StepAudio 3 Music Music generation AI music creation Current
Step-2 Large language model General language, reasoning and coding Earlier generation
Step-1.5V Multimodal understanding Image and video understanding Earlier generation
Step-1X Image generation Text-to-image generation Earlier generation
Step R-mini Reasoning Planning, reasoning and reflection Earlier generation

Step 5 Preview

Step 5 Preview is currently StepFun’s flagship model for complex real-world work.

It is built using a sparse Mixture-of-Experts architecture with 600B total parameters and 27B active parameters per token. It also supports a 1M-token context window and vision input.

StepFun positions Step 5 Preview around several areas:

  • Software engineering
  • Coding agents
  • Long-running development tasks
  • Professional knowledge work
  • Financial analysis
  • Research
  • Document analysis
  • Multimodal workflows
  • Tool use
  • Computer and software interaction

One of the more important differences from traditional chat-oriented models is the emphasis on long-horizon execution.

For example, StepFun describes experiments in which Step 5 Preview was given extended periods to optimize an inference kernel, improve a post-training pipeline, and operate a game environment over thousands of interactions. These examples are intended to demonstrate sustained planning, execution, feedback, and correction rather than one-shot question answering.

Step 5 Preview also supports vision. StepFun demonstrates workflows involving screenshots, documents, spreadsheets, software interfaces, and visual development environments.

Step 3.7 Flash

Step 3.7 Flash is designed around efficient agentic execution.

It adds capabilities including:

  • Native multimodal understanding
  • Visual input
  • Web search
  • Visual search
  • Tool orchestration
  • Coding
  • GUI interaction
  • Agent frameworks
  • Long-running workflows

StepFun says Step 3.7 Flash can understand images such as product interfaces, documents, charts, and natural scenes and then use code or tools to act on that information. It is also designed to work with agent frameworks including Claude Code, KiloCode, Hermes Agent, and OpenClaw.

The model uses a sparse architecture and is positioned as an efficient option for real-world agents.

StepFun reports 196B+ total parameters and approximately 11B active parameters, with multimodal capabilities and support for local, cloud, and data-center deployment.

This makes Step 3.7 Flash particularly relevant when an application needs a model to repeatedly 

"observe → reason → use a tool → inspect the result → act again."

Step 3.5 Flash

Step 3.5 Flash preceded Step 3.7 Flash and established StepFun’s current emphasis on efficient agentic reasoning.

It uses a sparse MoE architecture with approximately 196B total parameters and 11B active parameters per token. StepFun reports a 256K context window and generation speeds of approximately 100–300 tokens per second in typical usage, with higher throughput reported for some coding tasks.

The model was specifically optimized for:

  • Coding
  • Reasoning
  • Tool use
  • Agent workflows
  • Long-context processing
  • Local deployment

StepFun reports 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0 for Step 3.5 Flash. These are vendor-reported benchmark results and should be interpreted in the context of each benchmark’s methodology rather than treated as a universal measure of model quality.

One notable technical feature is Multi-Token Prediction (MTP), which StepFun uses to increase decoding efficiency.

The model can also be deployed locally. StepFun documents INT4 GGUF deployment on high-memory systems including Mac Studio M4 Max, NVIDIA DGX Spark, and AMD Ryzen AI Max+ 395 systems.

What Is StepAudio 3?

StepAudio 3 is StepFun’s newer family of speech-focused models.

Rather than treating speech as simply another input or output format, the family is designed for applications where AI needs to listen, understand, speak, and interact in real time.

The current StepAudio 3 family includes:

  • StepAudio 3 ASR — speech recognition
  • StepAudio 3 Realtime — real-time voice interaction
  • StepAudio 3 Gen — voice generation
  • StepAudio 3 TTS — text-to-speech
  • StepAudio 3 Music — music generation

The official platform describes StepAudio 3 as a family covering speech recognition, real-time interaction, voice generation and music generation.

StepAudio 3 Realtime

StepAudio 3 Realtime is particularly focused on conversational voice agents.

StepFun describes it as supporting:

  • Full-duplex interaction
  • Natural turn-taking
  • Interruptions
  • Audio understanding
  • Emotion and conversational cues
  • Reasoning during speech
  • Tool use

The full-duplex design means a user does not necessarily need to wait for the model to completely finish speaking before interacting again.

This architecture is relevant to customer-service agents, voice assistants, call-center applications, conversational interfaces, and other systems where low-latency interaction matters.

How Has StepFun AI Evolved?

StepFun’s development can be understood in several stages.

Step-1 and Step-1V

The early Step series established StepFun’s work in large language and multimodal models.

Step-1 focused on language capabilities, while Step-1V expanded into multimodal understanding.

Step-2

Step-2 represented a major increase in scale and adopted a Mixture-of-Experts architecture.

StepFun introduced a preview of Step-2 in 2024 and later released the formal version alongside other models at the 2024 World Artificial Intelligence Conference.

Step-1.5V and Step-1X

Step-1.5V focused on multimodal understanding, including images and video.

Step-1X moved into image generation and represented StepFun’s expansion beyond language models.

Step R-mini

Step R-mini marked a shift toward dedicated reasoning models.

It was designed around the ability to plan, reason, attempt solutions, and reflect before producing an answer.

Step-Video, Step-Audio and Other Multimodal Models

StepFun subsequently expanded into:

  • Video generation
  • Speech recognition
  • Voice generation
  • Audio interaction
  • 3D generation
  • Multimodal understanding

The company’s earlier public model portfolio therefore covered considerably more than conventional chatbots.

Step 3.5 Flash and Step 3.7 Flash

The 3.x generation represents another strategic shift.

Instead of optimizing primarily for chatbot responses, StepFun increasingly emphasizes agents.

That means the model is expected to work with:

  • Tools
  • Browsers
  • Terminals
  • APIs
  • Search
  • Code execution
  • GUI environments
  • Other agents

Step 3.5 Flash introduced this direction strongly, while Step 3.7 Flash expanded multimodal and tool-oriented capabilities.

Step 5 Preview

Step 5 Preview moves further toward professional and autonomous knowledge work.

The model combines:

  • Reasoning
  • Coding
  • Vision
  • Tool use
  • Long context
  • Research
  • Financial analysis
  • Professional workflows

StepFun currently describes it as its flagship model for agentic work.

What Technology Does StepFun Use?

Mixture of Experts

Several StepFun models use Mixture-of-Experts (MoE) architectures.

Instead of activating the entire model for every token, an MoE system routes each token through a subset of specialized experts.

The practical objective is to achieve a large model’s representational capacity without paying the full computational cost of activating every parameter on every token.

Step 3.5 Flash, for example, has approximately 196B total parameters but activates around 11B per token.

Step 5 Preview similarly uses sparse MoE architecture, with 600B total parameters and 27B active parameters per token.

Long Context

Long context is increasingly important for agents because real tasks may involve:

  • Large codebases
  • Multiple documents
  • Research materials
  • Spreadsheets
  • Conversation history
  • Tool outputs
  • Search results

Step 3.5 Flash supports a 256K context window, while Step 5 Preview expands that to 1 million tokens.

Multimodal Processing

Newer StepFun models increasingly combine text with visual information.

Step 3.7 Flash, for example, is designed to interpret screenshots, documents, charts and interfaces and then act through tools.

Step 5 Preview also supports vision input.

Agentic Tool Use

A major part of StepFun’s current technology direction is tool orchestration.

An agent can potentially:

  1. Understand a user’s objective.
  2. Break the objective into smaller tasks.
  3. Search for information.
  4. Call APIs or other tools.
  5. Execute code.
  6. Inspect results.
  7. Correct mistakes.
  8. Continue until the task is complete.

This is fundamentally different from a model that only generates a text response.

What Can StepFun AI Be Used For?

StepFun’s current models have applications across both consumer and professional workflows.

1. AI Coding

Step 5 Preview and Step 3.7 Flash can be used for:

  • Code generation
  • Debugging
  • Refactoring
  • Documentation
  • Feature development
  • Testing
  • Software maintenance
  • Coding agents

StepFun’s StepCodeBench evaluation specifically covers tasks such as feature modification, bug repair, refactoring, documentation, performance tuning, code generation, CI/CD operations and environment setup.

2. AI Agents

Agentic applications are one of StepFun’s clearest areas of focus.

A developer can use a StepFun model as the reasoning layer behind an agent that interacts with tools, websites, software, databases or APIs.

Potential applications include:

  • Research agents
  • Coding agents
  • Productivity agents
  • Data-analysis agents
  • Business workflow automation
  • Customer-support agents
  • Browser agents
  • Computer-use systems

StepFun also operates an Agent Builder Program aimed at developers building agents, workflows, tools and integrations.

3. Research and Deep Research

Long-context reasoning and tool use make StepFun models suitable for research workflows.

A research agent could:

  • Search multiple sources
  • Extract information
  • Compare documents
  • Organize evidence
  • Perform calculations
  • Produce a structured report

Step 5 Preview has been tested by StepFun on large-scale research workflows, including a climate study involving 1,000 locations and hundreds of thousands of records.

4. Financial Analysis

Finance is an explicit focus of Step 5 Preview.

StepFun describes internal evaluations covering:

  • Live financial search
  • Corporate valuation
  • Financial deep research
  • Financial data analysis
  • Evidence reconciliation

The company also reports results on the external FrontierFinance benchmark.

For real financial decisions, however, model-generated analysis should still be independently checked against primary financial documents and current market data.

5. Document and Spreadsheet Analysis

StepFun’s newer models can be applied to knowledge-work tasks involving documents, tables and structured information.

Potential workflows include:

  • Summarizing reports
  • Extracting structured data
  • Comparing documents
  • Building spreadsheets
  • Financial analysis
  • Creating presentations and reports
  • Research synthesis

StepFun demonstrates a Step 5 Preview workflow producing a multi-sheet analytical workbook containing source data, formulas, charts and regional analysis.

6. Multimodal Applications

With vision-enabled models, developers can build applications that understand both text and images.

Examples include:

  • Screenshot analysis
  • UI understanding
  • Chart interpretation
  • Document processing
  • Product analysis
  • Visual search
  • Image-based research
  • Visual coding workflows

7. Voice AI

StepAudio 3 opens another category of applications.

Potential use cases include:

  • Voice assistants
  • Customer-service agents
  • Call-center automation
  • Conversational AI
  • Speech transcription
  • Voice interfaces
  • Voice generation
  • Music applications

The real-time model is especially suited to applications requiring natural conversational turn-taking and interruption handling.

How Can Developers Use StepFun AI?

Developers can access StepFun through its official Open Platform.

The platform provides an API with an OpenAI-compatible interface. StepFun’s documentation shows developers using the OpenAI SDK while changing the API base URL to StepFun’s endpoint.

A basic workflow is:

  1. Create a StepFun account.
  2. Generate an API key.
  3. Select a supported model.
  4. Use the StepFun API endpoint.
  5. Send prompts or structured requests.
  6. Process the model’s response inside your application.

The international API endpoint is:

https://api.stepfun.ai/v1

StepFun also operates a China-specific platform and endpoint.

Because the API is designed to be compatible with the OpenAI SDK, developers familiar with OpenAI-style integrations can generally adapt existing application architecture rather than learning an entirely different interface.

Does StepFun AI Have a No-Code Option?

Yes.

StepFun currently provides AI Studio, allowing users to interact with its models through a browser without building an API integration first.

This is useful for:

  • Testing prompts
  • Comparing model behavior
  • Experimenting with workflows
  • Prototyping AI applications
  • Evaluating a model before development

StepFun describes Studio as a no-code environment where users can test models before connecting them to the API.

The company also says that capabilities previously associated with its Experience Center have been migrated to AI Studio.

Is StepFun AI Free?

Some StepFun access is available through free or promotional options, but StepFun is not simply a permanently free AI service.

The current platform uses several access models:

  • AI Studio access
  • Step Plan subscriptions
  • Usage-based API billing
  • Developer programs
  • Startup programs

The API platform says users can sign up and call public models using usage-based billing.

StepFun’s Step Plan currently includes several subscription tiers, with different monthly credit allocations and benefits such as flagship-model access, smart routing, MCP tool coverage and priority API access on higher tiers.

Because subscription pricing and promotional offers can change, users should check the current official StepFun plan page before purchasing rather than relying on an older article.

What Are the StepFun Step Plan Options?

The current Step Plan structure includes:

Plan Positioning Listed monthly credits Additional benefits
Flash Mini Entry-level 400M Flagship models, smart routing, MCP tools
Flash Plus Regular use 1,600M Priority API access, priority support
Flash Pro Professional use 8,000M Priority API access, priority support
Flash Max Heavy use 40,000M Priority API access, priority support

StepFun’s current plan page also says Studio receives additional creation credits equivalent to 40% of the plan quota.

The actual monetary prices should be checked directly on the current subscription page because plan pricing and promotions can change.

How Does StepFun Compare With Traditional Chatbots?

The biggest distinction is where the product is headed.

A conventional chatbot primarily focuses on answering a user’s message.

StepFun’s newer models are increasingly designed around:

Reason → Search → Use tools → Execute → Inspect → Correct → Continue

This makes them more relevant to agentic workflows.

For example, a traditional chatbot might explain how to build a website.

An agentic StepFun workflow could potentially:

  1. Understand the website requirements.
  2. Generate the code.
  3. Run the application.
  4. Inspect the result.
  5. Identify errors.
  6. Modify the code.
  7. Test it again.
  8. Produce the final project.

Step 5 Preview’s published evaluations specifically emphasize these longer workflows and sustained execution.

Step 3.5 Flash vs Step 3.7 Flash vs Step 5 Preview

Feature Step 3.5 Flash Step 3.7 Flash Step 5 Preview
Primary focus Efficient reasoning and agents Multimodal agents Complex professional work
Total parameters ~196B ~196B +600B
Active parameters ~11B ~11B 27B
Context 256K Not highlighted as 1M 1M
Vision No Yes Yes
Coding Yes Yes Yes
Tool use Yes Yes Yes
Web/visual search Limited compared with 3.7 Yes Yes
Long-horizon work Yes Yes Strong focus
Finance General General Explicit focus
Local deployment Yes Yes Current deployment details differ
Status Current Current Preview flagship

The progression is important: Step 3.5 Flash emphasizes efficient reasoning and agent execution; Step 3.7 Flash adds stronger multimodal and visual-agent capabilities; Step 5 Preview expands toward high-end professional and long-horizon work.

The specifications above are based primarily on StepFun’s published model information and should not be interpreted as a standardized third-party benchmark comparison.

What Are the Limitations of StepFun AI?

StepFun’s models are capable, but there are several reasons not to treat benchmark results as proof that every task will work reliably.

1. Model performance varies by task

A model can perform extremely well on one benchmark while struggling with another.

Step 5 Preview’s own evaluation table demonstrates this variation across reasoning, coding, computer-use, multimodal and finance benchmarks.

2. Agent reliability is different from chat quality

An agent must repeatedly make correct decisions.

A single incorrect tool call can cause a long workflow to fail even if the underlying model is strong at ordinary question answering.

3. Current information still requires verification

Models should not automatically be treated as authoritative sources for:

  • Current financial information
  • Legal advice
  • Medical decisions
  • Business-critical data
  • Live pricing
  • Current regulations

Applications should connect models to reliable data sources and implement verification where accuracy matters.

4. Some capabilities are still evolving

Step 5 Preview is explicitly a preview model, and StepFun plans to release its open weights on October 15, 2026.

That means the capabilities, availability and deployment options can change.

Who Should Use StepFun AI?

StepFun is particularly relevant to:

i) Developers

Developers can use the API to integrate reasoning, coding, vision and speech capabilities into their applications.

ii) AI Agent Builders

The platform’s current direction strongly emphasizes agents, tool use and autonomous workflows.

iii) Startups

StepFun has a Startup Program offering selected early-stage AI teams potential API credits, early model access, ecosystem support and promotional opportunities.

iv) Researchers

Open models such as Step 3.5 Flash and Step 3.7 Flash can be relevant for researchers experimenting with local inference, agent architectures and model behavior.

v) Businesses

Organizations can explore StepFun for document processing, research, coding, finance, customer support, automation and other knowledge-work applications.

What Is the Future Direction of StepFun AI?

StepFun’s recent releases indicate a clear movement from large language models toward multimodal, agentic and professional AI systems.

The progression can be summarized as:

Language → Multimodal → Reasoning → Agents → Multimodal Agents → Professional Autonomous Work

Step 5 Preview represents the latest stage of this progression.

The model is not presented merely as a chatbot. StepFun is positioning it as an AI system capable of handling complex tasks across software engineering, research, finance and professional knowledge work.

The planned open-weight release of Step 5 Preview on October 15, 2026 could also be significant for developers and researchers interested in self-hosting and customization. Until that release occurs, however, it should be described as a planned release rather than an already available open-weight model.

Conclusion

StepFun AI has evolved from a Chinese large-language-model developer into a broader AI platform covering reasoning, coding, multimodal understanding, agents, professional knowledge work and voice AI.

Its current model lineup is led by Step 5 Preview, while Step 3.7 Flash and Step 3.5 Flash provide more efficient options for agentic and coding workloads. The StepAudio 3 family extends the ecosystem into speech recognition, real-time conversation, voice generation and music.

For developers, StepFun is also more than a collection of models. Its ecosystem now includes AI Studio, an OpenAI-compatible API, Step Plan subscriptions, developer programs and startup support.

The most important development to watch next is Step 5 Preview’s planned October 15, 2026 open-weight release, which could make its latest flagship technology available to a much wider developer and research community.

Frequently Asked Questions

1. What is StepFun AI?

StepFun AI is an AI company developing large language models, multimodal models, agentic systems and speech technologies. Its current lineup includes Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash and StepAudio 3.

2. What is the latest StepFun AI model?

As of September 21, 2026, Step 5 Preview is StepFun’s newest flagship model. It has 600B total parameters, 27B active parameters per token, a 1M-token context window and vision input.

3. Is StepFun AI open source?

Some StepFun models are open source or available with open weights, including Step 3.5 Flash and Step 3.7 Flash. StepFun says Step 5 Preview is currently available through its products and API, with an open-weight release planned for October 15, 2026.

4. Does StepFun AI have an API?

Yes. StepFun provides an OpenAI-compatible API through its Open Platform. Developers can use the StepFun API to integrate its public models into applications and workflows.

5. Is StepFun AI free?

StepFun provides different access methods, including Studio, subscription plans and usage-based API access. Some free or promotional access may be available, but pricing and promotions can change. The current Step Plan page should be checked for the latest subscription details.

6. What is StepFun mainly used for?

StepFun’s newer models are designed for coding, AI agents, research, document analysis, financial workflows, multimodal applications, automation and voice AI. StepAudio 3 extends the platform into real-time voice interaction and speech applications.

Also Read –

Step 5 Preview: StepFun’s 600B Agentic AI Model

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top