Vibeleaderboard
← Back to Vibers
StepFun
Builder

StepFun

1 Tool · 26 Intel

Is this you? Sign in with X to claim this profile.

Tools

Step-3.5-Flash(huggingface.co)

Apache 2.0 conversational text generation model with 11B active parameters, servable via vLLM or SGLang.

Developer ToolsOpen SourceAI Modelsbuilt by @StepFun_ai

Intel

StepFun released the Step 3.5 Flash pretrained base and a midtrain checkpoint aimed at code, agents and long context, plus the SteptronOSS training code and, later, the SFT dataset. The post training pipeline is reproducible, not just the final weights.

AI Toolsbuilt by @StepFun_ai

StepFun's 2603 refresh of Step 3.5 Flash adds reasoning effort control: a low mode that spends 56% fewer tokens on routine calls and a high mode that keeps full reasoning while running 14% leaner. Live for Step Plan users.

AI Agentsbuilt by @StepFun_ai

Kilo, the open-source coding agent, has turned on StepFun's Step 3.7 Flash at no cost, aimed at multi-step orchestration and tool use across a real codebase rather than quick replies. A free path to try an agent-oriented open model inside an editor.

AI Toolsbuilt by @StepFun_ai

NextStep-1 drops vector quantization from autoregressive image generation, pairing a causal transformer with a lightweight flow matching head to predict continuous tokens, and reports quality comparable to Flux and SD3.5. Accepted as an ICLR 2026 oral, with code and weights.

AI Toolsbuilt by @StepFun_ai

A deployment recipe for self-hosting the open-weight Step 3.7 Flash: SGLang serving on Modal's serverless GPUs, eight H100s, Modal Volumes for weight caching, and an OpenAI-compatible chat completions endpoint in front.

Developer Toolsbuilt by @StepFun_ai

NextStep-1.1 reworks StepFun's autoregressive image generation series, using extended training with flow-based reinforcement learning for higher fidelity and addressing numerical instability in autoregressive RL. Weights and code are published on Hugging Face and GitHub.

AI Toolsbuilt by @StepFun_ai

An experimental vLLM plugin applies Attention-FFN disaggregation to mixture-of-experts serving, running the attention path and the expert path as separately scheduled workloads. Contributed by the Ascend and vLLM teams with StepFun, Ant Group, and FastAFD.

Developer Toolsbuilt by @StepFun_ai

StepFun's ASR update takes 30 minutes of continuous audio in one pass, decodes 4x faster than the previous generation and prices the API at $0.022 per hour, an 80% cut in inference cost, with bilingual English and Chinese accuracy.

AI Toolsbuilt by @StepFun_ai

StepFun put Step 3.5 Flash behind a monthly subscription in four tiers from $6.99 to $99, usable from Cursor, Windsurf, Cline or your own stack. Paid channels, including the OpenRouter paid API, now get a dedicated lane above 100 tokens per second.

AI Toolsbuilt by @StepFun_ai

Fireworks now hosts Step 3.7 Flash, a 198B sparse mixture-of-experts vision-language model with a 196B language backbone and a 1.8B vision encoder. Multi-token-prediction decoding pushes output to about 400 tokens per second on agent workloads.

AI Toolsbuilt by @StepFun_ai

Step-Audio-R1.1 is an open-weight speech model that allocates test-time compute to audio reasoning with a chain-of-thought scheme built for audio understanding. StepFun reports 96.4% on BigBench Audio and 1.51s time to first audio, published on Hugging Face and ModelScope with a live demo.

AI Toolsbuilt by @StepFun_ai

Step-Audio-EditX adds paralinguistic tags, finer speech rate control and stronger emotion fidelity. The release includes SFT, DPO and GRPO training code with vLLM inference, making the model tunable on your own data rather than fixed at the shipped checkpoint.

AI Toolsbuilt by @StepFun_ai

ACE-Step 1.5 XL is a 4B DiT decoder for music generation in three variants, running from 12GB VRAM with INT8 offload up to 24GB for full quality. The MIT license and commercially safe training data matter here as much as the audio quality.

AI Toolsbuilt by @StepFun_ai

StepAudio 2.5 TTS takes direction in plain language: describe the emotion, pacing and pauses you want instead of assembling tags or preset combinations. Zero shot voice cloning keeps timbre and emotion controllable separately. Available pay as you go or under Step Plan.

AI Toolsbuilt by @StepFun_ai

StepFun reports Step 3.5 Flash at the top of MathArena with 96.11% overall and 97% on AIME 2026 I, at roughly $0.40 per run from a model activating 11B parameters. MathArena is run by ETH Zurich's SRI Lab and INSAIT on uncontaminated competition sets, which makes the cost-per-accuracy point the nota

AI Toolsbuilt by @StepFun_ai

StepAudio 2.5 Realtime handles live voice conversation with attention to paralinguistic signal: tone, pacing, pauses and micro emotions. Personas are defined through the API, with 10,000+ native personas, five presets and RLHF tuning to hold character.

AI Toolsbuilt by @StepFun_ai

Cline has made Step 3.7 Flash free for a month, selectable through /model after installing the CLI. Cline's team places the open-weight, 256k-context model ahead of Gemini and DeepSeek flash tiers and near frontier performance on SWE-Bench, a claim worth checking on your own repos.

AI Toolsbuilt by @StepFun_ai

Step 3.7 Flash is a 198B sparse MoE with roughly 11B active parameters, vision input, 256K context and three reasoning levels, released under Apache 2.0. It reports 67.1 on ClawEval-1.1, 56.3 on SWE-PRO and above 98% on tool use, and runs on desktop class hardware.

AI Agentsbuilt by @StepFun_ai

StepFun raised the free daily request cap for Step 3.5 Flash on OpenRouter tenfold, from 1,000 to 10,000 calls per day, after usage pushed against the old ceiling.

AI Toolsbuilt by @StepFun_ai

STEP3-VL-10B is a compact open multimodal model trained fully unfrozen on 1.2T tokens, pairing a language-aligned perception encoder with a Qwen3-8B decoder, then scaled through 1k+ RL iterations and Parallel Coordinated Reasoning at inference. The full model suite is released.

AI Toolsbuilt by @StepFun_ai

StepFun released STEP3-VL-10B, an open-weight vision language model that reports beating GLM-4.6V and Qwen3-VL on MMMU, MathVision and MathVerse despite activating a fraction of their parameters, plus strong AIME, LiveCodeBench and spatial results. It comes from 1.2T token pre-training, 1,400+ RL it

AI Toolsbuilt by @StepFun_ai

StepFun opened beta access to Step-DeepResearch, a single agent research system that plans, verifies across sources and cites, backed by a 20M paper index. Reported 61.42% on research rubrics and 67.1% win/tie on ADR-Bench, with a technical report and GitHub release.

AI Agentsbuilt by @StepFun_ai

Step 3.5 Flash is a 196B sparse MoE activating 11B parameters per token, with 256K context, 100 to 300 tok/s throughput and agentic scores of 74.4% SWE-bench Verified and 51.0% Terminal-Bench 2.0. Open weights on Hugging Face, plus OpenRouter and local deployment.

AI Agentsbuilt by @StepFun_ai

Step Image Edit 2 is a 3.5B model for text to image and instruction based editing, reported first on KRIS-Bench across overall, factual and conceptual categories. It renders Chinese and English text accurately, edits in 1.6 seconds and costs $0.003 per image.

AI Toolsbuilt by @StepFun_ai

StepFun published an official ClawHub plugin, so OpenClaw v3.28 and later can point at Step models through either the standard API or a Step Plan subscription, with both the Chinese and international endpoints supported.

AI Agentsbuilt by @StepFun_ai

StepFun open sourced CoPaRe, or Coordinated Parallel Reasoning, a test time compute method that scales parallel reasoning chains rather than a single longer trace. The headline claim is an 8B model beating GPT-5 Thinking on math, with the release on Hugging Face.

AI Toolsbuilt by @StepFun_ai