Vibeleaderboard
← All Intel
Intel / post

Qwen3.8-Flash API goes live with 262K context and $0.016 cache reads

Source
x.com
Date
Qwen@Alibaba_Qwen

Qwen3.8-Flash on @qwen_cloud: $0.15/1M input tokens, $0.47/1M output tokens, and just $0.016/1M on cache hits. ☁️ Come give it a try! 👇 https://t.co/FQWEVpkUyt

Why it matters

Cache-hit at $0.016 per million change the arithmetic for loops that replay long every turn, and 262K native context sets the ceiling before you need the 1M extension.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
More from Qwen
Recommended reads
Comments

Checking sign-in…

Loading comments…