
Model release
DeepSeek releases V4.1 Flash under an MIT licence
DeepSeek released V4.1 Flash on September 10, 2026 and put it behind the deepseek-flash API name, routing the older deepseek-v4-flash names to it. DeepSeek calls it the smallest model in a new architecture family with native multimodal visual understanding, built for faster inference and higher throughput. Its pricing table lists a 1M context window, 384K max output, and standard rates of $0.30 per million input on a cache miss, $0.006 on a hit, and $1.20 output, halved off-peak.
- Why it matters
- It scores 40 on the current Intelligence Index while shipping under a permissive licence, which is an unusual combination: most models at that level cannot be self-hosted at all. The peak and off-peak split also means the advertised price depends on when your jobs run.
- Who should care
- Teams self-hosting models, and anyone running batch work that can be scheduled
- What you can do
- If your pipeline is not latency-bound, move it outside 01:00-04:00 and 06:00-10:00 UTC on weekdays and pay the off-peak rate, which is half.

