Skip to this page
THE SUPER INTELLIGENCE TIMESBY OPCELERATE NEURAL RSS
← Back to The Super Intelligence Times
Models Desk / DeepSeek / Source-backed briefing / 2026-09-10
← Back to The Super Intelligence Times
The Super Intelligence Times
Source Notes Desk
An emerald-lit processor connected by light trails to a dark server corridor.
Editorial illustration · AI-generated
Models Desk / DeepSeek / Source-backed briefing / 2026-09-10

DeepSeek V4.1-Flash Is $0.15 Off-Peak Input. Pro Routes To Flash Sept 14.

DeepSeek’s Sep 10 note puts DeepSeek-V4.1-Flash live as deepseek-flash—552B MoE, native multimodal, 1M context. Cache-miss input is $0.15 off-peak / $0.30 peak per million tokens. From 04:00 UTC Sept 14, deepseek-v4-pro routes to Flash at Flash rates.

Quick answerDeepSeek-V4.1-Flash launched Sep 10, 2026. Call it deepseek-flash. Cache-miss input: $0.15 off-peak / $0.30 peak per 1M; output $0.60 / $1.20. Off-peak is 50% of peak. Starting 04:00 UTC Sept 14, all deepseek-v4-pro requests route to V4.1-Flash at Flash rates until V4.1-Pro launches. USD as published; CAD on the invoice.

On September 10, 2026, DeepSeek introduced DeepSeek-V4.1-Flash—the smallest model in its new architecture family, with native visual understanding. The API name is deepseek-flash. The pricing page lists cache-miss input at $0.15 off-peak and $0.30 peak per million tokens, with output at $0.60 / $1.20. Those figures are USD as published. CAD is on the invoice. This desk is not inventing a Canadian list price.

AI automation Albertaprivate AI securityAI consulting Alberta
API namedeepseek-flash. Model version DeepSeek-V4.1-Flash.
Architecture552B-parameter MoE. Causal Encoder–Decoder: 8B active input, 16B output.
Cache-miss input$0.15 off-peak / $0.30 peak per 1M tokens. USD.
Output$0.60 off-peak / $1.20 peak per 1M tokens. Off-peak = 50% of peak.
Context / output1M context length. Max output 384K. Concurrency limit 2500.
Pro routingFrom 04:00 UTC Sept 14, 2026, deepseek-v4-pro routes to V4.1-Flash at Flash rates.

What DeepSeek listed for V4.1-Flash

The DeepSeek news post and the matching API docs news note are the primaries for product framing. DeepSeek calls V4.1-Flash the smallest model in its new architecture family, with native multimodal support. It is a 552B-parameter Mixture-of-Experts model with a Causal Encoder–Decoder design: 8B active parameters for input, 16B for output. Compared with the previous generation, DeepSeek says the KV cache needs 1/4 the HBM and 1/8 the SSD storage.

“Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.”DeepSeek news, Introducing DeepSeek-V4.1-Flash, September 10, 2026

V4-Flash and V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Official partners WorkBuddy (including CodeBuddy) and OpenCode are listed as fully supporting V4.1-Flash. New pricing took effect at 04:00 UTC on Sept 10, 2026.

The Hugging Face model card lists contexts up to one million tokens, MIT license on the repository and weights, and a controllable reasoning_effort integer from 1 to 100. On Terminal-Bench 2.1 at max reasoning effort, the card lists Pass@1 of 90.6. That is the HF number; this desk is not inventing other benches.

What it costs

These are DeepSeek’s published USD rates from the Models & Pricing page for deepseek-flash / DeepSeek-V4.1-Flash. Do not quote them as CAD. CAD is on the invoice. Off-peak rates are half of peak. Peak hours (pricing footnote): 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; all other hours are off-peak.

USD per 1M tokens, DeepSeek pricing page. Effective 04:00 UTC Sept 10, 2026. CAD on the invoice.
LineOff-peakPeakNotes
Cache hit input$0.003$0.006deepseek-flash
Cache miss input$0.15$0.30deepseek-flash
Output$0.60$1.20deepseek-flash
vs flagship tiersAlready covered on this deskCompare task fit, not a fake vendor war.

Context length is 1M; maximum output is 384K; concurrency limit for flash is 2500—all from the pricing page. Schedule flexible AI automation Alberta workloads off-peak when the clock allows; peak windows are narrow UTC mornings on weekdays.

What Alberta operators should do before Sept 14

Alberta work here means AI automation Alberta and private AI security. Cheap Flash API pricing is the shop path for agent and automation workloads versus flagship tiers already covered on this desk. Pick on task fit and data path, not a screenshot of someone else’s bill.

Action item: teams hard-coding deepseek-v4-pro must decide before the 04:00 UTC Sept 14 routing change. Docs also state 12:00 Beijing Time on September 14 for the same cutover. After that point, Pro requests route to V4.1-Flash and bill at Flash rates until V4.1-Pro launches. Update clients, agent configs, and runbooks now—not after the first surprise invoice.

Foreign hosted API is not private or on-prem. For regulated Alberta data—tenders, health-adjacent records, client credentials—prefer a private AI security path. If the workflow has to stay on the desk with clear ownership, that is AI consulting work, not a public paste into a third-country endpoint.

The guardrail

Price is not permission. A $0.15 off-peak cache-miss card does not put inference in Alberta, and it does not authorize shipping regulated client data offshore. For AI automation Alberta work, use deepseek-flash when the job fits a cheap multimodal Flash path, migrate Pro hard-codes before Sept 14 UTC, and keep sensitive workloads on a private agents path.

Opcelerate recommendationTreat DeepSeek-V4.1-Flash (deepseek-flash) as a low-cost API option for AI automation Alberta agent and automation workloads—cache-miss input $0.15 off-peak / $0.30 peak, output $0.60 / $1.20, USD as published. Inventory any hard-coded deepseek-v4-pro before 04:00 UTC Sept 14, when Pro routes to Flash at Flash rates. Foreign hosted API ≠ private/on-prem; for regulated Alberta data prefer private agents and the Opcelerate security path. Compare against flagship tiers already covered on this desk on task fit, not invented CAD list prices.