On September 10, 2026, DeepSeek introduced DeepSeek-V4.1-Flash—the smallest model in its new architecture family, with native visual understanding. The API name is deepseek-flash. The pricing page lists cache-miss input at $0.15 off-peak and $0.30 peak per million tokens, with output at $0.60 / $1.20. Those figures are USD as published. CAD is on the invoice. This desk is not inventing a Canadian list price.
deepseek-flash. Model version DeepSeek-V4.1-Flash.deepseek-v4-pro routes to V4.1-Flash at Flash rates.What DeepSeek listed for V4.1-Flash
The DeepSeek news post and the matching API docs news note are the primaries for product framing. DeepSeek calls V4.1-Flash the smallest model in its new architecture family, with native multimodal support. It is a 552B-parameter Mixture-of-Experts model with a Causal Encoder–Decoder design: 8B active parameters for input, 16B for output. Compared with the previous generation, DeepSeek says the KV cache needs 1/4 the HBM and 1/8 the SSD storage.
“Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.”DeepSeek news, Introducing DeepSeek-V4.1-Flash, September 10, 2026
V4-Flash and V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Official partners WorkBuddy (including CodeBuddy) and OpenCode are listed as fully supporting V4.1-Flash. New pricing took effect at 04:00 UTC on Sept 10, 2026.
The Hugging Face model card lists contexts up to one million tokens, MIT license on the repository and weights, and a controllable reasoning_effort integer from 1 to 100. On Terminal-Bench 2.1 at max reasoning effort, the card lists Pass@1 of 90.6. That is the HF number; this desk is not inventing other benches.
What it costs
These are DeepSeek’s published USD rates from the Models & Pricing page for deepseek-flash / DeepSeek-V4.1-Flash. Do not quote them as CAD. CAD is on the invoice. Off-peak rates are half of peak. Peak hours (pricing footnote): 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; all other hours are off-peak.
| Line | Off-peak | Peak | Notes |
|---|---|---|---|
| Cache hit input | $0.003 | $0.006 | deepseek-flash |
| Cache miss input | $0.15 | $0.30 | deepseek-flash |
| Output | $0.60 | $1.20 | deepseek-flash |
| vs flagship tiers | Already covered on this desk | Compare task fit, not a fake vendor war. | |
Context length is 1M; maximum output is 384K; concurrency limit for flash is 2500—all from the pricing page. Schedule flexible AI automation Alberta workloads off-peak when the clock allows; peak windows are narrow UTC mornings on weekdays.
What Alberta operators should do before Sept 14
Alberta work here means AI automation Alberta and private AI security. Cheap Flash API pricing is the shop path for agent and automation workloads versus flagship tiers already covered on this desk. Pick on task fit and data path, not a screenshot of someone else’s bill.
Action item: teams hard-coding deepseek-v4-pro must decide before the 04:00 UTC Sept 14 routing change. Docs also state 12:00 Beijing Time on September 14 for the same cutover. After that point, Pro requests route to V4.1-Flash and bill at Flash rates until V4.1-Pro launches. Update clients, agent configs, and runbooks now—not after the first surprise invoice.
Foreign hosted API is not private or on-prem. For regulated Alberta data—tenders, health-adjacent records, client credentials—prefer a private AI security path. If the workflow has to stay on the desk with clear ownership, that is AI consulting work, not a public paste into a third-country endpoint.
The guardrail
Price is not permission. A $0.15 off-peak cache-miss card does not put inference in Alberta, and it does not authorize shipping regulated client data offshore. For AI automation Alberta work, use deepseek-flash when the job fits a cheap multimodal Flash path, migrate Pro hard-codes before Sept 14 UTC, and keep sensitive workloads on a private agents path.
deepseek-flash) as a low-cost API option for AI automation Alberta agent and automation workloads—cache-miss input $0.15 off-peak / $0.30 peak, output $0.60 / $1.20, USD as published. Inventory any hard-coded deepseek-v4-pro before 04:00 UTC Sept 14, when Pro routes to Flash at Flash rates. Foreign hosted API ≠ private/on-prem; for regulated Alberta data prefer private agents and the Opcelerate security path. Compare against flagship tiers already covered on this desk on task fit, not invented CAD list prices.