llmrelay
// other

DeepSeek V4 Flash

Previous Flash id. Still callable here. DeepSeek retired it; new work should use V4.1 Flash.

DeepSeek V4 Flash is the previous Flash id. DeepSeek retired the weights on 10 September 2026 and temporarily routes first-party calls named deepseek-v4-flash to V4.1 Flash. On llmrelay this id still works, at the same $0.19 / $0.54, mapped to the current upstream model. For new work, use deepseek-v4.1-flash.

This page stays up because the old slug already ranks and because pinned configs matter. If you already validated output against deepseek-v4-flash, keep calling it. If you are starting a Flash job, the current id is deepseek-v4.1-flash.

Same 1M context and 384k max output as V4.1 Flash. We serve it through a third-party hosted deployment, not DeepSeek’s first-party API.

Previous generation

DeepSeek V4 Flash still runs on llmrelay at the rates below, and we keep selling it — pinned versions matter when you have output you do not want to re-validate. For new work, DeepSeek V4.1 Flash is the current pick at $0.19/M input.

Official list
llmrelay
Input / M tokens
$0.15
$0.19
Output / M tokens
$0.60
$0.54
Context window
1,000,000 tokens
Max output
384,000 tokens

What it costs you per month

Real-world budget scenarios. Numbers are simple sums — official list price vs llmrelay's 50% off tier.

Usage scenario
Official
llmrelay
Light coding (1M in / 200K out per month)
$0.27
$0.30save $-0.03
Heavy Cursor / Cline user (50M in / 5M out per month)
$10.50
$12.20save $-1.70
Production RAG (500M in / 20M out per month)
$87.00
$105.80save $-18.80

What it's good at

Best for

Pick DeepSeek V4 Flash when

  • +You already validated a pipeline against deepseek-v4-flash and do not want to change the model string.
  • +A contract or eval set names this exact id. Switching snapshots is a new test, not a free upgrade.
  • +You need the previous Flash id on the same prepaid key as Opus 5 and GPT-6 Astra.

Choose something else when

  • !New work. V4.1 Flash is the current Flash at the same $0.19/$0.54. Id: deepseek-v4.1-flash.
  • !You typed this URL looking for the 09-08 expires-on-0910 preview. That name is dead. Use deepseek-v4.1-flash.
  • !You need DeepSeek’s first-party API or native vision billed as DeepSeek. This is a third-party text deployment.
  • !Hard reasoning where being wrong is expensive. Use Claude Opus 5 or GPT-6 Astra.

Questions people ask about DeepSeek V4 Flash

How much does DeepSeek V4 Flash cost?

We charge a flat $0.19 per million input tokens and $0.54 per million output, the same as V4.1 Flash. DeepSeek’s current Flash list is $0.15/$0.60 off-peak and $0.30/$1.20 at peak. We do not currently pass cache-hit discounts through.

Should I use V4 Flash or V4.1 Flash?

V4.1 Flash for new work: deepseek-v4.1-flash. Use this id only to keep a config you already validated. Both cost $0.19 / $0.54 here and hit the same upstream model.

Is this the expires-on-0910 preview?

No. That community id died on 10 September 2026. DeepSeek GA’d V4.1 Flash the same day. Call deepseek-v4.1-flash.

What is the context window on DeepSeek V4 Flash?

1,000,000 tokens input, with up to 384,000 output tokens — the same window DeepSeek publishes for V4.1 Flash.

Try DeepSeek V4 Flash at half the price

Free to create an account, no subscription, $10 minimum top-up — enough to run this model against a real task and judge quality yourself.

Get API key →

Compare DeepSeek V4 Flash against the alternatives

The comparisons this model appears in, the models nearest it on price, and the full rate card.