DeepSeek V4.1 Flash: image understanding and what to check before switching
DeepSeek can now work with images as well as text. Its coding scores improved on some tests, but existing users should check which version their tools are using.
By Monolith

DeepSeek released V4.1-Flash on September 10, 2026. It can interpret images as well as text, which could help when a task involves screenshots or scanned material. Existing Flash API names, the labels software uses to request a model, now select this version. DeepSeek also changed its published plan for V4 Pro: the launch post announced retirement, but its pricing page later said Pro would remain available.12
A practical use for image understanding
Suppose a customer sends a screenshot of a broken booking form. With text-only input, someone first has to describe the screenshot. An image-capable model can receive the screenshot directly and help explain the visible problem. A developer can then check that explanation against the actual page.
A second possible trial is extracting information from a scanned document into a draft table. Keep the original beside the result and check names, dates and numbers. The value would be less retyping and a quicker first pass. A convincing-looking table is useful only when the entries match the document.
What changed versus V4
DeepSeek changed the way the model reads information and prepares a response. The technical name is a causal encoder-decoder architecture. For most business owners, the useful addition is that a request can include a picture as well as written instructions. You can point to a problem in a screenshot or ask about material in a scanned document.14
| Detail | What it means |
|---|---|
| Main model | 552 billion parameters, the learned settings inside a model. |
| Memory component | A separate 196 billion-parameter component. |
| Active processing | Only part of the system is active for each piece of input or output. These counts are not a quality score. |
DeepSeek says the new design uses less storage during processing. That may help the provider operate the service, but it does not tell you what your completed project will cost. The model accepts text and images and returns text. Its listed capacity is 1M tokens of context, with up to 384K tokens of output.12
Context means the material available while the model works on a request. A larger limit can help when several documents belong together. The important question is still whether it finds the right detail and explains where it came from.
What the comparable benchmark rows show
| Measure | V4-Flash | V4-Pro | V4.1-Flash |
|---|---|---|---|
| Terminal-Bench 2.1, Pass@1 | 82.7 | 87.9 | 90.6 |
| DeepSWE v1.1, resolved | 54.4 | 62.7 | 74.2 |
| GPQA Diamond, Pass@1 | 89.9 | 92.4 | 90.9 |
Read each row as a different kind of test. Terminal-Bench and DeepSWE test software work; GPQA Diamond tests difficult science questions. V4.1-Flash improves on the coding rows shown here, but scores below V4-Pro on GPQA Diamond. The software wrapped around the model also matters: V4.1-Flash scores 84.1 on Terminal-Bench with Codex and 90.6 with DeepSeek's Minimal setup. That surrounding software is often called a harness.4
For an agency, better software-test results might make the model worth trying on a small website repair. For a business that mostly summarizes customer notes, the same chart is less directly relevant. Choose a trial that resembles the work you need done rather than choosing from the highest number in the table.
The scores also depend on how the model was instructed to work. Here it used the maximum reasoning setting. Treat the table as a comparison under those conditions, rather than a promise for every speed or cost setting.
| Detail | What it means |
|---|---|
| Response settings | Temperature 1.0 and top-p 0.95, settings that influence variation in generated answers. |
| Excluded comparison | The official sources disagree on NL2Repo, so this article leaves it out. |
Practical uses and migration checks
A useful trial could be finding a website problem from a screenshot, fixing one known software bug or extracting information from a document image. Work on copies first. Give the reviewer the original material and a clear way to check the result. Ask your developer to record which model version handled the task, because a familiar service name can start selecting a newer version.
Existing users should have their developer check which model their software is requesting. A familiar name can start pointing to a new version, rather like a saved shortcut that opens an updated application. That change can affect responses even when your instructions stay the same.
| Detail | What it means |
|---|---|
| New model | deepseek-flash selects V4.1-Flash. |
| Older Flash names | deepseek-v4-flash and deepseek-v4-flash-vision-exp also select V4.1-Flash. |
| V4 Pro announcement | The September 10 post proposed a redirect on September 14 at 04:00 UTC. |
| Later guidance | By the September 11 check, the pricing page said V4 Pro would continue with unchanged billing. Confirm current guidance before changing a live system. |
Access, limits and cost at September 11
The service charges for the information it reads and the answer it generates. Reusing eligible stored information can have a lower rate, and the published schedule also distinguishes peak from off-peak hours. These are processing fees; they do not include the time your team spends checking the work.
| Detail | What it means |
|---|---|
| Peak | Input $0.30; cached input $0.006; output $1.20. |
| Off-peak | Input $0.15; cached input $0.003; output $0.60. |
| Peak schedule | 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. Other hours are off-peak. |
| Effective date | September 10, according to the release notice. |
The model files, also called weights, are available under the MIT license. Downloading them and operating a reliable service are separate jobs. A team considering its own installation needs to assess the equipment and support involved. For an existing hosted setup, retain a way to return to the previous working configuration if the new version fails your checks.3
Terms used in this article
| Term | Plain-language meaning |
|---|---|
| Token | A small unit of information counted by the model. Usage fees often depend on how many are read and generated. |
| Input / output | The information you send / the response the model generates. |
| Cached input | Eligible information stored for reuse, sometimes charged at a lower rate. |
| Benchmark | A defined test. Its score depends on the tasks, settings and supporting tools. |
| Context window | How much information fits into one request; it is not a guarantee of perfect recall. |
Which API name selects V4.1-Flash?
Is DeepSeek retiring V4 Pro on September 14?
Does V4.1-Flash beat V4-Pro on every test?
Read for this feature. The numbers match the markers in the text.
- DeepSeek V4.1 Flash announcementapi-docs.deepseek.com
- DeepSeek current models and pricingapi-docs.deepseek.com
- DeepSeek V4.1 Flash MIT licensehuggingface.co
- DeepSeek V4.1 Flash model cardhuggingface.co
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.