DeepSeek introduced V4 Flash Vision Exp on August 21, marking its first model capable of processing image inputs. This new capability allows the model to interpret charts, screenshots, and photo documents, similar to how it handles text. The model became available on API gateways like OpenRouter by August 27.
DeepSeek V4 Flash Vision Exp integrates image understanding into the existing V4 Flash model, retaining its low pricing structure of $0.22 per million input tokens and $0.66 per million output tokens, with prices doubling during peak weekday hours. This positions it as a budget-friendly option. In comparison, Google's Gemini 3.7 Flash, released on August 13, is priced at $0.75 and $3.75 per million on OpenRouter, serving as a common benchmark for budget vision tasks.
DeepSeek promotes its new model for document and chart understanding, alongside visual question answering. Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet." Both companies make strong claims regarding their models' capabilities, prompting a direct comparison to assess their practical utility for image input tasks.
To evaluate the models, three real-world back-office scenarios were simulated: chart reading, invoice auditing, and incident diagnosis. The chart reading test involved a stacked bar chart with a dual y-axis to challenge interpretation. The invoice audit included three deliberate errors in a vendor invoice. The incident diagnosis task required identifying the root cause of a payment service crash from production logs. Both models were tested using identical images and prompts via OpenRouter, with accuracy, token usage, cost, and speed recorded for each call.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
DeepSeek released V4 Flash Vision Exp, its first vision model, which is compared against Google's Gemini 3.7 Flash for image input capabilities. The comparison focuses on performance, cost, and speed for tasks like chart reading, invoice auditing, and incident diagnosis. This analysis helps developers choose between two budget-friendly vision models based on specific application needs.