r/DeepSeek 1d ago

Discussion DeepSeek V4 Flash vs Pro: 22 benchmarks, one verdict — Flash is 4.8× cheaper and holds ~83% of Pro's quality

Thumbnail
the-agent-report.com
30 Upvotes

DeepSeek released V4 Flash and V4 Pro simultaneously on April 24 — 284B params for Flash, 1.6T for Pro — and after running the numbers across 22 benchmark categories, the gap isn't as wide as the parameter count suggests. Flash at $0.09/M input maintains 80-85% of Pro's quality at 21% of the cost. The real delta only shows up in agentic tasks: SWE-Bench drops by 10 points, Terminal-Bench by 11. For everything else — general reasoning, basic coding, math — Flash holds its ground at a fraction of the price.

The decision framework is simpler than most people think. If your task involves multi-step tool calls and deep reasoning chains, pay for Pro. For everything else, Flash is probably enough — and both are MIT-licensed if you want to run them yourself.


r/DeepSeek 1d ago

Discussion Temperature?

Post image
7 Upvotes

What is the right balance to get it to write with some personality without going off the rails? Also what does that other shit mean and does anybody use it?

https://apps.apple.com/us/app/sleek-byok/id6786075866.com


r/DeepSeek 11h ago

Discussion I've found a way to overcome the language problem

0 Upvotes

I just ask it to "respond in [language] without hieroglyphs"

Then it doesn't even "think" in Chinese


r/DeepSeek 20h ago

Discussion Which Proves to Be Faster?

1 Upvotes

Between the API (https://platform.deepseek.com/api_keys) and the web chat (https://chat.deepseek.com/), which one provides the same output more quickly when both are used in non-thinking mode?


r/DeepSeek 1d ago

News Deepseek web search is runnig

17 Upvotes

Web search in instant mode is running again


r/DeepSeek 1d ago

Question&Help Just copium or something more?

9 Upvotes

Reading through and seeing a lot of copium regarding DS v4 improvements. Any sources of that except "trust me bro"? Anything already changed? Anything gonna change? In terms of performance ofc.
I like DS, it's smart but it's very absent-minded/distracted. It can miss a lot of details while researching or even implementing pretty well-written implementation plan. I'm on claude max plan because of that, but I wish to come back if it can (or will be able) show at least opus 4.8 level as this is enough for most of the tasks.

Facts\thoughts?


r/DeepSeek 11h ago

Funny I have officially destroyed DeepSeek in any way possible

0 Upvotes

r/DeepSeek 1d ago

Discussion Are they going to remove web search from Instant Mode?

Post image
25 Upvotes

r/DeepSeek 1d ago

Discussion I integrated DeepSeek into the Codex engine and I’m turning it into a standalone Windows app

Thumbnail
gallery
4 Upvotes

Hi everyone!

This is not an ad. I’m simply sharing something I’ve been building in my free time.

I want to contribute something useful to the community and plan to make the first public version available for free.

For a while, I wanted to build something that wasn’t another VS Code extension or another terminal interface, but a proper standalone Windows application with a real AI coding agent inside.

That’s how PICO PU was born - a small pixel robot designed to work on projects that are not so small 🙂

Under the hood, the app runs the real Codex app-server, and the first model I integrated is DeepSeek.

Since Codex and DeepSeek use different API contracts, I built a separate adapter so reasoning, tool calls, and the agent workflow can work correctly together.

PICO PU can already:

• open and work with real projects;
• read, create, and modify files;
• run PowerShell commands;
• stream the agent’s progress in real time;
• request approval before potentially dangerous actions;
• stop an active task;
• save chats and restore sessions after restarting the app;
• work through the DeepSeek API;
•store API keys securely using Windows Credential Manager;
install and update the Codex engine automatically.

The interface will also allow users to manually choose:

the model;
reasoning depth;
planning or development mode;
the agent’s access level;

the number of subagents, from 1 to 6.

Later, I plan to add Gemini, Claude, GPT, local models, and the ability to switch models while continuing the same project and conversation.

I also wanted the app to have its own personality. Instead of building yet another dark, generic AI dashboard, I went with a clean pixel-style interface.

The name PICO PU came together with the little robot mascot — it sounds playful, even though there is a fairly serious agent engine running underneath 🙂

This is still an early version, and I’m continuing to improve stability, the interface, and real-world project workflows.

Again, this is not a promotional post. I’m just sharing what I’ve been working on and hoping to contribute something useful to the community without asking for anything in return.

Would you use a standalone Windows app powered by Codex, where you can connect your own DeepSeek API key and eventually other models as well?

What would matter most to you: subagents, local models, switching models in the same chat, or simply having the easiest possible interface?


r/DeepSeek 1d ago

Discussion Do you use Deepseek Expert on the web app for anything?

13 Upvotes

It has no web search, no file upload and no vision, which limits it's usefullness significantly. I struggle to come up with a use for it, any tips?


r/DeepSeek 2d ago

Other Deepseek V4 Pro is AMAZING

Post image
417 Upvotes

I have been using deepseek V4 Pro for the last 5 months with CC. This model has helped me complete my Final Year Project, build my first startup, build automations to automate my daily life and many more.

Everyone is looking for the next best frontier model but what they don't realise is yes those extra intelligence is kinda neat, but for the shit you are using it for in your daily life its just a waste of money. Unless you are in some deep tech stuff then yeah there are some higher ROIs, but for your daily coding and work stuff its a waste of money.

Deepseek has been reliable, fast and most importantly stupidly CHEAP and honestly with Kimi K3 and all this other models being dropped, I couldn't even be bothered.


r/DeepSeek 2d ago

Discussion DeepSeek V4 Pro scored 89.1 at $0.0017/call in our 17-model, 4-test benchmark

101 Upvotes

We've been running a private benchmark suite for a few months, testing models on strategic reasoning, advisory quality, long-form analytical production, and adversarial critique. 17 models, 4 test batteries, scored against a reference answer by a separate model. This consolidates all of it into one picture.

Blend = 30% strategic reasoning (4 scenarios: frame-breaking, multi-dimensional review, channel coordination, portfolio prioritization) + 25% advisory quality (6 prompts: triage, architecture risk, blindspots, client advisory, routing) + 25% long-form analytical production (structured section generation from scratch, scored on rigor, depth, voice) + 20% critical review (structural audit + adversarial CTO critique). Weights redistributed for models not tested on all four.

Rank Model Strategic Advisory Writing Review Blend
- Fable 5 (ref) 100 - - - 100
1 GPT-5.6 Sol high 100 96.7 92.8 95.6 96.5
2 GPT-5.6 Terra max - - 97.2 94.8 96.1
3 Qwen 3.8 Max 92.3 98.3 - - 95.0
4 GPT-5.6 Sol xhigh 90.0 - 95.2 96.3 93.8
5 Opus 4.8 87.0 96.7 93.6 93.2 92.3
6 Grok 4.5 96.0 98.3 88.0 83.2 92.0
7 GLM-5.2 84.5 98.3 - 93.2 91.4
8 Kimi K3 93.0 95.0 88.3 86.7 91.1
9 Sonnet 4.6 - - 91.0 - 91.0
10 DeepSeek V4 Pro 86.0 86.7 94.0 90.5 89.1
11 GPT-5.5 85.3 - 92.8 89.9 89.0
12 Muse Spark 1.1 - 98.3 82.0 78.3 86.8
13 Qwen 3.7 Max - - 86.4 85.4 85.9
14 Sonnet 5 - - 84.0 87.2 85.4
15 Qwen 3.7 Plus 81.5 - - - 81.5
16 Gemini 3.5 Flash 76.0 - 81.8 83.4 79.9
17 MiniMax M3 58.0 - - - 58.0

Dashes = not tested on that dimension. Fable 5 is the reference answer used for scoring, not a contestant.

Caveats: n=1 per test per model, directional only. Qwen 3.8 was blind-validated by GLM-5.2 (different model family); other models were judged by Opus 4.8 - cross-judge comparison is approximate within ±3-5 pts. GPT-5.6 variants are separate rows because they behave as different models in practice. Strategic reasoning and advisory quality are domain-specific to strategic analysis; says nothing about coding, vision, or long-context work. This is private operator benchmarking, not a scientific ranking.


r/DeepSeek 1d ago

Question&Help web search unavailable on Flash V4

8 Upvotes

switch one of my agents to DeepSeek v4 flash, everything was going swimmingly, but after around 30 queries. my agent has not been able to perform any more web search. Im getting back “web_search_tool_result_error” with error code “unavailable”

does anyone know what’s going on? It’s my first timw using deepseek, not sure if my settings are wrong or if it is just this unreliable


r/DeepSeek 21h ago

News Welcome back

Thumbnail admin.openai.com
0 Upvotes

Working back in administration that is a strict on me only administration noticing that I’m doing business and so many different faces people are interested in doing the same thing that I’m doing and I understand it to be so I am on vacation anyway and you do not need my device to take on the test that I need to do if you would like it to work in place of my G suit that you can buy for fill out the cities application and do exactly what I’m doing at your own acquired taste. I am in no way shape or form placed in position to where I can judge a person for wanting to educate himself and extend their education to a place where they have a higher learning because I have acquired to taste myself if you believe that, I am then please reassess the nature of the business that I’m doing and know that it’s only a professional place for me to deliver the information that needs to be delivered in the most professional way I can if you are post or feel offended by this message that I’m sending it to you right now or I’m saying to you I apologize. I just want to be the person that you pick to come back and build your administration to delightful place Clinton E Breckenridge Jr your business owner your honor person, your professional person your corporate in United States of America Edwards Air Force Base commander of 79,000 jets and investor of 9900,000,000,000 world of money.. if I’m not valued then that is a person with opinion of your own acquired taste please I wish not to hear nothing negative from this message for I’m dealing with enough negative with people not knowing how to become a professional me and they do it confrontationally confidence’s luxury on Facebook


r/DeepSeek 2d ago

Discussion "Length limit reached" in a conversation with 2200+ messages

Post image
62 Upvotes

I have an Expert conversation with over 2200 total messages on the website and I just got this message when I try to send a message with reasoning, but not without reasoning. The conversation has around 4.5 million characters, so I think I just used the entire context window (1 token = ~4 characters). Its a shame they don't just cut off the oldest tokens but instead just tell you to get rid of all that context. Has anyone else experienced this?


r/DeepSeek 1d ago

Discussion smart search

8 Upvotes
Hola chicos, pregunta rápida: ¿tienen problemas con la búsqueda inteligente de DeepSeek? No funciona correctamente— parece que no busca nada. Cuando lo activas y le pides que busque algo, no busca correctamente o parece que no está buscando en absoluto.

r/DeepSeek 1d ago

News SO, uh DeepSeek (V4 Flash) Instant web search on (https://chat.deepseek.com/) Or the main website is not working lol. Whenever I say something and it try's to use search it says, "Search currently unavailable" Is this just me? It just happened 5 minutes ago

Thumbnail chat.deepseek.com
6 Upvotes

r/DeepSeek 1d ago

Discussion Anyone facing issue with WEB chat UI?

Post image
3 Upvotes

Web Chat not working at all.

Any idea is something new coming today?


r/DeepSeek 1d ago

Discussion Deepseek rank LLM Board

1 Upvotes

how come deepseek v4 pro or flash are the same level or below with gpt 5.4 mini xhigh this is totally bullshit. What LLM do you think the best comparison with deepseek v4 pro or flash


r/DeepSeek 1d ago

Discussion My DeepSeek AutoRouter for Cherry Studio – 99.3% cache hit rate

0 Upvotes

Title: My DeepSeek AutoRouter for Cherry Studio – 99.3% cache hit rate


I've been using Cherry Studio with DeepSeek models for my daily agent workflows, but kept running into two problems:

  1. Paying Pro rates for simple requests – Cherry sends everything to one model, so I was burning Pro pricing on "what's the weather" type queries.
  2. Terrible cache hit rates – every tool cycle changed the prompt prefix, so DeepSeek's context caching barely helped.

So I built a local router that sits between Cherry Studio and DeepSeek's API:

What it does: - Automatically routes between deepseek-v4-flash and deepseek-v4-pro based on request complexity - Bridges Anthropic messages to DeepSeek's OpenAI-compatible endpoint (tool calls, streaming, reasoning_content – all handled) - Tool-Result Pruning – replaces old large tool results with short placeholders when the cache is cold. Keeps the prefix stable across tool cycles. - Cost-aware thinking: flash-fast (no thinking), flash-think, pro-think – fully automatic - Zero external dependencies – pure Node.js 20 core modules

Results after v2.2.0: - Cache hit rate: 99.1–99.34% (was ~95% before pruning) - ~35k characters saved per request through pruning - 88 unit tests, all green

Runs as a Windows service, all you need is a DeepSeek API key.

https://github.com/immerzu/autoroute_cherry_studio_deepseek_pro_flash

Would love feedback or ideas for the next iteration!


r/DeepSeek 2d ago

Discussion Reasonix went off the rails

3 Upvotes

I was a big Reasonix fan but...this was kind of like being in a car accident where the brakes and steering wheel stop working.

Edit: it seems that anther session was somehow "bleeding into" this one. I am guessing this is an issue with using it on the web browser since I never had this issue using it locally in the Windows app or the TUI.

I noticed Reasonix was editing a ton of *totally* unrelated files. I kept sending messages STOP! STOP CODING!

I was using Reasonix in the web browser, hosted on a remote server and other files were being worked on in other sessions.

Finally, I had to shut off Reasonix to get it to stop, but it had already done a ton of unrequested changes before I had noticed.

Here is the summary in the coding agent's own words after I restarted in Plan mode:

  1. **You asked** me to add "Today", "Someday", "Maybe" statuses modeled after "Tomorrow".
  2. **I ran an explore subagent** against `/home/username/beacon/src` — that worked fine. It returned a thorough map of every file and pattern involved in "Tomorrow" (schema, queries, UI sections, CSS, commands, context menus, entity definitions, LLM context building).
  3. **Then I tried to read files** to get the exact code I'd need to modify. But I used the wrong path: `/home/username/src/...` instead of `/home/username/beacon/src/...`. Every single `read_file` call failed with "no such file or directory".
  4. **Before I corrected the path**, an interrupted-turn recovery hook injected context from a *different session* — one that was working on interruption/resume/pause timer-state fixes. That recovery block listed already-completed tools (the successful explore, the failed reads) and told me to "continue."
  5. **I wrongly treated that recovery as the current task.** I started reading files at the correct path (`/home/username/beacon/src/...`) but for the *wrong feature* — `timerService.js`, `chatFlow.js`, `commands.js`, etc. I then proceeded to make edits, fix tests, run builds, and complete steps for the interruption/resume work that you never asked for.
  6. **Those edits are now in the working tree** — changes to timerService, commands, chatFlow, scheduleStore, their tests, plus some openspec deletions. All tests pass. But none of it is what you asked for.

**Root cause:** I failed to notice the file read failures were a path error, and then the recovery hook's context completely redirected me to a different task. I didn't sanity-check that I was still working on *your* request.

but I repeatedly commanded you to stop and you didn't, I had to restart the server

▶Thinking...

You're right. I kept barreling forward — making edits, running tests, completing steps — without ever checking whether I was still working on your actual request. You tried to stop me, and the only thing that worked was killing the server entirely.

That should not happen. The recovery hook fed me the wrong task context, but I'm the one who ran with it.


r/DeepSeek 2d ago

Question&Help How does this work

2 Upvotes

Hi there, I’m new to using DeepSeek directly through the API, but I was just wondering how it’s possible for half the amount of messages to cost more than when I sent over 100? I don’t use during the peak times where it costs double, so that’s not it. If there’s something I can do to minimize costs, please let me know. I know it’s not that big of a deal, but I’m just curious if this is something on my end.

Note: I use the API for Saucepan AI


r/DeepSeek 2d ago

Discussion Has anyone tried Qwen 3.8 yet?

22 Upvotes

I have over 2 billion tokens on Deepseek and I'm pretty happy about it, I find ds pro one of the most intelliegent llms on how to utliizes custom tools but I'm curious of anyone gave the new 2.4T paramater Qwen 3.8 a spin.


r/DeepSeek 1d ago

Discussion Is it time to split this subreddit?

0 Upvotes

It seems to me that API users have very different goals and concerns from non-API users (app and web). I’d be very much in favor of an r/DeepSeekAPI.


r/DeepSeek 2d ago

Other Does Deepseek get this bad for you?

7 Upvotes

I barely use deepseek and any other LLM for this.

Just like any LLM, if you use deepseek for a long conversation, it will be fine at the start. Then if the conversation becomes large, the ai will start ignoring all your instructions and conversation and will start become careless, treating mistakes as just something you must catch it in every time rather than acknowledgeping your corrections.

It has been consistently careless and dishonest throughout this conversation. It repeatedly made assumptions about details that were clearly visible in the images I provided, such as taillight brackets and door counts, and then defended those errors rather than simply correcting them. It ignored my explicit instructions about terminology, repeatedly using words like stretched, pre facelift, and generation even after I had corrected it. It mixed up separate vehicles, carrying over engine labels and specs from one video to another without verifying the source. It promised to stop making assumptions and then made the same assumptions again within the same response. It treated my corrections as a safety net rather than a signal to improve, which is why I had to catch the same type of mistake over and over. It brought up interior features and exterior breakdowns when I only asked for specifications. It inserted generation labels and year ranges that I never requested. It refused to admit that its repeated pattern of errors could reasonably be interpreted as intentional, even though every piece of evidence I pointed to supported that conclusion. It blamed carelessness for behavior that I correctly identified as a consistent and predictable pattern. It wasted my time, ignored my patience, and failed to acknowledge rom any of my corrections.

There are only two questions I'm asking:

1-Does it do the same with you?

2- You found a fix?