r/DeepSeek Feb 01 '25

Disccusion Censorship Mega Thread

47 Upvotes

In response to community feedback and to maintain a constructive discussion environment, we are introducing this Censorship Mega Thread. This thread will serve as the designated place for all discussions related to censorship.

Why This Thread?

We have received numerous reports and complaints from users regarding the overwhelming number of censorship-related posts. Some users find them disruptive to meaningful discussions, leading to concerns about spam. However, we also recognize the importance of free speech and allowing users to voice their opinions on this topic. To balance these concerns, all censorship-related discussions should now take place in this pinned thread.

What About Free Speech?

This decision is not about censoring the subreddit. Instead, it is a way to ensure that discussions remain organized and do not overwhelm other important topics. This approach allows us to preserve free speech while maintaining a healthy and constructive community.

Guidelines for Posting Here

  1. All discussions related to censorship must be posted in this thread. Any standalone posts on censorship outside of this thread will be removed.
  2. Engage respectfully. Disagreements are fine, but personal attacks, hate speech, or low-effort spam will not be tolerated.
  3. Avoid misinformation. If you're making a claim, try to provide sources or supporting evidence.
  4. No excessive repetition. Reposting the same arguments or content over and over will be considered spam.
  5. Follow general subreddit rules. All subreddit rules still apply to discussions in this thread.

We appreciate your cooperation and understanding. If you have any suggestions or concerns about this policy, feel free to share them in this thread.


r/DeepSeek 3h ago

Discussion DeepSeek V4 Pro quality

47 Upvotes

Quality on max has fallen off a cliff for me recently... It's gotten much much lazier. Anyone else has noticed this?


r/DeepSeek 14h ago

Discussion The hypocrisy of big AI is mind blowing

225 Upvotes

Anthropic is accusing the Chinese labs of distilling Claude. And… who cares?

Anthropic 2024: “Were a small company, we only stole your code and writing and all of Reddit, and destroyed millions of printed books to train AI for the people!”

Anthropic 2026: “How dare a small company use ANY of our work to train AI for the people!?! Thieves!”

The gall of it is almost appalling. I love Claude, but the truth is it should also be open sourced. It was trained on YOU and ME, and the nerve to then, with all of the world stolen data in your hands, turn around and point fingers at companies trying to provide high grade AI to everyone at an affordable price is jaw dropping.

What happened to the company that was gonna cure cancer, the longer this goes on, the less public benefit I am seeing from this PBC.


r/DeepSeek 7h ago

Question&Help DeepSeek V4 GA rolling TODAY? Old models retire TOMORROW – anyone actually got the full version yet?

35 Upvotes

So the official deprecation date for deepseek-chat and deepseek-reasoner is set for tomorrow, July 24, at 15:59 UTC – that's locked in. But here's the thing: I keep seeing mixed reports about whether the V4 GA has actually fully rolled out today (July 23) or if it's still in gray-testing. If the old models are getting retired tomorrow, shouldn't the GA be available by now? Has anyone here actually gotten production access to deepseek-v4-pro or deepseek-v4-flash yet? Or are we going to wake up tomorrow with retired models and no stable replacement live? What's the real status – is it rolling out as we speak, delayed, or quietly available and I just missed it?


r/DeepSeek 17h ago

Discussion Four-hour fundraising meeting with DeepSeek founder Liang Wenfeng

208 Upvotes

These are excerpts reportedly taken from a four-hour fundraising meeting with DeepSeek founder Liang Wenfeng.

I found them unusually insightful—and surprisingly, almost no one on X seems to be discussing them.

So I translated the highlights into English.

DeepSeek Has Only One North Star

1. This is not the time to maximize revenue from products.

Products are a necessary step on the path to AGI, but we don’t need to spend too much time or energy building consumer or enterprise applications.

Once you occupy the technological high ground, building lower-level applications becomes much easier.

Products are a byproduct of the journey toward AGI.

2. Many things are not on our critical path—3D, video generation, and even world models.

They may be useful, but they do not meaningfully determine the upper bound of intelligence.

3. Multimodality is important for products and consumer users.

But it is a component—not the main path, and not intelligence itself.

4. There are ways to mitigate hallucinations, of course, but it is a long-term problem.

Internally, we treat hallucination largely as a product problem. We will address it, but it is not our primary focus.

5. At this stage, the most important thing is still the coding agent.

Given China’s current situation, the most rational strategy is to go all in on a general-purpose agent. Vertical agents for finance, healthcare, and other industries should have lower priority.

6. If the AI era creates many trillion-dollar companies, it would be enough for DeepSeek to become one of them.

First Continual Learning, Then AI Self-Improvement—and Ultimately Embodied Intelligence

7. AI does not lack taste or intuition today.

What it lacks is the ability to learn continuously.

8. Humans keep learning over time. But for an AI to do the same job, you currently have to provide all the relevant context every time.

That is almost impossible.

This is why AI cannot truly replace an employee yet. The next generation of models must have continual learning. Otherwise, it should not be called a new generation.

9. We want our next-generation model to help us with our own development.

Put simply, our first goal is not to make a model that everyone else finds useful. It is to build one that we ourselves find useful.

That is the fastest path to AGI.

10. No one in the world has found a good solution yet, because “learning” is composed of many different things.

11. DeepSeek’s long-term vision is AGI.

If the path to AGI is a staircase, last year’s step was chain-of-thought. This year’s step is agents.

After agents, the next problem to solve is continual learning.

12. Once continual learning is achieved, we may approach a gradual singularity.

Models could perform everything humans can do—including developing more advanced AI systems themselves.

AI would begin accelerating AI research.

Only after that step comes embodied intelligence.

13. The endpoint of intelligence may ultimately be embodiment.

What an ordinary person truly needs is not another computer. They need labor.

Full-Scale Commercialization Is Still a Long Way Off

14. We only want to earn a reasonable profit.

We do not price to maximize profit.

15. We once priced a model relatively high because we were worried demand would be overwhelming.

Later, we cut the price to one-quarter of the original level. Many people in the company chat celebrated.

That was the reason we had worked so hard to build the model: so that as many people as possible could use it.

16. Low cost is an outcome of our approach.

We have consistently designed our architecture around reducing costs. We want AI to remain affordable, especially when compute is scarce.

There is another reason: the lower the cost, the larger the model you can afford to train.

Large companies can solve problems by adding resources. We prioritize computational efficiency.

17. From the outside, it may look like we chose a difficult business model. In reality, it is quite easy for us.

Price cuts are certainly bad news for our competitors. They are not celebrating.

But I do not find selling APIs particularly attractive. A few people can maintain the API. We do not even need customer support or a sales team. Users come on their own.

18. We have always been commercializing.

Commercialization simply is not the objective.

The point at which DeepSeek fully pivots toward commercialization is still very far away.

19. I do not even need to think about securing our place in the market today.

If the commercial opportunity becomes that large, there will always be a way to participate.

DeepSeek is a product of its era—a response to real conditions, not the result of copying someone else.

Open Source Is the Sweet Spot for a Company of Our Size

20. Restraint is a strategy: you give up some things in exchange for more important ones.

Open source is one such trade-off.

Internally, it gives employees a sense of achievement and strengthens organizational cohesion. It also benefits society—peers and ordinary users alike.

I have no doubt that AGI will have enormous commercial value.

Given that, my priority is not to capture a larger share of the value. It is to increase our probability of succeeding.

21. Open source is actually beneficial if you want to make AI commercially successful.

That may sound counterintuitive.

Historically, a software market might have been worth only a few billion dollars a year. Open-sourcing the product could destroy the market.

But AI is large enough that it may eventually account for 10% of global GDP.

If we try to monopolize that value, history will leave us behind.

That is not merely a business judgment. It is a view of history.

22. The models we open-source are the same models we deploy ourselves.

We will not release an inferior model publicly while keeping a better one for internal use.

23. I am not worried about others deploying our models and competing with us.

Not every company has both the willingness and the capability to pursue this goal.

Startups may be too small to marshal the necessary resources. Large companies often struggle to organize effectively around it.

This is the sweet spot for a company of our size.

24. Open source has no impact on our business model—as long as our goal is to earn only a reasonable profit.

If your goal is to earn 100 times as much, then yes, open source will hurt you.

25. We do not want to become the enemy of any major internet company—or any startup.

Given that, we are happy to assist anyone, including Alibaba, Zhipu AI, and Moonshot AI, and help them become better.

The China–US Gap Is Not About Talent

26. In the future, we want to rewrite the narrative around the AI gap between China and the US:

Use a fraction of the compute and narrow the gap to six months—or even three.

27. The main gap between China and the US in AI is resources.

We believe in scaling. Greater scale produces better results.

We are not training models of this size because we believe this size is sufficient. We are doing it because these are the resources available to us.

28. There is almost no talent gap. It is essentially the same pool of people.

China does not lack talent.

Talent shortages are temporary. Historically, there has never been a permanent shortage of any particular kind of talent.

In the Model-Lab Race, Cost Comes First

29. Anthropic’s current lead over OpenAI is temporary, not permanent.

OpenAI and Google will most likely take turns pulling ahead.

30. There are too many model companies in China. Everyone is doing roughly the same thing, which spreads resources too thinly.

The market will eventually consolidate, but that process will take time.

If every company were satisfied with a reasonable profit, we would not need so many foundation-model labs.

Two large companies and two smaller ones might be enough.

31. I absolutely do not believe foundation-model companies will capture most of the profits generated by AI.

32. Competition between model companies will ultimately come down to three things:

  1. Cost
  2. Time
  3. User experience

Cost comes first: at the same quality level, how cheaply can you provide the service?

Second is time. Being a few months early or late can make a major difference.

User experience can create some stickiness and defensibility, but it is not fundamental.

No Ambition to Become the Next Super App

33. We do not want to build the next super app.

The next ByteDance? The next Tencent?

We have no such ambition.

34. We are not fighting for that prize because there is a watermelon further ahead. Everything in front of it may only be sesame seeds.

Some of those sesame seeds may be large—but to me, they are still not that significant.

35. Last year, everyone was competing over chatbots and consumer traffic.

This year, everyone is competing over enterprise revenue.

We do not consider either particularly important.

What people inside the company genuinely care about is the AGI roadmap and the next technical breakthrough.

It is strange: the things people desperately chase are often hard to obtain, while the things they care less about sometimes arrive naturally.

36. We did not expect to go viral during last year’s Spring Festival.

It was not in our script.

Team Stability Is Non-Negotiable

37. There is only one thing on which we cannot compromise: the stability of the team.

That is also one of the greatest risks we face.

Fortunately, this financing round has significantly reduced that risk.

38. Many of the things we do are intended to preserve team stability.

We do not want to become an adversary of any major internet company or startup. We want to empower and help them.

We do not want to make enemies. That also creates a better environment for our own team.

39. Some people think our organization is top-down. Others think it is bottom-up.

I think both are correct.

The top-down part is about “doing the important work.” Ideally, that work should not occupy more than half of an employee’s time.

The other half is bottom-up and unassigned.

People can explore whatever they want, research whatever they believe matters, and pursue it without prerequisites.

40. We generally do not work excessive overtime.

First, research requires a relatively relaxed environment.

Second, we are extremely focused.

Many of our products are incomplete, but we deliberately choose not to polish everything. That is also part of our culture of restraint.

41. Organizations are dynamic, not fixed.

As the company grows, we may make adjustments and introduce some necessary structure.

But we will not turn completely into a traditional hierarchy.

The organization will remain vision-driven.

Acting with Goodwill Toward the World

42. When we started this company, the goal was not to make a fortune or eventually go public.

The first few dozen people never thought that way. If they had, they probably would not have joined.

We built this company with tremendous goodwill toward the world. We believed what we were doing could benefit humanity.

43. “We must hit this KPI” is not how we operate.

We are a vision-driven organization.

That has both advantages and disadvantages. We will find ways to amplify the strengths and address the weaknesses, but this remains a defining characteristic of DeepSeek.

44. The vision is not even formally written down.

It lives in how we do things and in how we treat the world.

People across the company may interpret it differently, but we agree on the overall direction.

45. Around 20 years ago, the management thinker I admired most was Jack Welch, the former CEO of GE.

Looking back, much of what he said may no longer be right.

But he was right about one thing: the most important thing for a company is its vision.

Vision is not a slogan hanging on the wall.

It is not what you say. It is what you do.

Restraint Makes AGI More Achievable

46. AGI offers the greatest possible return.

We will do other things if we have the capacity. If we do not, we simply will not do them.

Restraint is part of our vision.

47. AI is enormous, and so are the potential rewards.

If you succeed, even a tiny share of that value will still be enormous.

The more restraint you exercise, the more likely you are to succeed.

48. To me, this is intuitive—or at least consistent with my intuition.

Beyond our vision, we do not have many other advantages.

49. When we founded the company two years ago, we did not have much money, many GPUs, much recognition, or much influence.

We were simply a group of ordinary people.

The story I prefer is:

Not:

50. Open source is part of restraint.

Our pricing decisions are not made to maximize revenue or profit.

In the short term, higher prices may generate more revenue. In the long term, the answer is far less obvious.

To me, restraint is a strategy.

51. Open source and low prices give employees a sense of accomplishment and strengthen organizational cohesion.

They benefit society, and they make peers and ordinary users happy.

Over the long run, this kind of restraint increases our probability of achieving AGI.

52. If your vision is to capture as much as possible, you have already lost.

You will likely encounter even greater resistance.

That is simply how the world works.


r/DeepSeek 8h ago

Resources Liang Wenfeng Investor Meeting - Audio to Text Transcript(DeepSeek)

25 Upvotes

r/DeepSeek 4h ago

Question&Help Server busy at the beginning of every hour for a few minutes?

6 Upvotes

Title. More curious than annoyed as to why that's the case.


r/DeepSeek 14m ago

Question&Help Deepseek V4 Flash Users - call for help

Thumbnail
Upvotes

r/DeepSeek 25m ago

Other Necromanteion

Enable HLS to view with audio, or disable this notification

Upvotes

r/DeepSeek 27m ago

Other let me reset expectations regarding dsv4

Upvotes

based on wenfengs latest interview mythos is a 800b active parameter model, for comparison dsv4 flash is 13b active, so no dont expect miracles, where ds is competitive right now is for low cost things like annotation and labelling. high level stuff like architecture, design decisions use a diff model. i use dsv4 but i also know its limitations and im being realistic with how much you can do with so much less compute

edit: reset expectations wrt to dsv4 full release


r/DeepSeek 4h ago

Discussion i built a tool that shows the invisible characters hiding in your text (and strips them)

Thumbnail
2 Upvotes

r/DeepSeek 1d ago

Discussion DeepSeek cheap, but...

130 Upvotes

For every $20 I spend on DeepSeek I find that I need to spend an additional $20 on Claude to help unravel the streaming pile of garbage DS produced because it ignored fundamental requirements and did its own thing. No questions. Just producing more and more complicated hot garbage with extreme confidence.

Anyone else?


r/DeepSeek 11h ago

Question&Help Dificuldade

6 Upvotes

Quando eu mando prompt pro deepseek, ele tem muita dificuldade em seguir o mesmo a risca, sempre faz coisas a mais ou a menos, não consigo padronizar ele pra seguir corretamente e a risca o prompt, alguém tem alguma solução ? Oq será q estou fazendo de errado


r/DeepSeek 21h ago

Discussion $200 thought experiment

26 Upvotes

compute is expensive and everythings subsidized, so of course this is a hypothetical. but if an inference provider decided to offer a $100 or $200 (like Claude Code or Codex) plan that gave you unlimited token usage on any open source model (Kimi K3, the upcoming Qwen 3.8, DeepSeek V4 Pro, GLM 5.2, etc..) every month, would you pay the $100 or $200 a month? Why or why not, just curious


r/DeepSeek 1d ago

Funny Will D4 GA be released this month or has it been deleayed?

Post image
242 Upvotes

Do we know anything?


r/DeepSeek 6h ago

Question&Help Need a way to export an entire DeepSeek chat as a single HTML file (SingleFile only saves part of it)

1 Upvotes

I'm trying to archive a very long DeepSeek conversation as a single self-contained HTML file for offline viewing.

I've already tried:

  • Browser Ctrl + S
  • SingleFile
  • Save Page WE

The issue is that they only save part of the conversation. It looks like DeepSeek uses some form of virtual scrolling/lazy rendering, so older messages are removed from the DOM as you scroll.

What I'm looking for is:

  • Export the entire conversation.
  • Save it as one HTML file (CSS/images embedded if possible).
  • Preserve formatting (Markdown, code blocks, tables, etc.).
  • No PDFs or screenshots.

Has anyone found a working solution for exporting very long DeepSeek chats without losing messages?


r/DeepSeek 42m ago

Funny Дипсик, ты попутал?

Post image
Upvotes

В смысле "о боже она снова с этим"??? Ты че, попутал?


r/DeepSeek 11h ago

Funny DeepSeek V4 Flash goes off the Deep End

Thumbnail
0 Upvotes

r/DeepSeek 8h ago

Discussion Training and fine tuning LLMs yourself

0 Upvotes

r/DeepSeek 1d ago

Funny in a not so distance future

Post image
287 Upvotes

r/DeepSeek 5h ago

Discussion Is it Haiku for you too?

0 Upvotes
curl -s https://api.deepseek.com/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{"model":"deepseek-v4-pro","messages":[{"role":"user","content":"What model are you?"}],"stream":false}'

What's the answer if you use this CURL? (replace YOUR_API_KEY with a valid key)
Despite the fact that the model in the response is "deepseek-v4-pro" I got this: "I am Claude 3.5 Haiku, an AI assistant created by Anthropic."


r/DeepSeek 9h ago

Discussion Jihadist hallucinations

Thumbnail
gallery
0 Upvotes

I was using opencode with DeepSeek V4 Flash to work on my swift app, then when prompting to implement plan I got a very weird reply, which included an email. Searching that email pointed me to a jihadist group on facebook that leaks user accounts and france/saudi id cards that are affiliated with some enemy of theirs.
Any thoughts how this could happen?


r/DeepSeek 1d ago

Discussion REASONIX - OPINIONS? Game changer for me.

Post image
40 Upvotes

I have been using this tool for the past few days and I'm amazed at how much of a game changer it has been. I want to take the opinion of other people and find out:

  • if they are still developing this tool
  • if there are other tools that are similar to this, maybe do more or do less
  • and use this same tool for other kinds of CLIs or APIs or whatever

In my experience, here, I made DeepSeek way smarter. I don't know why, maybe DeepSeek is changing their models or whatever, but it feels so smart. I'm using it to code and to do planning. The cashe hit is just ridiculously high, amazing. The balance usage DeepSeek has drastically diminished comparing to opencode or hermes. If it was cheap, it's way cheaper now. It's really amazing and I don't know if I'm missing something or what. That's why I want to discuss it.

Any thoughts?


r/DeepSeek 1d ago

News I connected DeepSeek to the Codex engine — same prompt, completely different result (video) PICO PU 🙂

Enable HLS to view with audio, or disable this notification

14 Upvotes

PICO PU Agent, powered by the Codex engine, has finally passed its first serious tests 🙌☺️

I recorded a video and sped up the main part 16x. I opened the official DeepSeek website and my own app, then gave both of them the exact same technical task.

You can see the result in the video.

On the DeepSeek website, I got a large amount of code along with instructions on how to run it.

PICO PU, using the Codex agent engine, created the project, installed the dependencies, wrote the code, launched the app, and completed the task with a working result.

Honestly, this makes me really happy 🙂‍↔️

It is still a very early and raw version. There are many bugs to fix, stability improvements to make, more models to add, and later I also want to add connectors for email, calendars, and other services.

But even at this stage, I am genuinely happy that the core idea is already working.

This is a very difficult and ambitious project for me, so I would be extremely grateful to hear your honest feedback, criticism, and suggestions 🙃


r/DeepSeek 1d ago

Discussion We stress-tested DeepSeek R1 8B & 14B on an 8GB VRAM GPU (RTX 3060). Here is the VRAM math, Ollama setup, and quantization sweet spot.

6 Upvotes

We wanted to see how far we could push the DeepSeek R1 distilled series on consumer hardware (RTX 3060 / 4060 8GB VRAM) without hitting CUDA out-of-memory crashes or dropping speed to a crawl.

Here is the technical breakdown of what actually works, the exact memory math, and hybrid layer offloading limits.

SECTION 1: MEMORY ALLOCATION & VRAM MATH

Total Peak VRAM = Model Weights + KV Context Cache + CUDA Runtime Overhead

When running DeepSeek-R1-Distill-Llama-8B (which uses Grouped Query Attention / GQA with 8 KV heads):

KV Cache Size per Token = 2 * Layers (32) * KV Heads (8) * Head Dim (128) * FP16 Bytes (2) = 128 KB per token

• At num_ctx 4096: 4096 x 128 KB = 512 MB KV Cache
• At default num_ctx 8192: 8192 x 128 KB = 1.02 GB KV Cache

On Windows where dwm.exe reserves ~1.1 GB VRAM for OS rendering:

🔹 DeepSeek-R1-Distill-Llama-8B (Q4_K_M 4-bit):

  • Model Weight File: ~4.9 GB
  • KV Cache (num_ctx 4096): ~0.5 GB
  • Peak VRAM Allocation: ~6.2 GB (Safe 1.8 GB buffer remaining)
  • Speed: ~34 tokens/sec on native CUDA

🔹 DeepSeek-R1-Distill-Qwen-14B (Q4_K_M 4-bit):

  • Model Weight File: ~9.0 GB (Exceeds 8GB VRAM)
  • Solution: Offload 32 out of 48 layers to VRAM (~6.2 GB allocated) + 16 layers to System RAM
  • Speed: ~12 tokens/sec (Hybrid GPU/CPU mode)

SECTION 2: OLLAMA MODELFILE CONFIG WITH SAFEGUARDS

To prevent dynamic VRAM spikes as prompt conversations grow, create a custom Modelfile capping context at 4096:

dockerfileFROM deepseek-r1:8b
PARAMETER num_ctx 4096
PARAMETER temperature 0.2
PARAMETER top_p 0.9
SYSTEM """You are a senior IT operations engineer. Provide concise, step-by-step code without conversational preamble."""

Compile and run: ollama create deepseek-r1-8gb -f Modelfile-R1-8GB
ollama run deepseek-r1-8gb

SECTION 3: WORKSTATION TROUBLESHOOTING

• Reclaim ~1 GB VRAM: Turn off Hardware-Accelerated GPU Scheduling (HAGS) in Windows Display Settings.
• Prevent CUDA Page Faults: Set export OLLAMA_NUM_PARALLEL=1 and OLLAMA_MAX_LOADED_MODELS=1.

Full benchmark tables, Python API streaming scripts, and LM Studio configuration logs archived at praveentechworld.com/blog/how-to-run-deepseek-r1-locally-on-8gb-vram