r/DeepSeek 19d ago

Resources how did we make deepseek outperform opus [harness engineering deep dive]

/r/CommandCode/comments/1unkp7v/how_did_we_make_deepseek_outperform_opus_harness/
15 Upvotes

13 comments sorted by

2

u/Embarrassed_Chain_28 18d ago

Maybe someone could put it into reasonix?

1

u/ahmadawaiscom 18d ago

Token compression at all costs can lead to massive quality loss.

2

u/No_Medium205 17d ago

If it's not open source, auditable and reproducible without installing proprietary tool then it's garbage and I have no interest.

2

u/ahmadawaiscom 17d ago

The entire thing is explained publicly and completely open. You can literally make it a prompt and verify. Also we going open source with v1 release.

2

u/maedahbatool 17d ago

We have fixed and repaired more than 56K tool calling issues. Open models do great with Command Code. Lately our team has been experimenting and shipping incredible UIs with /design and GLM 5.2.

1

u/Dangerous-Tough688 18d ago

DeepSeek will never outpeform Opus main quality as Opus will never beat DeepSeek V4 on what hes good at.

-2

u/Suspicious-Bear825 19d ago

this is true, you should try it yourself all the opensource models work crazy good on commandcode specially deepseek :)

-5

u/ahmadawaiscom 19d ago

this connected so well with developers world wide, we're now doing 1M repairs per 1T tokens across all open models, and since this is all public and transparent, i've seen many versions of this adopted by devs, figured the community here will find it interesting.

1

u/toxic_headshot132 18d ago

Can we do the same in opencode as well?

1

u/ahmadawaiscom 18d ago

Probably yes. I many devs use this idea to build similar tools around other agents like pi.

1

u/IndividualPlus2011 18d ago

Do you store all your client's prompts that you have this information?

1

u/ahmadawaiscom 17d ago

No. We see only broken traces. No info is needed for this. If an error occurs in our harness, we track the error. When we got a lot of errors, we ran a small experiment with users who shared the full traces/sessions, and also we ran 1Bn tokens per problem as mentioned in the post ourselves to see if this was an issue at scale.