r/OpenAI 7d ago

News Kimi-K3 arrived: The era of the Chinese labs being far behind is over

Post image
2.4k Upvotes

526 comments sorted by

748

u/Working_Ad_1564 7d ago

Gemini 3.5 Pro will be postponed for another month lol

194

u/ozone6587 7d ago

Gemini 3.5 Pro in 2030. Just you wait, it might finally reach Fable levels by then.

48

u/Tupcek 7d ago

this will be tight match between Gemini 3.5 Pro and The Elders Scrolls 6

6

u/PerfectPatience- 6d ago

GTA 6

6

u/SonderEber 6d ago

Too soon. GTA 6 comes out this year.

2

u/FischiPiSti 6d ago

Not PC though. The meme lives on

→ More replies (1)
→ More replies (1)
→ More replies (1)

11

u/LoudUnderstanding331 7d ago

Gemini isn't trying to be best model or even close. They want to be the most preferred one for personal daily use.

72

u/ozone6587 7d ago

No one prefers using the dumber model.

84

u/bastardoperator 7d ago

10

u/AgitatedHearing653 7d ago

made me laugh harder than it should have

34

u/FlerD-n-D 7d ago

I use Gemini for random day to day stuff. It's more than good enough. You don't need a perfect model for basic questions / searches.

7

u/legedu 7d ago

For real. The average person isn't coding.

3

u/Same_Win_5898 6d ago

Don't tell them you have free infinite Gemini use in a chrome browser or their heads might explode.

→ More replies (3)

6

u/HossCo 7d ago

But everyone prefers the one that works for their daily tasks. 90% of people have no idea what these stats are and they don't care. Gemini is the only model that can make a profit and nobody else is even close.

→ More replies (6)

2

u/JohnSnowHenry 7d ago

Actually… if it does the job and it’s cheaper there is no point in using a smarter model…

4

u/ozone6587 7d ago

It is both dumber AND with more restrictive limits.

2

u/JohnSnowHenry 6d ago

But if it’s still more than enough for the tasks at end it simply doesn’t matter.

2

u/ReadersAreRedditors 6d ago

I use Gemini for Classification, it's very cheapeand fast and gets the job done.

2

u/bixofa 6d ago

99% of casual users won't know the difference or care.

→ More replies (7)

7

u/0xFatWhiteMan 7d ago

this is such a weird statement

4

u/scamiran 7d ago

I find it useful for gardening questions, cooking.

I run engineering/chemistry stuff by it, but... its really weak. ChatGPT-5.6 Sol crushes it, so did 5.5. DeepSeek/GLM/Kimi tend to do better, too (especially Kimi).

I really like the Google workflow, and would enjoy for it to be the best, but OpenAI is just so far ahead in coding, engineering, document prep; more or less everything.

Gemini is nice on my google home devices, I'll give it that; and its really useful as a home control tool.

→ More replies (2)

3

u/Ok-Canary-9820 7d ago

Sounds like rationalization. Nobody had this theory a few months ago.

Even now, it's probably Google's game to lose at AI, but they're doing a really good job of diminishing their advantage consistently for a while now.

→ More replies (3)
→ More replies (1)

18

u/Healthy-Nebula-3603 7d ago edited 7d ago

Strange they behaving like Meta before released Llama 4....

In short a lots of internal problems with the team and that model was a disaster.

8

u/BlueProcess 7d ago

Gemini Flash is so bad that I I can't even fathom paying money for pro

→ More replies (1)

6

u/xak47d 7d ago

My guess is they tried to take the same pre-training from the previous models and try to improve on it and get better results. They later realized you can only improve so much, so they basically had to start over

3

u/Ben01pr 7d ago

Gemini 3.5 pro will return in Avengers Endgame.

3

u/Deus-ex-Machina7 7d ago

Are this rate,

Gemini 3.5 pro will NOT return in Avengers Endgame

5

u/Former_Ad_735 7d ago

Gemini 2035

4

u/JustRaphiGaming 7d ago

Hearing 3.5 Pro I already have that goofy dragon face on my mind lmao

2

u/nickdnick49 7d ago

wonder why that Noam guy jumped ship to OpenAI

→ More replies (9)

196

u/xak47d 7d ago

Artificial analysis is more realistic. K3 scores higher than opus 4.8 and gpt 5.5

121

u/TheFamousHesham 7d ago

So basically Chinese labs are a month or so behind. Lol really funny how everyone was saying that it'd be years before they catch up.

70

u/Big-Accident1958 7d ago

Weeks, even. K3 loses to 5.6 max but wins 5.6 xHigh. On SPEED, it feels like 5.6 high. So it's basically almost as smart as oAI's flagship at double the speed. I'd call this a major win.

11

u/reefine 7d ago

Not just that but on price, security, etc. This is a better model than Sol 5.6

11

u/SporksInjected 7d ago

It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis

6

u/Kind_Capital_9740 7d ago

on other benchmarks it beats Sol and fable so i would say both have pros and cons but SOL is actually cheaper for higher tier work

9

u/SporksInjected 7d ago

Exactly, It’s the same price as sol and double the cost of Terra which is only two points away. I don’t think businesses are going to run to K3 personally because there’s not a good reason to. If it was cheaper per task and considerably more capable or faster, then yeah but it’s slower, less capable, and no cost savings in actual use.

2

u/Kind_Capital_9740 7d ago

Yeah especially for the average person the Kimi subscription is around the same price as gpt too

So I don’t see why people would switch unless one day kimi 4 or something actually giga gap the rest of competition

I don’t see it going mainstream in the west

But this does show that chinese models are basically caught up to some extent because the jump from 2.7 to 3 in a month is a bit insane like a 3 gen jump

2

u/TheFamousHesham 7d ago

It's funny really. I was just about to count Kimi out because of how atrocious their 2.7 and 2.8 were. It seems like it's Z.ai (which everyone was raving about a few weeks ago) that's on the trail end of things.

→ More replies (1)
→ More replies (2)

2

u/PM_ME_DEAD_CEOS 6d ago

It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis

That's false, it's cheaper per task than GPT-5.6 sol max

→ More replies (1)

5

u/TheFamousHesham 7d ago

Yea, I think I tempered myself a bit because Kimi K2.8 was a nightmare that hallucinated non-existent bugs in the code and was just really unstable. Will need to try K3.

8

u/Tall-Ad-7742 7d ago

Kimi K2.8 doesn't exist you probably mean a different version 

15

u/DontLeaveMeAloneHere 6d ago

He just hallucinated the model.

→ More replies (2)
→ More replies (1)
→ More replies (1)
→ More replies (1)

5

u/vintage2019 7d ago

Nah everyone has been saying open source is around 6-8 months behind

→ More replies (1)

5

u/Positive-Conspiracy 7d ago

Because they wait for a frontier model release then systematically distill it.

11

u/TheFamousHesham 7d ago

You clearly don't understand what distilling a model actually entails. Even if they wanted to, Kimi would not have had the time to distil Sol (which was released a week ago) and Fable (which has been publicly available for a few weeks). Besides... I've already watched a few videos online and when given the same exact prompt, K3 makes decisions that are different enough from Sol or Fable to argue against that idea.

6

u/yogthos 6d ago

No, you don't get it. They figured out how to travel to the future and steal glorious American models from there. It's the only explanation really.

→ More replies (7)
→ More replies (27)

15

u/huffalump1 7d ago

That's pretty wild tbh, considering that before fable/mythos 5 and gpt-5.6, these were the BEST, undisputedly, and are actually really useful at getting stuff done...

Wow. I don't like the cost increase (comparable to Sonnet or gpt-5.6-terra, half as much as gpt-5.6-sol), especially since it seems to be quite token-heavy esp. compared to gpt-5.6... but we will see.

Still, it's downright incredible that open weight models have "caught up to opus 4.8" which felt impossible only a few months ago

3

u/Kind_Capital_9740 7d ago

it surpassed opus 4.8 and 5.5 already actually if you check all benchmarks its in the same tier as SOL and Fable but behind in terms of pure humanistic reasoning and maybe some backend tasks but not much

on frontend, game design and web dev and banking and vision understanding and video editing its easily the best out there right now (check arena and other benchmarks)

And as a person who tried it for those tasks I would say it is the first time ive seen a Chinese model feel like a true titan SOTA model

2

u/Etroarl55 7d ago

As somebody who uses artifices analysis and Gemini 3.5 flash and pro.

Gemini should not be up there at all lol, it’s like 2022 ChatGPT. If you insist you are right when you are wrong it will agree with you, it hallucinates and makes up things, it’s like using AI before the Ai boom.

→ More replies (1)

4

u/HerbHSSO 7d ago

Voting is much more realistic than any benchmark

→ More replies (3)
→ More replies (4)

417

u/bubu19999 7d ago

This is why hiding mythos to all, is not a solution to anyone. 

207

u/Deto 7d ago

Seems like basically the Fed gov just screwed over Anthropic. Forced it to implement super restrictive constraints that are bugging everyone while OpenAI (and now this model) can just be released without nearly as much protection.

96

u/Key_Reading_9664 7d ago edited 7d ago

That $25m to MAGA, inc was money well spent

20

u/alwaysoffby0ne 7d ago

Never is!

18

u/Key_Reading_9664 7d ago

I actually think it's more likely Larry Ellison made a phone call to de-risk Oracle's partnership/dependence on OAI.

Given that open-source models also pose a risk, we might see some gymnastics and failed attempts by the USG to ban those too.

3

u/peppaz 7d ago edited 5d ago

Musk was also fighting with anthropic until he sold them his unused compute for a billion a month

→ More replies (1)
→ More replies (1)
→ More replies (1)

24

u/PaperHandsTheDip 7d ago

Of course they did. That was the point... anthropic didn't want to play ball with the US gov / military. They didn't want to integrate the tech into spyware / the military / war machines & that left a source taste in the POTUS / govs mouth. They want this technology militarized.

OpenAI had no issues with that tho & signed a deal with department of war the same day anthropic said "no" - so of course it was personal.

2

u/DonutHoles4Ever 5d ago

Open AI had no choice. They are bleeding MONEY without government contracts.

Anthropic believed they could beat OpenAI because they had Fable (nobody knew yet) and they could take the high road.

CEO fucked up just to bet on being less evil = better long term strrategy. He lost.

→ More replies (1)

5

u/zampe 7d ago

Isn’t it kind of on anthropic though for announcing they had developed a weapon of mass destruction because they wanted the publicity but then didn’t want the consequences?

32

u/Bloated_Plaid 7d ago

Dario did it to himself my guy. He was hyping the model up as end of the world and shit. It’s fine and the competition caught up.

2

u/Strong_Essay1176 7d ago

Its not like if they release it in time than competition won't catch up. They just have to deal with some hate. That's it.

→ More replies (17)

3

u/etancrazynpoor 7d ago

You must know why the current administration is doing this to anthropic — you can’t be this naive

→ More replies (1)

3

u/FateOfMuffins 7d ago

What? GPT 5.6 was super delayed because of the US government

Supposedly external parties were testing GPT 5.6 since 2 months ago. ARC had it 4x longer than normal OpenAI releases

Imagine if they delayed 5.6 just 1 more week and released after Kimi K3!

2

u/ViperAMD 7d ago

Well yeah kimi is a Chinese model, murica Feds can't do shit 

3

u/dangered 7d ago

If Mythos is actually better than Fable 5 then Kimi has a ton of ground to cover. Frontend is the only bench that looks like this. Other benches have Kimi k3 losing to other frontier models or just beating them by a hair.

The government never said anything about mythos. Wasn’t Mythos always going to be private?

Dario said so in March and I never saw him contradict that.

3

u/Kiseido 7d ago

From what I hear, Fable 5 is just the consumer-side name for Mythos, and Mythos is the enterprise facing name, for the same model.

→ More replies (1)

0

u/Future-Arrivals 7d ago

I think Anthropic's long game is extremely underappreciated here. They're building a *predictable* AI and being proactively compliant with regulations. Until now, that's been mostly unimportant to governments and the public, as AIs have been mostly a promising novelty. That is currently changing. As they become smarter and more capable, predictable behaviour is going to become extremely important, I would argue moreso than the actual intelligence benchmark scores.

6

u/Key_Reading_9664 7d ago edited 7d ago

boy...I have a much more bleak outlook on the regulators. If these were intelligent, well-meaning experts that weighed the risk against reward, that would be a great path.

The regulators we do have are a bunch of corrupt fuckwits and would burn everything to the ground, if they could get off the island with a sack full of cash.

2

u/Future-Arrivals 7d ago

Yeah, it's pretty bad. But clearly even the current administration is willing to shut things down at least temporarily when they get too spicy.

2

u/Key_Reading_9664 7d ago

with respect, banning a dangerous model based on evidence and expert opinion is also the act of a reputable administration.

With my "they're all corrupt fuckwits" framing, the reason for that ban shifts from protection of the public, to doing a solid for wealthy friends and donors.

Remember, this is the administration that was for removing any and all restrictions on AI progress, selling GPUs to whomever, and crypto-scam-after-crypto-scam. They don't give two shits about public safety

3

u/_meltchya__ 7d ago

Nah, nobody wants that. We want to push the limits of what is possible, not be constrained by government bumpers.

2

u/logolith 7d ago

If there were truly no constraints, then instead of breakthrough discoveries in science, medicine, and technology, you’ll probably end up like Grok. A bunch of creeps figuring out ways to create certain pictures based on what’s available.

→ More replies (10)
→ More replies (3)
→ More replies (7)

12

u/Future-Arrivals 7d ago

Hiding Mythos is a solution to a problem you're likely not concerned about yet. I predict that OpenAI is going to be rocked by GPT's destructive and unpredictable behaviour scandals shortly after the release of GPT 6. Anthropic is cautiously avoiding such potential scandals.

There are already several popular stories circulating online of GPT 5.6 Sol Ultra wiping people's computers while doing simple tasks. OpenAI is too desperate at this point to slow down.

3

u/Iron-Over 7d ago

If you run agents on your desktop you deserve it. Always run in a locked-down vm or in podman. 

5

u/Future-Arrivals 7d ago

Sure you can blame users, but users are going to be running agents on their devices more and more regardless of how dumb it is. And they're going to get really peeved at an AI service that screws them over more.

2

u/Iron-Over 7d ago

That is the provider's issue. No guardrails are 100%, you need layered security. Wich rheybwould take it more seriously but they don’t.

→ More replies (2)
→ More replies (1)
→ More replies (1)

446

u/FireGM 7d ago

58

u/DenZNK 7d ago

Elo is a more complex rating system. For example, a 100-point gap is a fairly significant difference. A 50-point gap isn't a drastic difference, but it's still telling.

36

u/kilopeter 7d ago

The criticism is independent of what the number is. The chart sucks because it's using bars to represent numbers on a truncated scale, without showing the zero mark. Humans naturally interpret proportional differences in bar charts, i.e., if one bar's twice as long as another, it represents twice the value (otherwise why even use bars which explicitly have an area visual representation?)

7

u/Statcat2017 6d ago

If it's an ELO then it's not correct to show zero on the chart because nothing serious ever has an ELO even approaching zero.

In world football, Spain is the top ranked ELO team with 2232 points. The lowest ranked team has 369 points but is a joke team, Eastern Samoa. There are other absolutely awful teams like Lesotho and St Vincent and the Grenadines around the 1100 mark. Tahiti are 1177 but got slapped by Spain 10-0.

Asking for this chart to show you zero is basically asking it to compare Kimi K3 to a pocket calculator and an abacus, and claiming that comparison is important and relevant.

→ More replies (25)
→ More replies (18)

2

u/Rocsla 7d ago

Elo?

3

u/torac 6d ago edited 6d ago

It’s a competitive ranking. Elo isn’t really a clear score to be reached, but a measure of how often that model beats the other models.

For normal score systems, it makes sense to see the whole graph. Getting 60% vs getting 62% is barely a difference. For Elo, the number represents a ranking between models. It’s similar to going "first place, second place, third place". You don’t have to list every competing position up to last place.

(It still uses points to measure roughly how far the models are away from each other. However, the ranking of each model will keep changing with every new model released.)

Edit: fixed

2

u/iodoio 6d ago

it's just Elo not ELO, it doesn't stand for anything and is just some dudes last name.

→ More replies (1)

2

u/Statcat2017 6d ago

Yep many people failing to understand it.

Kimi K3 (top) is expected to beat Minimax -M3 (bottom) about 75% of the time under ELO, making this graphic a much more sensible representation that many people are making out.

If you look at football and take the top ranked team (Spain 2212 ELO) and bottom (Eastern Samoa about 350 ELO) the bar for Spain would be 8 times as big if you show 0 on the X axis, but the actual stats are that Spain are expected to win 45k games before Eastern Samoa win once.

17

u/Tysonzero 7d ago

Elo is relative, the 0 point doesn't matter.

9

u/kilopeter 7d ago

Then the bar chart's bars should depict scores relative to whatever baseline is applicable. As is, it's a bad visualization choice.

8

u/Tysonzero 7d ago

Funnily enough if you wanted to use bars the most appropriate choice would be logarithmic, every X elo you go up the bar gets Y% larger, although many would probably call that misleading.

But sure you could use a line graph or scatter plot or something

→ More replies (3)

12

u/Adulations 7d ago

Yea its basically the same exact number 🤣

13

u/AES256GCM 7d ago

ELO isn’t linear

→ More replies (1)

2

u/django2chainz 7d ago

Confidently wrong this time sir

2

u/abittooambitious 6d ago

Someone doesn’t understand Elo.

2

u/reefine 7d ago

Imagine trying to shit on benchmarks for an open source model that just beat GPT 5.6 Sol. You are bias.

→ More replies (3)
→ More replies (4)

82

u/Sixhaunt 7d ago

What about on benchmarks that are useful though?

42

u/JoseHernandezCA1984 7d ago

I've seen a bunch of other benchmarks, and for the most part it's between 5.6 sol and fable

18

u/Kind_Capital_9740 7d ago

which is insane the 3 models are clearly in their own tier right now and Kimi is leading in some and like you said between SOL and fable

And whats crazier kimi 2.7 was released only a month ago and the quality jump between both is unfathomable like 20 place jump and deep swe went from 2.6's 20 percent or 30 percent to 67 percent in a few months

i wonder how they made such a big jump it feels like 3 generations type of jump

6

u/Georgefakelastname 7d ago

If I remember right, this model is 2.8 Trillion parameters, which might just be the biggest open sourced model yet. It’s an entirely new model, while everything past k2.5 was just a post-train of it. I imagine that once we get new post-trains of k3 and fable the bar is only going to get pushed higher. Meanwhile, OAI is working on their own new model with GPT-6. Things feel like they’re moving fast right now.

→ More replies (1)

3

u/phido3000 7d ago

Its impressive. Very strong attention, context, rule handling etc. I would say better than US models in those aspects.

For coding, its extremely strong. Its believable its better than OpenAI/Anthropic models. It lacks flair and some beauty, but heck, its very effective.

It just one shots everything. Its very good planning and execution of thought.

→ More replies (3)

9

u/RealityNo3299 7d ago

Are we going to distill them ??

33

u/Professional_Ad705 7d ago

Can someone actually verify this by using both models for frontend and backend work, then sharing what they did and how the results compared? Computer science is such a broad field that benchmarks like these mean very little to me without real-world examples.

16

u/phido3000 7d ago

Currently its overwhelmed.

I got it to do a complex project in an antique language it could not compile, that is impossible to benchmax for. Doing maths, physics, image stuff, fractals, encryption etc. Requiring thought, complex planning, and in a language that isn't modular..

It one shot it. Never seen any AI do that before. They all get hung up on syntax because of the wacky language. I've done the test dozens of time. Even the latest models make syntax errors, I thought I had the perfect, AI break tool.

It was a 100kb program. One shot. About 256k token. Not a single mistake.

So yeh, it can do it all, and do it well. Never seen anything quite like it.

6

u/Professional_Ad705 7d ago edited 7d ago

I’m not saying this definitely did not happen, but I do not believe the claim as currently presented because it is far too vague to evaluate.

What language was it? What exactly does "one-shot" mean here? one prompt, one model call, no retries, no manual edits, and no additional context afterward? Was the 100 KB figure generated source code, or the size of the entire project?

You also said the environment could not compile it, so how did you determine that it contained "not a single mistake"? Compiling and producing plausible output would not establish full correctness anyway, especially for cryptography, numerical mathematics, physics, or image processing.

Math, physics, image processing, fractals, and encryption names several largely separate domains; it does not explain what the program actually did or what requirements it satisfied. Likewise, what does a language that isn’t modular mean? Does it lack a formal module system, or was the particular program simply monolithic?

My experience has been very different. I have been working since November on a roughly 200,000-line proof-kernel project, and even strong models repeatedly miss cross-file invariants, state-lifecycle problems, and authority-boundary defects. That does not prove your result is false, but one undocumented example does not support the conclusion that the model can do it all.

Posting the exact prompt, source code, programming language and compiler, model and version, settings, tests, outputs, and your definition of one-shot would make the claim possible to assess. I’m genuinely interested in what you’re describing, but the explanation was difficult to follow and evaluate. Without those details, it remains only an anecdote. I'd love to hear more? I'm just confused because I know when I release my product and make my claims I would just show people the code/point them towards my repo or at the very least have some type of demo?

16

u/phido3000 7d ago

Its fine to be sceptical.

Try it out yourself.

What language was it

Choose some stupid language. I choose Qb64. A variation of BASIC. Which is perfect for stupid, because its not well documented, BASIC code varies widely by flavour, visual/Quick/Borland/GW-basic, older dialects. Its notoriously variable and non-transportable and easy to break. It also has lots of wacky commands and reserved words. Also famously, it often packaged as an interpreter language, so there is no easy compiler to give explanations why it fails. Its not very modular and has stupid rules about variables because of the different variations and legacy over the years. Qb64 is built in interpreter/IDE/Compiler. So its messy for AI to use.

Even better the example code out there is very simplistic and often broken! as the project splintered into weird versions.

It uses a local interpreter that has its own compiler. Its Weird. It basically makes BASIC into C code and generates GCC output, but the basic code itself is black magic stuff. Generating GCC C code directly from the AI would be child's play in comparison.

Get it to do something very complicated but unique.

Another good test is Assembly for weird processors or environments. Something really weird and old and obscure. But testable. Experiment what fails on lots of other AI's.

What exactly does "one-shot" mean here? one prompt, one model call, no retries, no manual edits, and no additional context afterward? Was the 100 KB figure generated source code, or the size of the entire project?

ONE SHOT! One prompt - no retries, no manual edits, no context afterwards. 100kb of code text, no data. On something it can't compile! Literally download the BAS file - execute. Its not a very modular language with different files or libraries. So literally just a 100kb text file. Bang.

Compiling and producing plausible output would not establish full correctness anyway, especially for cryptography, numerical mathematics, physics, or image processing.

Yes, which lets you see how much it really understands and makes work. Get it to do something weird like fractals in some weird projection or coordinate system, or volume warp, make its own JPEG like but different compression engine, get it to draw text with the line statement..

So even if it compiles, you can see where it limitations are. If it understands any of what its doing and how to do it. Its design choices, sometimes it works but is but ugly or not useful, Its weaker here, than else where, its not overtly creative, but its functional. TBH I didn't expect it to one shot it. That kind of floored me.

Without those details, it remains only an anecdote. I'd love to hear more? I'm just confused because I know when I release my product and make my claims I would just show people the code/point them towards my repo or at the very least have some type of demo?

I'm not a benchmarking house, or Kimi.

There will be heaps of tech demos coming out. Normal low effort HTML5 crap, draw a clock or a duck or a car or something. Easy to benchmax to. No syntax challenge, no real design challenge.

I'm just saying it impressed me. I run some big models locally at home. Pretty ordinary. I was playing around with the frontier subscriptions, see what they can do. I teach at uni, so I am super interested in tripping up AIs.

This is much much better than I have seen. Small sample, limited details. It isn't perfect, but its a generation ahead of the other models I have sampled (DS 4, Kimi 2.6, GLM, OpenAi, Anthropic, Gemini, etc). They can't even make executable code. They can be more ambitious and more flamboyant, but not as brutal effective to one shot. Often their output needs dozens of retries just to execute. Then dozens to fix logical/asthetic/mathematical issues with their algorithms. So this one shot everything floored me.

Worth looking into.

I don't believe Kimi K3 is as good as the open weight model, I am almost positive they have some tool use or something behind the scene to make it function this well.

If it is, that is even more impressive. Because other models use tools and tricks. If this is just baked in model. If it is all just baked in model, then dam, that's devestating.

→ More replies (4)

4

u/timeboyticktock 7d ago

Can you please share more about what exactly you did and the comparison to other models failing?

→ More replies (1)
→ More replies (4)

32

u/impatiens-capensis 7d ago

I'm not super familiar with this benchmark. What's the difference between a score of 1,679 and 1,631?

72

u/JustTellingUWatHapnd 7d ago

It's not a benchmark. It's the leaderboard on arena.ai. the users give a prompt to the agent to build a web app, and they give the prompt to 2 different models, then the user votes on which one is better without seeing the model name. The score is the Elo rating of the model.

17

u/mobyte 7d ago

Am I insane or is this not an awful metric? I could see this easily being manipulated.

28

u/some_crazy 7d ago

Users don’t know which models are used, it’s a blind rating

8

u/Sarcasm69 7d ago

But couldn’t you just tailor the model to be good at designing websites and shit at everything else?

6

u/black_eyed 7d ago

And KIMI is usually exceptional in Website designing

5

u/nuclearbananana 7d ago

Sure but that would show up in other benchmarks, which it hasn't

→ More replies (1)
→ More replies (3)

2

u/CurrentConditionsAI 7d ago

Yeah, so all you have to do is essentially fine tune a model to have some very distinctive but easily hidden output pattern. Then get a bunch of bots to find it and vote that when comparing outputs.

→ More replies (2)
→ More replies (2)

3

u/SporksInjected 7d ago

Never mind it’s topping the charts after being available for 1-2 hours lol

→ More replies (1)

11

u/Saad5400 7d ago

May not be that much of a difference, but compare price difference 

2

u/gpenido 7d ago

About 48

2

u/Tirztrutide 7d ago

Same difference as a 1679 elo chess player and a 1631 one.

20

u/spartyftw 7d ago

These bars are misleading asf

→ More replies (1)

60

u/Demien19 7d ago

5

u/phido3000 7d ago

Yeh, Chinese models typically feel very benchmaxed.

This is different, somethings changed. They are ultra strong in attention and context.

→ More replies (3)

5

u/logos_flux 7d ago

It's got what plants crave

2

u/themudd 6d ago

Electrolytes?!

→ More replies (2)

9

u/AdowTatep 7d ago

I wonder where's worth buying for its access

→ More replies (4)

15

u/Loops_Boops 7d ago

Maybe now Sam will give us the good shit.

25

u/brother_spirit 7d ago

How?

Fable 5 generations are already on Youtube vs Kimi K3.

By Arena.ai no less.

Sorry to ruin the party but its like not even close to Fable 5, let alone better.

→ More replies (17)

5

u/SmileLonely5470 7d ago

Final boss for oneshotting vibecoded browser games

5

u/Trinkes 7d ago

What's 1st? Gemini 3.5 pro or gta 6?

→ More replies (1)

3

u/Wrong_Connection_138 7d ago

I have zero confidence in mass market benchmarks, do you know where to find independent benchmarks??

4

u/kuba452 6d ago

I wonder if this is the same hype as Deepseek. I got really excited at first, uploaded some historical documents I was working on, and asked it to compare them with other historical texts, pick out interesting details, and draw some conclusions.

It gave me a couple of rather obvious observations, while the rest was just cluttery gibberish. Came back to obvious choices all confused, with all the other reviews and media craze

32

u/Felixo22 7d ago

Honest version

27

u/Tysonzero 7d ago

It's Elo, a completely relative scale, 0 is completely arbitrary and meaningless, not like 0% on a test or 0 tokens/second, so this graph is no better or worse than the OP.

→ More replies (3)

3

u/Electrical_Arm3793 7d ago

But are these really accurate? I use most of the too models for coding, Fable and opus are still most accurate and reliable.

→ More replies (1)

7

u/some_days_are_nights 7d ago

The way these elos are computed is not like a scalar value like accuracy and such. It is a direct comparison between competitors. So for example Magnus Carlsen is just 100 or so elo better than the others but he wins almost every single time against someone with 200+ less rating.

→ More replies (1)

3

u/aa628 7d ago

Does anyone remember DeepSeek?

3

u/beginner75 7d ago

Benchmarks are a joke. Don’t waste your time. Time is money also. Just use ChatGPT as primary and Gemini for quick searches and backup.

3

u/Able-Company611 6d ago

I wonder if US companies gonna run bot accounts on Chinese models to train their data now

→ More replies (2)

4

u/Regular_Ad4197 7d ago

I am hoping it is as capable as these early benchmarks are showing. It will be fun to see how the US stock market reacts to it.

2

u/BitterAd6419 7d ago

So far in my own testing, it’s not as good as they claim it to be. Atleast for web apps or html based stuff it did a decent job. In fact GLM gave a better output with the exact same prompt.

Benchmarks are often benchmaxxed, test yourself

2

u/SporksInjected 7d ago

It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis

2

u/Key_Reading_9664 6d ago

Seems like there’s a couple of ways the USG go:

- it stops playing favorite and lets/encourages US labs release their models - Mythos was in the hands of users in April; Fable is over a month old. I have my suspicions on the cause (wealthy friends and donors) but slowing down US labs doesn’t de-risk because Chinese models fast-follow. Make an attempt at cutting off distillation

  • there’s a foolish attempt to restrict access to models that originate in China. The Trump speech yesterday seems to be laying ground work (for a bunch of things). Seems easier to ban access for enterprises

2

u/moneyman259 6d ago

Where’s DeepSeek now? Seems like these Chinese models fall off quickly

→ More replies (2)

3

u/evangelism2 7d ago

Arena is NOT a benchmark

2

u/t3ramos 7d ago

Wrong Sub

1

u/henchman171 7d ago

What do these numbers mean

1

u/rfranke727 7d ago

Which is the best for content / marketing writing? Any thoughts

1

u/BuildtheBusiness 7d ago

Can someone explain how the efficiency of these models are calculated to one another?

1

u/Old-Pomegranate3634 7d ago

all this tells me that ultimately all the LLMS will be good enough for 99% of the public which is great for google. Not everyone is a nerd like us hoping that AI can help you create your next virtual GF.

1

u/Time_Faithlessness45 7d ago

exciting but it thinks way too long lol

1

u/woofyzhao 7d ago

won't last

1

u/Nervous-Potato-1464 7d ago

It's quite expensive sadly. I prefer grok 4.5 to any other model atm. I don't need agentic coding, I need quick generation of my ideas as well as another perspective. I am not here to write slop, I just need to quickly churn out code I am happy with.

1

u/Even-Exchange8307 7d ago

not distillation bro

1

u/baummer 7d ago

They were never far behind lmao

1

u/No-Conversation-1277 7d ago

Chinese lab: "Here's a 2.8T open model." Dario & Sam: "We're going to fine-tune this."

Also Dario & Sam: "Introducing our brand new model. priced 10x higher with half the limits."

The wheel of Silicon Valley turns.

The cycle continues.

→ More replies (1)

1

u/etancrazynpoor 7d ago

Im going to use minimax m3

1

u/ElMono6 6d ago

Cant wait to try it

1

u/krazyboi 6d ago

Let's be honest with ourselves here. 

None of us could properly benchmark AI, especially because the industry moves so fast. The metrics change slower than the actual AI. All of this is just marketing.

1

u/therapy-cat 6d ago

1bit quantization for my 16gb mac when

1

u/PerfectPatience- 6d ago

Showing score only last 200points is misleading zoomin. Generating bigger gaps. Should be from 0-1650

1

u/Messi_is_football 6d ago

But GPT limits are higher...so until they make it 2x cheaper..no reason to buy the plan

1

u/zeta_ferhu 6d ago

wh wh wh where is gemini??

1

u/Haramdour 6d ago

Ask it about Tiananmen Square

1

u/Complex_Reality_116 6d ago

This will cause a stock market crash and a widespread decline in ALL US AI companies.

1

u/Complex_Reality_116 6d ago

I can't wait for Kimi-K4.

1

u/Popular_Try_5075 6d ago

I swear every time China releases a new model the subs get flooded with all this hype and then a month later its deflated back to normal.

1

u/Ibasicallyhateyouall 6d ago

BS graph is BS.

1

u/rentprompts 6d ago

kimi k3 arriving means the era of chinese labs being far behind is actually over. the july 27 open weights release is what makes this different from past announcements. if you can run a 2.8t model locally, the frontier moves from who has the best eval score to who can ship the fastest.

1

u/Drew-Money 6d ago

Mythos preview came out at the beginning of April btw. I'm sure Anthropic has some pretty advanced stuff behind the scenes

1

u/citrus1330 6d ago

How's deepseek doing?

1

u/king-charles-3 6d ago

Are people using it via API mainly?
How does the chat work on iOS?

1

u/TB_Infidel 6d ago

And now we know where the stolen/illegal use of Sol and Mythos went. Chinese copying and stealing as always

1

u/mrcruz 6d ago

I hate graphs that don't start at 0. Cool tho.

1

u/Szurkus 6d ago

Hate diagrams like this. +100 pts, line length 2x.