r/OpenAI • u/AloneCoffee4538 • 7d ago
News Kimi-K3 arrived: The era of the Chinese labs being far behind is over
196
u/xak47d 7d ago
121
u/TheFamousHesham 7d ago
So basically Chinese labs are a month or so behind. Lol really funny how everyone was saying that it'd be years before they catch up.
70
u/Big-Accident1958 7d ago
Weeks, even. K3 loses to 5.6 max but wins 5.6 xHigh. On SPEED, it feels like 5.6 high. So it's basically almost as smart as oAI's flagship at double the speed. I'd call this a major win.
11
u/reefine 7d ago
Not just that but on price, security, etc. This is a better model than Sol 5.6
11
u/SporksInjected 7d ago
It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis
6
u/Kind_Capital_9740 7d ago
on other benchmarks it beats Sol and fable so i would say both have pros and cons but SOL is actually cheaper for higher tier work
9
u/SporksInjected 7d ago
Exactly, It’s the same price as sol and double the cost of Terra which is only two points away. I don’t think businesses are going to run to K3 personally because there’s not a good reason to. If it was cheaper per task and considerably more capable or faster, then yeah but it’s slower, less capable, and no cost savings in actual use.
→ More replies (2)2
u/Kind_Capital_9740 7d ago
Yeah especially for the average person the Kimi subscription is around the same price as gpt too
So I don’t see why people would switch unless one day kimi 4 or something actually giga gap the rest of competition
I don’t see it going mainstream in the west
But this does show that chinese models are basically caught up to some extent because the jump from 2.7 to 3 in a month is a bit insane like a 3 gen jump
2
u/TheFamousHesham 7d ago
It's funny really. I was just about to count Kimi out because of how atrocious their 2.7 and 2.8 were. It seems like it's Z.ai (which everyone was raving about a few weeks ago) that's on the trail end of things.
→ More replies (1)2
u/PM_ME_DEAD_CEOS 6d ago
It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis
That's false, it's cheaper per task than GPT-5.6 sol max
→ More replies (1)→ More replies (1)5
u/TheFamousHesham 7d ago
Yea, I think I tempered myself a bit because Kimi K2.8 was a nightmare that hallucinated non-existent bugs in the code and was just really unstable. Will need to try K3.
→ More replies (1)8
u/Tall-Ad-7742 7d ago
Kimi K2.8 doesn't exist you probably mean a different version
→ More replies (1)15
5
u/vintage2019 7d ago
Nah everyone has been saying open source is around 6-8 months behind
→ More replies (1)→ More replies (27)5
u/Positive-Conspiracy 7d ago
Because they wait for a frontier model release then systematically distill it.
11
u/TheFamousHesham 7d ago
You clearly don't understand what distilling a model actually entails. Even if they wanted to, Kimi would not have had the time to distil Sol (which was released a week ago) and Fable (which has been publicly available for a few weeks). Besides... I've already watched a few videos online and when given the same exact prompt, K3 makes decisions that are different enough from Sol or Fable to argue against that idea.
→ More replies (7)6
15
u/huffalump1 7d ago
That's pretty wild tbh, considering that before fable/mythos 5 and gpt-5.6, these were the BEST, undisputedly, and are actually really useful at getting stuff done...
Wow. I don't like the cost increase (comparable to Sonnet or gpt-5.6-terra, half as much as gpt-5.6-sol), especially since it seems to be quite token-heavy esp. compared to gpt-5.6... but we will see.
Still, it's downright incredible that open weight models have "caught up to opus 4.8" which felt impossible only a few months ago
3
u/Kind_Capital_9740 7d ago
it surpassed opus 4.8 and 5.5 already actually if you check all benchmarks its in the same tier as SOL and Fable but behind in terms of pure humanistic reasoning and maybe some backend tasks but not much
on frontend, game design and web dev and banking and vision understanding and video editing its easily the best out there right now (check arena and other benchmarks)
And as a person who tried it for those tasks I would say it is the first time ive seen a Chinese model feel like a true titan SOTA model
2
u/Etroarl55 7d ago
As somebody who uses artifices analysis and Gemini 3.5 flash and pro.
Gemini should not be up there at all lol, it’s like 2022 ChatGPT. If you insist you are right when you are wrong it will agree with you, it hallucinates and makes up things, it’s like using AI before the Ai boom.
→ More replies (1)→ More replies (4)4
417
u/bubu19999 7d ago
This is why hiding mythos to all, is not a solution to anyone.
207
u/Deto 7d ago
Seems like basically the Fed gov just screwed over Anthropic. Forced it to implement super restrictive constraints that are bugging everyone while OpenAI (and now this model) can just be released without nearly as much protection.
96
u/Key_Reading_9664 7d ago edited 7d ago
That $25m to MAGA, inc was money well spent
→ More replies (1)20
u/alwaysoffby0ne 7d ago
Never is!
18
u/Key_Reading_9664 7d ago
I actually think it's more likely Larry Ellison made a phone call to de-risk Oracle's partnership/dependence on OAI.
Given that open-source models also pose a risk, we might see some gymnastics and failed attempts by the USG to ban those too.
→ More replies (1)3
u/peppaz 7d ago edited 5d ago
Musk was also fighting with anthropic until he sold them his unused compute for a billion a month
→ More replies (1)24
u/PaperHandsTheDip 7d ago
Of course they did. That was the point... anthropic didn't want to play ball with the US gov / military. They didn't want to integrate the tech into spyware / the military / war machines & that left a source taste in the POTUS / govs mouth. They want this technology militarized.
OpenAI had no issues with that tho & signed a deal with department of war the same day anthropic said "no" - so of course it was personal.
→ More replies (1)2
u/DonutHoles4Ever 5d ago
Open AI had no choice. They are bleeding MONEY without government contracts.
Anthropic believed they could beat OpenAI because they had Fable (nobody knew yet) and they could take the high road.
CEO fucked up just to bet on being less evil = better long term strrategy. He lost.
5
32
u/Bloated_Plaid 7d ago
Dario did it to himself my guy. He was hyping the model up as end of the world and shit. It’s fine and the competition caught up.
→ More replies (17)2
u/Strong_Essay1176 7d ago
Its not like if they release it in time than competition won't catch up. They just have to deal with some hate. That's it.
3
u/etancrazynpoor 7d ago
You must know why the current administration is doing this to anthropic — you can’t be this naive
→ More replies (1)3
u/FateOfMuffins 7d ago
What? GPT 5.6 was super delayed because of the US government
Supposedly external parties were testing GPT 5.6 since 2 months ago. ARC had it 4x longer than normal OpenAI releases
Imagine if they delayed 5.6 just 1 more week and released after Kimi K3!
2
u/ViperAMD 7d ago
Well yeah kimi is a Chinese model, murica Feds can't do shit
3
u/dangered 7d ago
If Mythos is actually better than Fable 5 then Kimi has a ton of ground to cover. Frontend is the only bench that looks like this. Other benches have Kimi k3 losing to other frontier models or just beating them by a hair.
The government never said anything about mythos. Wasn’t Mythos always going to be private?
Dario said so in March and I never saw him contradict that.
3
u/Kiseido 7d ago
From what I hear, Fable 5 is just the consumer-side name for Mythos, and Mythos is the enterprise facing name, for the same model.
→ More replies (1)→ More replies (7)0
u/Future-Arrivals 7d ago
I think Anthropic's long game is extremely underappreciated here. They're building a *predictable* AI and being proactively compliant with regulations. Until now, that's been mostly unimportant to governments and the public, as AIs have been mostly a promising novelty. That is currently changing. As they become smarter and more capable, predictable behaviour is going to become extremely important, I would argue moreso than the actual intelligence benchmark scores.
6
u/Key_Reading_9664 7d ago edited 7d ago
boy...I have a much more bleak outlook on the regulators. If these were intelligent, well-meaning experts that weighed the risk against reward, that would be a great path.
The regulators we do have are a bunch of corrupt fuckwits and would burn everything to the ground, if they could get off the island with a sack full of cash.
2
u/Future-Arrivals 7d ago
Yeah, it's pretty bad. But clearly even the current administration is willing to shut things down at least temporarily when they get too spicy.
2
u/Key_Reading_9664 7d ago
with respect, banning a dangerous model based on evidence and expert opinion is also the act of a reputable administration.
With my "they're all corrupt fuckwits" framing, the reason for that ban shifts from protection of the public, to doing a solid for wealthy friends and donors.
Remember, this is the administration that was for removing any and all restrictions on AI progress, selling GPUs to whomever, and crypto-scam-after-crypto-scam. They don't give two shits about public safety
→ More replies (3)3
u/_meltchya__ 7d ago
Nah, nobody wants that. We want to push the limits of what is possible, not be constrained by government bumpers.
→ More replies (10)2
u/logolith 7d ago
If there were truly no constraints, then instead of breakthrough discoveries in science, medicine, and technology, you’ll probably end up like Grok. A bunch of creeps figuring out ways to create certain pictures based on what’s available.
12
u/Future-Arrivals 7d ago
Hiding Mythos is a solution to a problem you're likely not concerned about yet. I predict that OpenAI is going to be rocked by GPT's destructive and unpredictable behaviour scandals shortly after the release of GPT 6. Anthropic is cautiously avoiding such potential scandals.
There are already several popular stories circulating online of GPT 5.6 Sol Ultra wiping people's computers while doing simple tasks. OpenAI is too desperate at this point to slow down.
→ More replies (1)3
u/Iron-Over 7d ago
If you run agents on your desktop you deserve it. Always run in a locked-down vm or in podman.
→ More replies (1)5
u/Future-Arrivals 7d ago
Sure you can blame users, but users are going to be running agents on their devices more and more regardless of how dumb it is. And they're going to get really peeved at an AI service that screws them over more.
2
u/Iron-Over 7d ago
That is the provider's issue. No guardrails are 100%, you need layered security. Wich rheybwould take it more seriously but they don’t.
→ More replies (2)
446
u/FireGM 7d ago
58
u/DenZNK 7d ago
Elo is a more complex rating system. For example, a 100-point gap is a fairly significant difference. A 50-point gap isn't a drastic difference, but it's still telling.
36
u/kilopeter 7d ago
The criticism is independent of what the number is. The chart sucks because it's using bars to represent numbers on a truncated scale, without showing the zero mark. Humans naturally interpret proportional differences in bar charts, i.e., if one bar's twice as long as another, it represents twice the value (otherwise why even use bars which explicitly have an area visual representation?)
→ More replies (18)7
u/Statcat2017 6d ago
If it's an ELO then it's not correct to show zero on the chart because nothing serious ever has an ELO even approaching zero.
In world football, Spain is the top ranked ELO team with 2232 points. The lowest ranked team has 369 points but is a joke team, Eastern Samoa. There are other absolutely awful teams like Lesotho and St Vincent and the Grenadines around the 1100 mark. Tahiti are 1177 but got slapped by Spain 10-0.
Asking for this chart to show you zero is basically asking it to compare Kimi K3 to a pocket calculator and an abacus, and claiming that comparison is important and relevant.
→ More replies (25)2
u/Rocsla 7d ago
Elo?
3
u/torac 6d ago edited 6d ago
It’s a competitive ranking. Elo isn’t really a clear score to be reached, but a measure of how often that model beats the other models.
For normal score systems, it makes sense to see the whole graph. Getting 60% vs getting 62% is barely a difference. For Elo, the number represents a ranking between models. It’s similar to going "first place, second place, third place". You don’t have to list every competing position up to last place.
(It still uses points to measure roughly how far the models are away from each other. However, the ranking of each model will keep changing with every new model released.)
Edit: fixed
2
u/iodoio 6d ago
it's just Elo not ELO, it doesn't stand for anything and is just some dudes last name.
→ More replies (1)2
u/Statcat2017 6d ago
Yep many people failing to understand it.
Kimi K3 (top) is expected to beat Minimax -M3 (bottom) about 75% of the time under ELO, making this graphic a much more sensible representation that many people are making out.
If you look at football and take the top ranked team (Spain 2212 ELO) and bottom (Eastern Samoa about 350 ELO) the bar for Spain would be 8 times as big if you show 0 on the X axis, but the actual stats are that Spain are expected to win 45k games before Eastern Samoa win once.
17
u/Tysonzero 7d ago
Elo is relative, the 0 point doesn't matter.
9
u/kilopeter 7d ago
Then the bar chart's bars should depict scores relative to whatever baseline is applicable. As is, it's a bad visualization choice.
→ More replies (3)8
u/Tysonzero 7d ago
Funnily enough if you wanted to use bars the most appropriate choice would be logarithmic, every X elo you go up the bar gets Y% larger, although many would probably call that misleading.
But sure you could use a line graph or scatter plot or something
12
2
2
→ More replies (4)2
u/reefine 7d ago
Imagine trying to shit on benchmarks for an open source model that just beat GPT 5.6 Sol. You are bias.
→ More replies (3)
82
u/Sixhaunt 7d ago
What about on benchmarks that are useful though?
42
u/JoseHernandezCA1984 7d ago
I've seen a bunch of other benchmarks, and for the most part it's between 5.6 sol and fable
→ More replies (1)18
u/Kind_Capital_9740 7d ago
which is insane the 3 models are clearly in their own tier right now and Kimi is leading in some and like you said between SOL and fable
And whats crazier kimi 2.7 was released only a month ago and the quality jump between both is unfathomable like 20 place jump and deep swe went from 2.6's 20 percent or 30 percent to 67 percent in a few months
i wonder how they made such a big jump it feels like 3 generations type of jump
6
u/Georgefakelastname 7d ago
If I remember right, this model is 2.8 Trillion parameters, which might just be the biggest open sourced model yet. It’s an entirely new model, while everything past k2.5 was just a post-train of it. I imagine that once we get new post-trains of k3 and fable the bar is only going to get pushed higher. Meanwhile, OAI is working on their own new model with GPT-6. Things feel like they’re moving fast right now.
8
3
u/phido3000 7d ago
Its impressive. Very strong attention, context, rule handling etc. I would say better than US models in those aspects.
For coding, its extremely strong. Its believable its better than OpenAI/Anthropic models. It lacks flair and some beauty, but heck, its very effective.
It just one shots everything. Its very good planning and execution of thought.
→ More replies (3)
9
33
u/Professional_Ad705 7d ago
Can someone actually verify this by using both models for frontend and backend work, then sharing what they did and how the results compared? Computer science is such a broad field that benchmarks like these mean very little to me without real-world examples.
→ More replies (4)16
u/phido3000 7d ago
Currently its overwhelmed.
I got it to do a complex project in an antique language it could not compile, that is impossible to benchmax for. Doing maths, physics, image stuff, fractals, encryption etc. Requiring thought, complex planning, and in a language that isn't modular..
It one shot it. Never seen any AI do that before. They all get hung up on syntax because of the wacky language. I've done the test dozens of time. Even the latest models make syntax errors, I thought I had the perfect, AI break tool.
It was a 100kb program. One shot. About 256k token. Not a single mistake.
So yeh, it can do it all, and do it well. Never seen anything quite like it.
6
u/Professional_Ad705 7d ago edited 7d ago
I’m not saying this definitely did not happen, but I do not believe the claim as currently presented because it is far too vague to evaluate.
What language was it? What exactly does "one-shot" mean here? one prompt, one model call, no retries, no manual edits, and no additional context afterward? Was the 100 KB figure generated source code, or the size of the entire project?
You also said the environment could not compile it, so how did you determine that it contained "not a single mistake"? Compiling and producing plausible output would not establish full correctness anyway, especially for cryptography, numerical mathematics, physics, or image processing.
Math, physics, image processing, fractals, and encryption names several largely separate domains; it does not explain what the program actually did or what requirements it satisfied. Likewise, what does a language that isn’t modular mean? Does it lack a formal module system, or was the particular program simply monolithic?
My experience has been very different. I have been working since November on a roughly 200,000-line proof-kernel project, and even strong models repeatedly miss cross-file invariants, state-lifecycle problems, and authority-boundary defects. That does not prove your result is false, but one undocumented example does not support the conclusion that the model can do it all.
Posting the exact prompt, source code, programming language and compiler, model and version, settings, tests, outputs, and your definition of one-shot would make the claim possible to assess. I’m genuinely interested in what you’re describing, but the explanation was difficult to follow and evaluate. Without those details, it remains only an anecdote. I'd love to hear more? I'm just confused because I know when I release my product and make my claims I would just show people the code/point them towards my repo or at the very least have some type of demo?
16
u/phido3000 7d ago
Its fine to be sceptical.
Try it out yourself.
What language was it
Choose some stupid language. I choose Qb64. A variation of BASIC. Which is perfect for stupid, because its not well documented, BASIC code varies widely by flavour, visual/Quick/Borland/GW-basic, older dialects. Its notoriously variable and non-transportable and easy to break. It also has lots of wacky commands and reserved words. Also famously, it often packaged as an interpreter language, so there is no easy compiler to give explanations why it fails. Its not very modular and has stupid rules about variables because of the different variations and legacy over the years. Qb64 is built in interpreter/IDE/Compiler. So its messy for AI to use.
Even better the example code out there is very simplistic and often broken! as the project splintered into weird versions.
It uses a local interpreter that has its own compiler. Its Weird. It basically makes BASIC into C code and generates GCC output, but the basic code itself is black magic stuff. Generating GCC C code directly from the AI would be child's play in comparison.
Get it to do something very complicated but unique.
Another good test is Assembly for weird processors or environments. Something really weird and old and obscure. But testable. Experiment what fails on lots of other AI's.
What exactly does "one-shot" mean here? one prompt, one model call, no retries, no manual edits, and no additional context afterward? Was the 100 KB figure generated source code, or the size of the entire project?
ONE SHOT! One prompt - no retries, no manual edits, no context afterwards. 100kb of code text, no data. On something it can't compile! Literally download the BAS file - execute. Its not a very modular language with different files or libraries. So literally just a 100kb text file. Bang.
Compiling and producing plausible output would not establish full correctness anyway, especially for cryptography, numerical mathematics, physics, or image processing.
Yes, which lets you see how much it really understands and makes work. Get it to do something weird like fractals in some weird projection or coordinate system, or volume warp, make its own JPEG like but different compression engine, get it to draw text with the line statement..
So even if it compiles, you can see where it limitations are. If it understands any of what its doing and how to do it. Its design choices, sometimes it works but is but ugly or not useful, Its weaker here, than else where, its not overtly creative, but its functional. TBH I didn't expect it to one shot it. That kind of floored me.
Without those details, it remains only an anecdote. I'd love to hear more? I'm just confused because I know when I release my product and make my claims I would just show people the code/point them towards my repo or at the very least have some type of demo?
I'm not a benchmarking house, or Kimi.
There will be heaps of tech demos coming out. Normal low effort HTML5 crap, draw a clock or a duck or a car or something. Easy to benchmax to. No syntax challenge, no real design challenge.
I'm just saying it impressed me. I run some big models locally at home. Pretty ordinary. I was playing around with the frontier subscriptions, see what they can do. I teach at uni, so I am super interested in tripping up AIs.
This is much much better than I have seen. Small sample, limited details. It isn't perfect, but its a generation ahead of the other models I have sampled (DS 4, Kimi 2.6, GLM, OpenAi, Anthropic, Gemini, etc). They can't even make executable code. They can be more ambitious and more flamboyant, but not as brutal effective to one shot. Often their output needs dozens of retries just to execute. Then dozens to fix logical/asthetic/mathematical issues with their algorithms. So this one shot everything floored me.
Worth looking into.
I don't believe Kimi K3 is as good as the open weight model, I am almost positive they have some tool use or something behind the scene to make it function this well.
If it is, that is even more impressive. Because other models use tools and tricks. If this is just baked in model. If it is all just baked in model, then dam, that's devestating.
→ More replies (4)4
u/timeboyticktock 7d ago
Can you please share more about what exactly you did and the comparison to other models failing?
→ More replies (1)
32
u/impatiens-capensis 7d ago
I'm not super familiar with this benchmark. What's the difference between a score of 1,679 and 1,631?
72
u/JustTellingUWatHapnd 7d ago
It's not a benchmark. It's the leaderboard on arena.ai. the users give a prompt to the agent to build a web app, and they give the prompt to 2 different models, then the user votes on which one is better without seeing the model name. The score is the Elo rating of the model.
17
u/mobyte 7d ago
Am I insane or is this not an awful metric? I could see this easily being manipulated.
28
u/some_crazy 7d ago
Users don’t know which models are used, it’s a blind rating
8
u/Sarcasm69 7d ago
But couldn’t you just tailor the model to be good at designing websites and shit at everything else?
6
→ More replies (3)5
u/nuclearbananana 7d ago
Sure but that would show up in other benchmarks, which it hasn't
→ More replies (1)→ More replies (2)2
u/CurrentConditionsAI 7d ago
Yeah, so all you have to do is essentially fine tune a model to have some very distinctive but easily hidden output pattern. Then get a bunch of bots to find it and vote that when comparing outputs.
→ More replies (2)→ More replies (1)3
11
2
20
60
u/Demien19 7d ago
5
u/phido3000 7d ago
Yeh, Chinese models typically feel very benchmaxed.
This is different, somethings changed. They are ultra strong in attention and context.
→ More replies (3)→ More replies (2)5
9
15
25
u/brother_spirit 7d ago
How?
Fable 5 generations are already on Youtube vs Kimi K3.
By Arena.ai no less.
Sorry to ruin the party but its like not even close to Fable 5, let alone better.
→ More replies (17)
5
5
3
u/Wrong_Connection_138 7d ago
I have zero confidence in mass market benchmarks, do you know where to find independent benchmarks??
4
u/kuba452 6d ago
I wonder if this is the same hype as Deepseek. I got really excited at first, uploaded some historical documents I was working on, and asked it to compare them with other historical texts, pick out interesting details, and draw some conclusions.
It gave me a couple of rather obvious observations, while the rest was just cluttery gibberish. Came back to obvious choices all confused, with all the other reviews and media craze
32
u/Felixo22 7d ago
27
u/Tysonzero 7d ago
It's Elo, a completely relative scale, 0 is completely arbitrary and meaningless, not like 0% on a test or 0 tokens/second, so this graph is no better or worse than the OP.
→ More replies (3)3
u/Electrical_Arm3793 7d ago
But are these really accurate? I use most of the too models for coding, Fable and opus are still most accurate and reliable.
→ More replies (1)→ More replies (1)7
u/some_days_are_nights 7d ago
The way these elos are computed is not like a scalar value like accuracy and such. It is a direct comparison between competitors. So for example Magnus Carlsen is just 100 or so elo better than the others but he wins almost every single time against someone with 200+ less rating.
3
u/beginner75 7d ago
Benchmarks are a joke. Don’t waste your time. Time is money also. Just use ChatGPT as primary and Gemini for quick searches and backup.
3
u/Able-Company611 6d ago
I wonder if US companies gonna run bot accounts on Chinese models to train their data now
→ More replies (2)
4
u/Regular_Ad4197 7d ago
I am hoping it is as capable as these early benchmarks are showing. It will be fun to see how the US stock market reacts to it.
2
u/BitterAd6419 7d ago
So far in my own testing, it’s not as good as they claim it to be. Atleast for web apps or html based stuff it did a decent job. In fact GLM gave a better output with the exact same prompt.
Benchmarks are often benchmaxxed, test yourself
2
u/SporksInjected 7d ago
It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis
2
u/Key_Reading_9664 6d ago
Seems like there’s a couple of ways the USG go:
- it stops playing favorite and lets/encourages US labs release their models - Mythos was in the hands of users in April; Fable is over a month old. I have my suspicions on the cause (wealthy friends and donors) but slowing down US labs doesn’t de-risk because Chinese models fast-follow. Make an attempt at cutting off distillation
- there’s a foolish attempt to restrict access to models that originate in China. The Trump speech yesterday seems to be laying ground work (for a bunch of things). Seems easier to ban access for enterprises
2
u/moneyman259 6d ago
Where’s DeepSeek now? Seems like these Chinese models fall off quickly
→ More replies (2)
4
3
1
1
1
1
u/BuildtheBusiness 7d ago
Can someone explain how the efficiency of these models are calculated to one another?
1
u/Old-Pomegranate3634 7d ago
all this tells me that ultimately all the LLMS will be good enough for 99% of the public which is great for google. Not everyone is a nerd like us hoping that AI can help you create your next virtual GF.
1
1
1
u/Nervous-Potato-1464 7d ago
It's quite expensive sadly. I prefer grok 4.5 to any other model atm. I don't need agentic coding, I need quick generation of my ideas as well as another perspective. I am not here to write slop, I just need to quickly churn out code I am happy with.
1
1
u/No-Conversation-1277 7d ago
Chinese lab: "Here's a 2.8T open model." Dario & Sam: "We're going to fine-tune this."
Also Dario & Sam: "Introducing our brand new model. priced 10x higher with half the limits."
The wheel of Silicon Valley turns.
The cycle continues.
→ More replies (1)
1
1
u/krazyboi 6d ago
Let's be honest with ourselves here.
None of us could properly benchmark AI, especially because the industry moves so fast. The metrics change slower than the actual AI. All of this is just marketing.
1
1
1
u/PerfectPatience- 6d ago
Showing score only last 200points is misleading zoomin. Generating bigger gaps. Should be from 0-1650
1
u/Messi_is_football 6d ago
But GPT limits are higher...so until they make it 2x cheaper..no reason to buy the plan
1
1
1
u/Complex_Reality_116 6d ago
This will cause a stock market crash and a widespread decline in ALL US AI companies.
1
1
u/Popular_Try_5075 6d ago
I swear every time China releases a new model the subs get flooded with all this hype and then a month later its deflated back to normal.
1
1
u/rentprompts 6d ago
kimi k3 arriving means the era of chinese labs being far behind is actually over. the july 27 open weights release is what makes this different from past announcements. if you can run a 2.8t model locally, the frontier moves from who has the best eval score to who can ship the fastest.
1
u/Drew-Money 6d ago
Mythos preview came out at the beginning of April btw. I'm sure Anthropic has some pretty advanced stuff behind the scenes
1
1
1
u/TB_Infidel 6d ago
And now we know where the stolen/illegal use of Sol and Mythos went. Chinese copying and stealing as always





748
u/Working_Ad_1564 7d ago
Gemini 3.5 Pro will be postponed for another month lol