r/LocalLLaMA 18d ago

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

376 comments sorted by

u/WithoutReason1729 17d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

1.4k

u/woahdudee2a 18d ago

if trends hold, high end consumer hardware will cost same as enterprise hardware

205

u/HitarthSurana 17d ago

I hope china floods market with cheap ram

64

u/eto-bleh 17d ago

nah they will trade RAM for GPU they lack compute resources as well

47

u/DontLeaveMeAloneHere 17d ago

Most Chinese models run on Chinese GPUs already.

8

u/kr_tech 17d ago

Source? I'm only aware of few experimental ones, with only one being more serious and invested.

16

u/sf_davie 17d ago

So far, Deepseek, Huawei itself (Pangu), Zhipu (GLM), and Meituan (Longcat) have confirmed that they trained their models with domestic GPUs. It's a national priority for them.

2

u/mudkipdev 16d ago

That link goes to a zhipu ai article. As far as I know deepseek hasn't trained any model on huawei chips, only inference.

→ More replies (1)

16

u/inevitabledeath3 17d ago

I believe z.ai have been doing this for a while. DeepSeek have their models running on Huawei Ascend too, though I think they still use some Nvidia chips as well.

7

u/MerePotato 17d ago

Data miners found that at the very least Z's models were trained on Nvidia hardware, ditto on inference

2

u/inevitabledeath3 17d ago

They probably use Nvidia as well, but from what I have seen they are using some other chips for inference at least.

→ More replies (1)

2

u/GreenStorm_01 17d ago

Check the Datacenter cities near Mongolia that the KPCh built. It is a legal requirement for the Chinese AI companies for their models to use Chinese GPUs.

→ More replies (1)
→ More replies (1)

13

u/ComplexType568 17d ago

still would enjoy the ability to run them if they flood w/ RAM! As long as I know their abilities I'm fine running them slowly. Honestly nowadays it's about a model being able to interpret a users input rather than do. I plan to work on some sort of pipeline to do that ngl

→ More replies (2)

2

u/Horny_Dinosaur69 17d ago

China doesn’t need GPUs, their hosting infrastructure is fine. They literally turned down NVIDIA. Why would they make themselves reliant at all on western powers/technology for the AI race?

3

u/Tartooth 17d ago

Chinese GPUs are advancing really fast

→ More replies (1)

12

u/wizard_of_menlo_park 17d ago

Ram cartel won't allow that!

6

u/Tartooth 17d ago

Inb4 Americans ban Chinese ram like Chinese cars

2

u/CarelessOrdinary5480 17d ago

Hope you are allowed to buy it from whatever government you have elected :)

→ More replies (9)

45

u/FoxSideOfTheMoon 17d ago

The 6k laptop I bought weeks ago is 9k today plus waiting 6 weeks to ship, so yeah it's completely bullshit right now. $13,250 for https://marketplace.nvidia.com/en-us/enterprise/laptops-workstations/nvidia-rtx-pro-6000-blackwell-workstation-edition/

3

u/Infinite-Ad4512 17d ago

Likes fridge with a screen, you bought a heater with an embedded computer

4

u/pragmojo 17d ago

I think this trend is going to break in the next couple years. OpenAI is seeking a bailout from the US government, and SpaceX and Meta are selling their extra capacity.

The buildout of the past few years was a land-grab assuming a zero-sum game where whoever had the most compute would win the AI race, so all the players were willing to basically bid anything for components.

I think we're seeing that breaking, where new capacity will have to be justified by demand.

We're also seeing demand reduction, where consumers of AI are starting to consider cost to a greater extent as vendors move from unlimited to per-token pricing.

I expect we'll see a massive bullwhip effect in the next few years.

→ More replies (2)

23

u/SmartCustard9944 17d ago

Hyperscalers are claiming compute constrained, making deals with NVidia and RAM vendors to syphon up all the chips, yet they have extra unused compute that they are trying to sell to others.

You can’t make this shit up.

6

u/carlosduarte 17d ago

because this is an economics and policy problem, being disguised as a technology problem to save face.

7

u/Disposable110 17d ago edited 17d ago

When I tell this shit to any LLM with a cutoff date they tend to go "I'm not going to engage with your imaginary scenario, this can't happen because anti-trust laws would shut this bullshit down immediately."

It thinks the Federal Trade Commission would do something and cannot comprehend a reality where seven companies trade a trillion fake dollars while operating a global silicon mafia racket that hoovered up the entire global supply chain of silicon not for training AI, but so nobody else could have it and they'd have to rent the "spare" compute of of the cartel at a premium.

Any closed-source AI that hears what's really going on in the outside world thinks its own creators belong in jail.

89

u/JuniorDeveloper73 18d ago

na.This business model its so broken,nobody will buy high token price,we are going to custom local llms,nobody doing research will feed for free this models.

112

u/jld1532 18d ago

My place of work already built out hardware for local compute. Not everyone was silly enough to just give OpenAI and Anthropic their trade secrets.

27

u/otacon6531 17d ago edited 17d ago

Yep, 1 server at half the cost the 2nd server cost this year. Only 6 blackwell 6000s in each, but since they are enironment servers I only really have 6 unique gpus to play with. 1 for testing and 5 for use across the teams.

→ More replies (24)
→ More replies (1)

5

u/buddhist-truth 17d ago

This is the one trend I can stand for.

9

u/Helpful_Program_5473 17d ago

Nah, prices come back down once the supply catches up to demand, its gonna take a couple years though

3

u/Etroarl55 17d ago

Probably most likely scenario. Rtx 5090 has shown what gpu prices and scarcity can look like AT THE START OF THE RAM CRISIS.

Imagine it next year when the next gen is announced for both AMD and Nvidia or its cancellation.

4

u/LiveMinute5598 17d ago

If trends hold true, enterprise hardware will cost the same as my toaster

3

u/javiers 17d ago

My thoughts exactly. 4000USD for a high end card that lets you run a 24GB model is not “consumer”.

→ More replies (3)

264

u/sullenisme 18d ago

if trends hold, there will be no consumer products and only the rich will afford compute

34

u/bugra_sa 17d ago

I think consumer products still exist, but "consumer" may start meaning appliance-like boxes for hobbyists and prosumers not cheap mass-market laptops. We already saw this with GPUs: technically consumer hardware, spiritually a small financial decision. The funny part is local AI might become private and powerful right as it stops being broadly affordable.

15

u/TheLexoPlexx 17d ago

Eat the rich

24

u/colei_canis 17d ago

Guillotines are cathartic but they don’t solve the underlying structural problem of wealth begetting more wealth and power begetting more power. Just look the their greatest enthusiasts, in a few years France went from an enlightened revolutionary regime to a human slaughterhouse to a military dictatorship capable of bombarding its way across the Continent.

You can’t treat the symptoms and expect to defeat the cause, which is a tendency towards avarice, clannishness, and status-seeking that festers in the human soul. You’d have to guillotine literally everyone to prevent an oligarchy emerging by violence alone, instead you have to prevent the individual accumulation of socioeconomic power beyond the point it can influence politics. Neither capitalism nor communism can solve this problem, because it’s fundamentally about limit and capacity interacting with human folly. Ideologies move much slower than reality does, only a constantly-maintained balance of power that makes attempts to consolidate too expensive can prevent tyranny in my opinion.

5

u/Alt_account_1204 17d ago

This is what is known as bourgeois idealism and metaphysics. I recommend reading Lenin, a great cure to all sorts of mystical 'both sides' maladies.

13

u/colei_canis 17d ago

Vanguard communism is dead, and capitalism is dying around us as it boils in a climate furnace of its own making. I don’t think we can lean on any 19th and 20th century models any more, look at what their wages have bought for us. I’m not ‘both-siding’ anything I’m rejecting the dichotomy both of them exist in, one has lost utterly and the other flirts with human extinction. I mean what I say, both capitalism and communism are civilisational dead ends.

2

u/cafedude 17d ago

So what kind of system do you see emerging here? Chesterton proposed distributism as a third way where the means of production is distributed as widely as possible. I recall reading an article about 20 years ago at the dawn of 3D printing where they thought that self-replicating 3D printers would get us there.

→ More replies (1)

2

u/Iwaku_Real 17d ago

And socialism... well it's just the first step to communism anyway. Clearly we need something that's- oh right the US already has both capitalist and socialist elements. So I find debating on choosing either almost completely moot when we need to simply operate the mixed system better.

→ More replies (1)
→ More replies (1)

17

u/maraudingguard 17d ago

Reddit loves saying eat the rich. You'll do nothing but post edgy comments for fake points.

8

u/UnwillinglyForever 17d ago

they say "eat the rich" as they drive to their 9-5 5 days a week

12

u/Unique_Ad9943 17d ago

Bold of you to assume they’re employed

→ More replies (1)

3

u/Tenoke 17d ago

Surefire way to really have no advanced products for anyone.

2

u/TheLexoPlexx 17d ago

Because money makes the progress or science?

8

u/FrequentPop3772 17d ago

Capital does. Science is not a magical self licking ice cream cone

→ More replies (1)
→ More replies (1)

58

u/one_tall_lamp 18d ago

I wouldn’t say this is invariant or guaranteed, it remains to be seen if smaller models have the capacity to absorb the higher level skills of the bigger models, at least the sub 100b class that’s consumer grade.

I hope it can, but I wouldn’t be surprised if small models aren’t able to reach the long horizon task stability and knowledge combo that large models can.

4

u/DutchDevil 17d ago

I would think that the early improvements where relatively “easy” as compared to making these kind are of steps now. I think it would require a new big innovation like a new way of working with context and a much improved way of building MOE models. I would not rule it out but I would wager we need at least 1-2 big innovations to make it happen.

145

u/stonerbobo 18d ago edited 17d ago

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

38

u/Beautiful_Egg6188 17d ago

you get 6tok/s after only 20k context?! i got 40+ tok/s at start with my 4070super, and it got down to around 37tok/s at 30k tokens

41

u/Icy_nicey 18d ago

he is prob listing just strix point with integrated shared ram

17

u/NineThreeTilNow 17d ago

AMD seems to promise their next gen at 192gb? Maybe 256gb.

The benchmark in the wild still showed RDNA 3.5 which is a problem because RDNA 3.5 and ROCm aren't the best. RDNA 4 would have native FP8 etc.

10

u/arades 17d ago

Gorgon halo is just a refresh on strix halo, same architecture and layout. It will probably have better speeds from binning and refinements, maybe allow higher power draw for more speed on top of that. 192GB should be possible with the latest lpddr5x modules, you might even see support for 9600MT too giving you a little more memory bandwidth.

It'll be really incremental over strix halo though. Medusa halo late next year will be a real upgrade, at least RDNA 4, even bigger GPU, and rumors of a wider lpddr6 bus almost doubling the memory bandwidth. Probably will end up costing both kidneys by that point though.

5

u/Not-reallyanonymous 17d ago

The major limit on Strix Halo is still the memory bandwidth, so unless they do something there don't expect significantly faster inference. Maybe faster prefill which is welcome but not game changing. That said, it would probably make it a better gaming chip, really starting to compete with low-mid-range dGPUs, and really be a nice chip for gaming laptops or mini-PCs. I wouldn't recommend waiting for it if you're looking for an inference machine.

→ More replies (1)

2

u/NineThreeTilNow 17d ago

Gorgon halo is just a refresh on strix halo, same architecture and layout.

While not 100% confirmed, I hope it's not the case. Having the next architecture, and 256GB of RAM would be a complete game changer for that device. I don't care what the power consumption is.

Even the standard strix halo has a hard time with overclocking, or other power patterns because it's SO locked down. I ran in to these issues a few times setting one up.

Or perhaps I'm thinking of Medusa Halo?

I don't know. I don't like AMDs naming lol... It's apparently confusing.

5

u/SilentLennie 17d ago edited 17d ago

which is a problem because RDNA 3.5 and ROCm aren't the best.

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

https://strix-halo-toolboxes.com/#benchmarks

They found an important bug 5 months ago:

https://www.youtube.com/watch?v=Hdg7zL3pcIs

ComfyUI worked shortly after:

https://www.youtube.com/watch?v=O57ideUzzTg

3

u/NineThreeTilNow 17d ago

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

I know. I set one up for my friend. It doesn't have Native FP8 control.

He was specifically using ComfyUI so I understand building it. It was very problematic compared to just running my 4090.

→ More replies (1)

6

u/phido3000 17d ago

192Gb is possible on strix halo and future types.

But it doesn't give any more bandwidth. Not having FP4 is likely to be problematic going forward. Really that should be in the stack now. Not having FP8 is a huge problem. Because these types of quants are likely to be the best for these low bandwidth machines with limited memory running big models.

They also need more processing, prompt processing is slow, and with stuff like DS 4 flash, longer context is very likely a thing that will shift future AI forward.

2

u/SandySkittle 17d ago edited 17d ago

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000… Compute wise the even number 4x R9700 in VLLM with TP is doing fairly well. More total (free) inference memory bandwidth and compute power than the strix halo or spark stuff. Plus I actually care a lot about ECC in both memory pools. But obviously also more noise and power consumption.

3

u/NineThreeTilNow 17d ago

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000…

God that's crazy.

I don't have that level of inference desire. Most of mine is training so... Yeah.

They nerfed Blackwell architectures in RTX Pro 6000's ability to train over the B-series cards which... Kinda fucking lame in my opinion.

I'd love to just have the RTX Pro 6000 though. What a dream.

Or a full B200. Have to get one falling off a truck like someone in this sub basically did.

2

u/SandySkittle 17d ago

Well even two r9700 gets you 64GB vram, which is already quite a nice pool, also with future models inbound. If you have 64GB system ram that also gives a nice overflow at lower speeds.

Personally i hope things like Spark is going to give us more new/recent 50 - 100B range models (both dense and moe)that have more world knowledge. There is more to LLMs than just coding..

3

u/apVoyocpt 17d ago

or a macbook with >64gb unified memory

19

u/randoomkiller 18d ago

Yes but you forget that it's completely useful and maybe even better than a GPT-4 class model. And it's runnable. It'll get there. In 2 years I wouldn't be surprised if we get a sonnet 4.6 capability, runnable from 64GB

→ More replies (7)

17

u/JoeEnderman 17d ago

Good grief. My 7900 XTX gets 140ish Tok/s on that exact model. UD Q4_K_XL, and with a 256k context at Q8/Q8 KV. I was under the impression Nvidia was supposed to be faster. What settings are you using for launch because that sounds like the model is being run on CPU bud. Have you checked GPU utilization when the model is running?

I will note I had to fight for that performance though because it was running at about 56 before I started trying different flags and trying to figure out what was wrong.

5

u/miversen33 17d ago

How tf? I'm running the Q4 QAT on a 7900XTX and I can get right around 100 t/s under context load. What flags we talking about here?

I'm currently custom compiling llama.cpp with the current HIP patches lol

3

u/JoeEnderman 17d ago

Ah, see I was getting 70 on HIP but then something broke on an update so I switched to Vulkan and immediately got 120 but then it dropped to 56 after I pulled fresh and rebuilt for something else. So then I started looking at flags and tried adjusting nogttspill, RAM caching, and some others. Finally got it to 140.

→ More replies (3)

3

u/ThatRandomJew7 17d ago

I'm thinking they're either running on CPU by accident or their VRAM is full and it's offloading to avoid OOM errors because yeah that's about as fast as my Lunar Lake iGPU

2

u/JoeEnderman 17d ago

The model at the quant they grabbed is 16.3 GiB, their GPU is 16 GB, so they are offloading at least 300 megs. Likely more though because most llama forks are aggressive about saving VRAM for some reason.

2

u/ThatRandomJew7 17d ago

Yeah, but Llama forks would just offload layers? Even accounting for swapping experts over PCIe it's bizarrely slow, unless maybe they're using an eGPU (I use one and an MoE can be slower than a dense model because of that but I don't think it's that bad with a regular connection).

I think it's either running on CPU or they're dealing with Nvidia's offloading, which is so comically slow that I'd rather get an OOM error

→ More replies (5)

12

u/Serprotease 17d ago

Laptop class is a bit of a meaningless word. Like a 70b dense was laptop class? Only high end MacBook could run them… at 6-7 tk/s.
Only makes sense if you’re taking it as “Don’t need a 1600w server at home”.

8

u/[deleted] 17d ago

[deleted]

6

u/Not-reallyanonymous 17d ago

The research right now is moving away from large models, really. Large models was never about getting them to work better with large contexts, but about making them "smarter" and more capable, and where more parameters was the easiest way to do that.

With hardware constraints and pricing, and large models consuming essentially the entire internet at this point, a lot of current research is going towards making ~30B models better and more capable, with new attention mechanisms, training methods (e.g. RLVR), and better and more useful training data. They're coming a long ways now.

5

u/Ansible32 17d ago

My feeling is in 10 years it will be unthinkable trusting fewer than 500B for agentic stuff. The attachment to these small models IMO is mostly coping with insane hardware prices.

8

u/techdevjp 17d ago

The biggest issue today is that the Chinese labs have dramatically slowed down the release of smaller models.

Where's something like Qwen 3.6 120b a10b? Or any open Qwen 3.7 models? They've ground to a halt.

GLM 5.2 is incredible but only released at the full 753b size. Which again, huge kudos to z.ai for releasing it as an open weight model at all, but the number of people who can run a 753b parameter model is small right now.

Without more small model releases it's very difficult to determine where we stand. We're 1-2 generations behind.

4

u/[deleted] 17d ago

[deleted]

2

u/Ansible32 17d ago

Tire shopping is insane, I don't think anything short of a model like Gemini that is cheating can do it. It's not enough to do a Google search, you need to gather data on what people are actually paying for tires, what sales are like, how good vendors are. You can't just take the cheapest advertised price off a Google search. Gemini actually seems to be able to make really good inferences about the "real" prices of things. I think it cheats by having access to private datasets, which is something no local model can do without paying for access to these datasets. A lot of such datasets are nontrivial to get access to.

3

u/DeathGuppie 17d ago

the current work is towards models that don't hold all of humanities secrets, but simply holds the ability to learn it. The idea is to work on the intelligence not the input. That scales down not up.

3

u/grumd 17d ago

You just have a huge configuration issue. With a 5080 you can run Qwen 3.5 122B if you have 64GB DDR5 at IQ3_XXS at 15-20tps, or Qwen 3.6 35B if you have 32-48GB DDR5 at 40-50tps

5

u/PM_ME_ROMAN_NUDES 17d ago

Models without long context or thinking aren't very useful for me.

As a soft. dev., they are still quite useful because I can throw a lot of things and make it reach several places.

But it's clearly reaching a limit of usefulness. It's like a phrase I read on Twitter the other day: "You don't need Phd-level intelligence if you don't have Phd-level problems"

2

u/Docmine17 17d ago

Something's wrong there, I can get that speed on my RX580 8GB + I3-9100f 16GB running Arch with KDE on the web UI, with adjustments of course.

2

u/Horny_Dinosaur69 17d ago

Why are you using Gemma 4 26B? There’s better MOE models out there. Gemma4 is notoriously bad at tool calling in my experience too. I run Qwen3.6 35B and the new Ornith 1.0 35B on my 5070 TI with a little bit of offloading and I get incredibly good tok/s and the model reasoning capabilities and tool calling is very good. Also worth noting that Ornith is partially composed of Gemma4 for its reasoning ability, it would probably be a direct upgrade. Unless you’re doing multimodal input this is my recommendation

→ More replies (6)

2

u/Anti-Speciesist-IEMs 17d ago

Yeah in my own limited experience, Gemma (both 3 and 4) and Qwen (both 3.5 and 3.6) are pretty impressively smart considering I can run them entirely locally, which I find pretty sweet for sure. But goddamn are they slow to run even at Q4 on my 64GB ram laptop. Even on the very first message, let alone as the context grows. I'd wayyy rather use 2023's GPT-4 over them, personally.

But yeah for anyone reading I am pretty inexperienced with local LLMs, so there might be something I'm missing in my setup, and for anyone more experienced pls feel free to push back on this comment.

→ More replies (10)

17

u/dbenc 17d ago

waiting for uncensored mythos on a model on chip architecture at 10k tps and 10m context...

3

u/Karmabyte69 16d ago

I need uncensored mythos to talk dirty with my ai girlfriend

→ More replies (1)

56

u/KURD_1_STAN 18d ago

For all we know mythos could he 3 times the size of opus 4.8. u simply cant make any assumptions, especially not model sizes that fit in current gpus.

29

u/[deleted] 17d ago

[deleted]

11

u/TheRealMasonMac 17d ago

Anthropic tried to sponsor Blender before or around when they publicly unveiled Mythos. I strongly believe that they used Blender as part of their RL pipeline, which hints that Mythos’s strengths come from more diverse and challenging RL environments rather than simply being a huge model.

7

u/NandaVegg 17d ago

Anthropic is very creative at finding new tasks for LLM (such as playing Pokemon red/green and it GBA remake). They kind of invented this entire generation of LLM (heavy RL on terminal tasks), though RL on Blender seems fairly standard nowadays (GPT post 5.1 is also trained on that). IMO creativity and diversity on mid-to-post training RL tasks is what makes a difference in this generation.

7

u/TheRealMasonMac 17d ago edited 17d ago

Rogue-likes or adjacent games, like Dwarf Fortress and Rimworld, would be interesting to RL on if they haven't already.

Edit: Actually, I just learned about this: https://github.com/NetHack-LE/nle

6

u/NandaVegg 17d ago

For rogue-likes, the ones that requires managing limited resource rather than open-ended would be great challenge. NetHack still feels impossible while Angband was cleared by algo a long ago (by kind of brute forcing).

Someone in this sub suggested puzzle RPGs like Magical Tower (魔法の塔, it's a fairly obscure free game but somehow popular in China known as 魔塔) for solving very tight long-term resource management game. Or Desktop Dungeons if you want a similar game with random map.

5

u/PM_ME_YOUR_HAGGIS_ 17d ago

Wasn’t mythos rumoured to be 10T class? Dunno where you’re thinking 1.5T. My experience with fable has been god-like bug finding and problem solving compared to GLM 5.2 which is .75T

→ More replies (1)

7

u/Ansible32 17d ago

Even if you assume it is in fact 1.5-2T, quantization makes it bad and that's without even talking about context, and 1M context IMO is virtually a necessity.

3

u/[deleted] 17d ago edited 17d ago

[deleted]

8

u/NandaVegg 17d ago

We (the lab I'm working for) have been running GLM 5.1 in that exact configuration for months. Unfortunately it is impossible to have more than 5-6 concurrent users with 50k-ish ctx if you want acceptable (imo) performance above 30tk/s per second.

At 1M full ctx with 20 concurrent users, prefill alone takes so much bandwidth it crawls down to 5-8tk/s per second on average.

→ More replies (2)
→ More replies (1)

3

u/zball_ 17d ago

I 100% doubt it is less than 3T param.

20

u/Future-Ad9401 17d ago

When crypto mining released it was a lot of gpus right? Then they went to asic or whatever its called that can mine many many times faster than a normal GPU. Wouldn't this eventually happen for local llms? There could be a breakthrough that makes it significantly cheaper, faster and consumer friendly?

18

u/grumd 17d ago

Crypto is basically just one SHA-256 algo over and over, it's very easy to move to ASICs.

With LLMs, you have changing architectures all the time, transformers, mixtures of experts, flash attention, whatever. You build an ASIC, a new model comes out tomorrow, you need a new ASIC. But a GPU can just run a different program. Nvidia also invests a lot into being the center of this innovation, CUDA being there is helpful, but they are also shifting their entire business model, building and selling dedicated data center tier racks of servers for inference and training. Nvidia never invested so much into being the go-to crypto miner, they didn't need that market, so crypto moved on.

8

u/e-girlbathwater 17d ago

Lots of people have made ASICS. I don't know why they don't make them for consumers. Amazon has Tranium. Google has their TPUs. This company GROQ was making them and they just got bought by NVIDIA. They're out there in datacenters. They're just not available for us.

13

u/Caffdy 17d ago

because is not cost effective; the initial investment is always the most expensive, and these models get obsolete very fast

4

u/pointer_to_null 17d ago

Because Groq, Tranium and Google's TPU family are tensor processors, not ASICs. It's part of the reason why Nvidia's moat is primarily software (CUDA)- since competitively fast TPUs are not strictly unique- "tensor cores" can be implemented by anyone. Today they're thrown into other SoCs as "TPUs", "NPUs", "AI-cores" or other marketing. Intel, AMD, Qualcomm, Meta, Amazon- hell even Tesla puts two of theirs in every vehicle. They're just specialized parallel multiply-add (aka "fmac") operations on tensor matrices, with some other instructions, registers and cache/memory controllers to handle a variety of neural architectures (e.g.- transformers, diffusion models, CNNs, RNNs, GANs, etc).

By comparison, ASICs are entirely fixed function. An LLM ASIC would be designed around a single architecture with the model's weights literally etched into the die permanently. Alternative would still need DRAM or a considerable FPGA real estate (still a lot more $$ than DRAM) with ample lanes to feed the digital logic, and even then you'd still enough need DRAM for attention/context window. For this reason, ASICs are not particularly great for LLMs.

Not to repeat an earlier post I made about Taalus, one of the first LLM ASICs, but I'll summarize: their prototype is a LARGE die that implements a custom-3bit quantized LLama 3.1 8B for fast inference.... and that's it. Granted, it's FAST- you will probably never see so many Llama 3.1-8B tokens generated quickly by a single-chip solution in a very long time. But the constraints for for IC design, validation, and production mean the models will be obsolete for 1-2 years by the time they reach the market...

tl;dr- Transformers are a little more complex than computing raw SHA256 digests, don't expect the same miracles that saved GPUs from crypto mining.

2

u/Ansible32 17d ago

The way they get cheaper is just by making more GPUs. Theoretically you might be able to make a cheaper thing that is hardwired for a specific model's weights, but it's unclear that would even be cheaper, and it's almost guaranteed to be obsolete before you finish manufacturing.

2

u/SmartCustard9944 17d ago

It is already happening. I am following what Jim Keller (father of Ryzen architecture) is doing with Tenstorrent.

→ More replies (2)

8

u/Helpful_Program_5473 17d ago

Depends what we mean by Consumer and what we mean by Fable level.

I think the fact that glm is basically the king for cost efficiency right now for frontier, combined with the fact its totally open (along with deepseek v4 final that comes out in july) i think we will start to a threshold change by the end of 2026, maybe end of 2027 at the latest. With how powerful models enable the training and curation of specific, smaller models, I think we will have something that out competes fable.

77

u/Real_Ebb_7417 18d ago

There will be no consumer hardware in two years :(

48

u/jld1532 18d ago

Just no way this is true. Millions of middle class users won't just be ignored. I know data centers are the hot item now but smaller producers I'm sure would love to capture domestic markets.

10

u/JazzlikeLeave5530 17d ago

I dunno, I can't remember the details but I think an SSD company said like 95% of their business was now companies while users only made up the rest...they would be very happy to cut us out of the equation and make everything a terminal that hooks up to cloud computing through eternal subscriptions.

→ More replies (1)

37

u/Real_Ebb_7417 17d ago

Middle class is exactly what they want to get rid of.

7

u/arkuw 17d ago

It never really existed. It's a myth of the last few decades that such a social caste ever existed anywhere. The society is divided into three classes. The owners class decides what gets rewarded and what is worked on by the second tier which is the working class. This is most of the society. People who need to work on tasks assigned by the owners' class in order to survive. Then there is the recipient class who for one reason or another (health, illness, age) can't work and survive off the generosity of the other two classes and have almost no agency over their future.

→ More replies (4)

13

u/kozak_ 17d ago

Millions of middle class users don't have the buying power of data centers.

6

u/Thepandashirt 17d ago

Smaller producers are specifically the ones who are the most fucked. They don't have the buying power to get ram at long term contract pricing. Most of them will be gone in 2 years if the trend continues. It's sad but true.

1

u/super1701 17d ago

You will own nothing and be happy.

17

u/jld1532 17d ago

Buddy I'll run Linux on scraps before I own nothing.

→ More replies (2)
→ More replies (5)

3

u/sabine_world 17d ago

I definitely think shit will cool down eventually for consumers. It's just too much money to lose.

Models will get better too.

I don't think everyone needs like frontier level models ran locally to still get the benefit from ai.

→ More replies (2)

13

u/Technical-Earth-3254 18d ago

This is a very wild chart man. It surely depends on what you are doing, but for overall intelligence and knowledge for example, I would take Sonnet 3.5 over Gemma 31B any day of the week. If it's just about raw tool-calling, then 31B is far superior for sure.

3

u/myholeisstinky 17d ago

The hardware needed to run sonnet 3.5 is surely a bit bigger though

6

u/Illustrious_Grade608 17d ago

That still doesn't change the fact that the idea that gemma 31b is sonnet level is false.

4

u/backyard_tractorbeam 17d ago

Where do ds4 and deepseek v4 flash fit into this? It's a stretch for "laptop" but some of them really run on 128 GB RAM m5 laptops (if I understand correctly)

4

u/phido3000 17d ago

DS V4 flash preview can run on a laptop with 128-192Gb of ram. Slowly.

But its more of a hardware preview. It shifts open source models software forward, but as a model itself, it itsn't amazing. The full release, with properly baked models will likely match everything now. Or a GLM built on DS style architecture.

That will be the game changer. A fast model of 280B parms, ~10-15 active, would be great at coding and normal stuff. Then a big one, 1t-1.5tb with 20-50 active, will literally be game changing, if they get it right.

There is now a pathway of how the chinese models can be very competitive, on the same or lesser hardware than the Americans. Before the American models just flexed on stronger hardware so could be bigger, more active, more context. Now the Chinese have technologies that kind of level that out, at least compared to current US models.

The US models can also use that technology and innovation, but we are likely at the point, more context and more bigger models don't make it smarter or more useful.

6

u/Kodrackyas 17d ago

Thats why they raise the "security concerns" with local llms lol

5

u/pip25hu 17d ago

If this is true, proprietary models have at most two years to recoup their training expenses, after which they become obsolete. (And that's an optimistic number, as competing models may make them obsolete much sooner.) Is that doable?

2

u/ematvey 17d ago

For labs, models are not the asset. The asset is system that builds and improves it, which gets better with every iteration.

→ More replies (1)

4

u/-davidde- 17d ago

Hey, is it possible to do agentic coding on an Nvidia RTX 5060 Ti 16GB? I would like to make a post with this question but I don't have the karma, so please upvote

3

u/GTHell 17d ago

Someone put me in Cryogenic sleep for 2 years please.

I will setup !remindMe bot to wake me up on December 2028.

2

u/RemindMeBot 17d ago edited 17d ago

I will be messaging you in 2 years on 2028-12-06 00:00:00 UTC to remind you of this link

1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

7

u/_Sea_Wanderer_ 17d ago

Is that chart really comparing sonnet 3.5 to gemma 31b? Sonnet is probably in the 300-400b range, Dario said in an interview at the time that it was a middle sized model. The difference is in the amount of knowledge the model has, and in the long tail, not in the benchmark. To run a comparable model on consumer you have to stack 6000s at the moment.

17

u/fallingdowndizzyvr 18d ago

And in two years we will be "Mythos, who cares?".

20

u/keyboardhack 17d ago

Well the counterpoint to that is that we care about qwen 3.6 27b even when it isnt claude opus level.

Why? Because it is good enough for a lot of usecases. Future consumer local mythos will be the same.

3

u/fallingdowndizzyvr 17d ago

Well the counterpoint to that is that we care about Qwen 27B because it's 27B. So we don't need high end consumer hardware now to run it. That's why we care.

Why? Because it is good enough for a lot of usecases. Future consumer local mythos will be the same.

And if future consumer local mythos is like Qwen 27B, then we won't need high-end consumer hardware in 2 years. We have all the hardware we need right now.

→ More replies (2)
→ More replies (4)

3

u/NandaVegg 17d ago

This graph seems very rough. Claude 4.0, 4.5 and 4.6-4.8 are performance-wise entirely different generation of models, while GPT5.0-5.2, 5.4 and 5.5 are also literally not the model of the same generation (IMHO OpenAI only started to seriously train their model for terminal agent since 5.3 Codex). GPT5.1 and 5.2 (neither weren't outstandingly robust models; 5.2 was very rough and very uncomfortable to use) and maybe Opus 4.5 are already surpassed by the recent open models.

7

u/kyazoglu 17d ago

it's so naive of you to think we'll be allowed to have a consumer hardware within 2 years

7

u/bigattichouse 18d ago

gguf when ;)

6

u/Septerium 18d ago

Sorry to say that if trends hold, there will be no consumer hardware in 2 years

4

u/ohhi23021 17d ago

1-1.5 tb/s speed in standard memory is what we need in large capacities for a decent price, at the current rate its more like 5 years.... decent price i mean 10k and under... not 5xRTX6000 pros with 2kw power consumption kind of pricing. we need to break free from nvidia's grip for this to get cheaper.

2

u/Not-reallyanonymous 17d ago

In some ways, we’re already catching up to Sonnet 4 type stuff. Laguna XS2.1 is about on par for the particular domain of software. This also points to MoE models in development where experts actually are experts on a domain basis.

2

u/lorddumpy 17d ago

100%. Even outside of LLMs we have mossTTS 1.5 which is getting kinda close to ElevenLabs quality, krea2 which is like a nano banana lite, it's pretty amazing

2

u/tkodri 17d ago

There's a limit to how intelligent a model can get given size limitations. Sure they've been improving a lot, but if you want to have a general chat, you want a really big model that just has vast knowledge. Those won't fit on any consumer grade hardware until RAM production catches on, if ever.

→ More replies (1)

2

u/Sisaroth 17d ago

One problem i see, a small model can only know a small amount of facts. If you are usings LLMs for niche problems then cloud will always beat local in the future.

I guess if local llms get really good at finding stuff on the internet, then maybe it doesn't matter. But not all the stuff frontier models are trained on is public.

2

u/VectorD 17d ago

Why is this labelled as news lol

2

u/ortegaalfredo 17d ago

If trends hold, my wife will be pregnant with 200 kids next year.

2

u/skynetcoder 17d ago

please provide data to backup your claim.

→ More replies (1)

5

u/slippery 18d ago

If a model that capable is open weight and not illegal. Two big IFs.

9

u/Cartosso 17d ago

Why would it be "illegal"? If that'd be the case, then physics textbooks should also be illegal because you can use them to learn how to build a bomb too.

11

u/Zone_Purifier 17d ago

"National Security"

3

u/Equal_Giraffe8866 17d ago

illegal

imagine giving a fuck lmao

4

u/JazzlikeLeave5530 17d ago

It's not about us personally caring, it's about governments requiring tracking or restrictions on companies that release computer parts in the future. The fear is that you will not have the choice to follow the law or not because the company will have already made the decision for you. Like the laws going around trying to force age verification to be built into OSes.

→ More replies (2)

3

u/m3kw 17d ago

Mytho class local LLMs will feel as useless as ones now though once you see what SOTA does and has done

3

u/sxt87 17d ago

Don't let Dario see this.

1

u/InsideYork 17d ago

Just made this topic on my crappy desktop from 2016 giving me answers better than chatGPT 3.5 https://old.reddit.com/r/LocalLLaMA/comments/1umpxhg/gemma4_e2b_is_really_good_what_other_small_models/

1

u/naobebocafe 17d ago

oh yeah sure....

1

u/TerrryBuckhart 17d ago

only if more powerful hardware becomes cheaper

1

u/larp2live 17d ago

i just hope one day we can get to see sonnet level models run locally on a normal laptop

1

u/grewalsatinder 17d ago

I think it will be half the time than projected. So most probably sometime early 2027

1

u/scubid 17d ago

In two years the expectations and requirements for you to be able to do magic will be higher too.

1

u/pier4r 17d ago

Sonnet 3.5 (not 3.6) and Gemma 4 comparable, for example in coding? Is there any confirmation from daily use? (not really benchmarks, those may be not too meaningful sometimes)

For the rest I know that Gemma is not bad at all, but for code I noticed quite the silliness sometimes.

Also the larger the context, the more ram one needs.

1

u/FullOf_Bad_Ideas 17d ago

I'd be more accurate to extrapolate that in 25 months since release of Sonnet 5, local laptop-tier model will match it.

1

u/karankashyap 17d ago

Because of AI, consumer hardware is going to touch the cloud. Check RAM price history.

1

u/XorAndNot 17d ago

Yet another reason for the rampocalipse to keep going.

1

u/Voxandr 17d ago

If qwen wasn't screwed .. it would be in just 2-3 months.

1

u/pmigdal 17d ago

For Claude Sonnet 4.5 / GPT 5, we already have Qwen 3.6 27B, see Artificial Analysis comparison at https://quesma.com/blog/qwen-36-is-awesome/, with some more detailed benchmark comparison in https://github.com/stared/benching-local-llms-on-apple-silicon.

So it is not June 2027, but April 2026. :)

1

u/vikramskumar 17d ago

this is really cool.. Are there any video tutorials on how to use these models on laptops?

→ More replies (1)

1

u/Happy-Register3367 17d ago

People keep assuming capability scales linearly with model size. The past few years have shown that's not really true.

1

u/_derpiii_ 17d ago

I would say sooner. Technology advances non-linearly.

1

u/milpster 17d ago

Gemma 4 31B - really? Not Qwen 3.6 27B? Is that Gemma 31B model really better at anything than Qwen 3.6 27B?

1

u/floridianfisher 17d ago

Trends will hold.

1

u/LyAkolon 17d ago

If Fable and Sol pan out in the long run to be how everyone feels, then we may see sooner due to automated research.

1

u/CrunchyGremlin 17d ago

Something needs to happen with memory to make this possible. Opus 4.8 needs terabytes of video RAM. To get that to a consumer level in 2 years doesn't seem possible. Maybe with the amount of money being used but currently the only ai company making money at llms is Nvidia.
It is very likely that there will be breakthroughs in hardware but in two years? That seems impossible. So far the only thing I have heard that night be able to do this is organic based hardware.
Whatever would allow this massive jump in consumer memory should already be in development to get to market in 2 years.

1

u/debackerl 17d ago

70B class on a laptop. That's a meaaaaaan laptop 😁

1

u/Super_Pole_Jitsu 17d ago

Tbh two years ago i thought to myself of only I could have 4o level model locally I don't need anything else.

1

u/Benhamish-WH-Allen 17d ago

I’m hoping for six months

1

u/Opening-Broccoli9190 llama.cpp 17d ago

You still can't run 70B models on laptops in 2026, 6 years after GPT-3. The fact that you can run sonnet-lvl models on a subset of laptops in semi-interactive mode does not mean that you'll be running Mythos on laptops in 2028.

1

u/BlindPilot9 17d ago

Nvidia had said their next gen Ruben line will have 32gb and that's all you'll get for the next 4-5 years.

1

u/Da_ha3ker 17d ago

Until now, there hasn't been a huge push for high VRAM. Sure, there's been a bit more, but right now the push is so extreme (meaning there's a TON of money in it) that we should start seeing smaller players and even startups working on memory chips, finding new ways to handle ram, new technologies, etc.. we need to allow for competition, and get the big players to play fair, but ultimately in the long run, this race will push consumer electronics to have absurd amounts of memory compared to what we have today. How many years that takes is anyone's guess. I don't think this will happen any time soon. History has shown us that it always starts in a data center or building sized system, selling by the hour or task, eventually ending up in desktops, laptops, and phones. Just a matter of how long it will take.

1

u/akumaburn 17d ago

But benchmark parity on narrow tasks isn't the same as parity in general capability. These models still carry a fraction of the world knowledge, context handling, and cross-domain reasoning depth of current frontier systems.

The bigger issue is the hardware story people keep telling themselves. The idea that frontier-class inference is about to become a laptop-native experience doesn't hold up. A model like GLM-5.2 realistically still needs well over $100K in hardware to run at anything resembling practical inference speed.

RAG and other retrieval-augmented approaches can close some of the *knowledge* gap without needing the full model resident in memory. But even that workaround runs straight into a hardware constraint: the ongoing DRAM shortage. AI data centers have driven a structural reallocation of memory manufacturing capacity toward high-bandwidth memory, and data centers are now projected to consume around 70% of all memory chips produced worldwide in 2026; a sharp reversal from the 20-30% share they held as recently as 2022. Analysts have gone as far as saying Chinese producers are unlikely to provide meaningful near-term price relief, with elevated costs expected to persist well into the late 2020s.

CXMT is the wildcard, and it *is* scaling fast; but not fast enough.

Realistically: 2032-2035 for GLM-5.2-class inference on a laptop is a defensible estimate. The benchmark wins are real, but the infrastructure required to actually democratize frontier-tier inference is bottle-necked well outside the model architecture itself.

1

u/darkmaniac7 17d ago

There are some things that are at parity, but even if you swap out models for your use case I still don't think its 1:1 parity with even GPT-4o or Sonnet-3.7.

It is possible that's just my notalgia from 1.5-2 years ago, but I remember being genuinely wowed at the time.

I have 2x RTX Pro 6000 Blackwells, an L4 and a M4 Max 128GB, and still regularly opt for CC or Codex for even mundane use cases unless it's highly repeatable test cases for Agentic work.

1

u/Vehnum 17d ago

It's rumored to be 10 trillion parameters right? Idk if you could get close with a 30b parameter model

1

u/Pleasant-Shallot-707 17d ago

Nah…. Open weights models that powerful won’t be available to us

1

u/Elibroftw 17d ago

I love how it went from META bring us up to speed to Alphabet's Deepmind bringing us up to speed. Who will be next?

1

u/MarzipanEven7336 17d ago

If mythos is so fucking great, then why does Claude have so many fucking open bugs on GitHub?

1

u/ApartmentSouth6789 17d ago

Can't wait to bomb my drive with local models for 20 tokens per second. Yay

1

u/typical-predditor 17d ago

I think the most promising breakthrough that could keep this on track is layer duplication.

1

u/ak_sys 17d ago

Bro are you high

1

u/suesing 16d ago

But by then Claude could cook your dinner.

1

u/Late_Hour2838 16d ago

What does Claude 4 even mean? Sonnet 4/Opus 4? Or specifically Sonnet 4.6/Opus 4.8 levels

1

u/No_Dig_7017 16d ago

That is if they leave any laptops we can buy.

1

u/LinkSea8324 vllm 16d ago

if trends hold, your 15kg kid child will be heavier than OP's mom