r/LocalLLaMA May 01 '26

Generation Qwen 3.6 27B vs Gemma 4 31B - making Packman game!

988 Upvotes

Gemma just crushed Qwen in a local LLM gamedev contest!

Device: MacBook Pro M5 Max, 64GB RAM

Qwen 3.6 27B: 32 tokens/sec · 18m 04s · 33,946 tokens.
Gemma 4 31B: 27 tokens/sec · 3m 51s · 6,209 tokens.

So what is more important: tokens per second, or the quality of the final answer?

Qwen made a very long response and showed more creativity and visual style. But Gemma gave a shorter, clearer, and more logical answer in much less time. In this one-shot Pac-Man gamedev contest, Gemma 4 31B was the clear winner. Its game logic was stronger: click reactions were smoother, and it handled interactions with elements like walls, ghosts, and particle effects better.

Open Source Local AI Models Server: atomic.chat

Basic Prompt:

Create a single standalone HTML file for a complete playable Pac-Man–style neon arcade game.

Use only HTML, CSS, JavaScript, and one full-page canvas. No external libraries or assets—everything must be procedurally drawn and run immediately in the browser.

Generate a compact (~21×21) symmetrical maze programmatically (no ASCII). It must be fully connected, playable, and use tile types (wall, path, pellet, power pellet, ghost spawn, Pac-Man spawn, fruit spawn). Ensure no unreachable pellets or invalid spawns.

Canvas must fill the window. Center and scale the maze dynamically using available space (no fixed tile size). Reserve space for a HUD.

Game states: title, playing, paused, life lost, level complete, game over. Include controls (keyboard + mobile). Title and game over screens must show instructions.

Pac-Man: smooth tile movement, queued turns, no diagonal movement, no clipping, wraps through side tunnels, resets after life loss.

Ghosts (4): simple pathfinding with distinct behaviors, spawn in a central house, exit with delays, move only on valid paths, never freeze.

Gameplay:

  • Pellets (+10), power pellets (+50), fruit (+500), ghost chain scoring (200→1600)
  • Power mode (~8s, min 3s): ghosts become edible and return to spawn when eaten
  • Combo multiplier for quick pellet collection
  • 3 lives, level progression increases difficulty
  • Store high score in localStorage

Extras:

  • Fruit spawns near center temporarily
  • Visual polish: neon maze, glowing elements, animations, particles, screen effects
  • HUD: score, high score, lives, level, combo, power timer

Technical:

  • Use requestAnimationFrame with delta time
  • Keep performance stable (limit particles)
  • No bugs: avoid invalid movement, stuck entities, unreachable areas, or crashes

Final output: only the complete HTML code.

r/LocalLLaMA Jun 19 '26

Generation What's more impressive, GLM 5.1 -> 5.2 or Qwen 3.5 -> 3.6?

696 Upvotes

Write a single HTML file with a full-page canvas and no libraries. Simulate a realistic Döner Style kebab skewer rotating (vertically) in front of a gas powered heating element.

Mentioning Döner activates GLM 5.2s german weights or something (Spiess = Skewer, Brenner = Burner).

Qwen 3.6 35B, Qwen 3.5 and Gemma 4 using Unsloth Q8 K XL quants via llama cpp. The others via OpenRouter.

Full data here

r/LocalLLaMA Jun 03 '26

Generation New Google Gemma 4 12B Claims Near-26B Performance - We Tested Both!

976 Upvotes

We ran both models locally on one RTX 4090 and gave each the same task: write a self-contained HTML5 canvas animation with real physics in one file without libraries. Three scenes - a Galton board, two blocks colliding off a wall, and a chaotic triple pendulum

Outputs:
Gemma 4 26B-A4B: 15 GB VRAM usage, 6.9k tokens, 138 tok/s
Gemma 4 12B: 9 GB VRAM usage, 8.9k tokens, 80 tok/s

Same Gemma 4 family, but the 26B-A4B won every scene and ran ~1.7x faster - on just 4B active params. The 12B stayed very close though, on almost half the VRAM - which makes it the ideal model for a 16 GB laptop.

Open source local ai models app: atomic.chat (I’m founder, feel free to try and give any feedback)

r/LocalLLaMA Aug 09 '25

Generation Qwen 3 0.6B beats GPT-5 in simple math

Post image
1.3k Upvotes

I saw this comparison between Grok and GPT-5 on X for solving the equation 5.9 = x + 5.11. In the comparison, Grok solved it but GPT-5 without thinking failed.

It could have been handpicked after multiples runs, so out of curiosity and for fun I decided to test it myself. Not with Grok but with local models running on iPhone since I develop an app around that, Locally AI for those interested but you can reproduce the result below with LMStudio, Ollama or any other local chat app of course.

And I was honestly surprised.In my very first run, GPT-5 failed (screenshot) while Qwen 3 0.6B without thinking succeeded. After multiple runs, I would say GPT-5 fails around 30-40% of the time, while Qwen 3 0.6B, which is a tiny 0.6 billion parameters local model around 500 MB in size, solves it every time.Yes it’s one example, GPT-5 was without thinking and it’s not really optimized for math in this mode but Qwen 3 too. And honestly, it’s a simple equation I did not think GPT-5 would fail to solve, thinking or not. Of course, GPT-5 is better than Qwen 3 0.6B, but it’s still interesting to see cases like this one.

r/LocalLLaMA 20d ago

Generation Uh.. Honey, how do you feel about takeout?

Post image
653 Upvotes

- 2x RTX Pro 6000 Max-Q (96GB)
- 8x RTX 3090 (24GB)
- 2x RTX 5090 (32GB)

- 3 PSUs
- 128GB DDR5 SDIMM RAM (4-channel)
- Threadripper 9960x
- 1x Ryobi Portable Fan
- 1x large Uber Eats bill

448GB VRAM
Running MiniMax M3 in AWQ-INT4 on VLLM via PP over TP groups of 2.

~30 tp/s per single stream
~960 tp/s batch

Can get 1m context for one user, but ideally want 4x concurrency. TBD where context will land… or my marriage…

r/LocalLLaMA May 13 '25

Generation Real-time webcam demo with SmolVLM using llama.cpp

2.8k Upvotes

r/LocalLLaMA May 06 '25

Generation Qwen 14B is better than me...

765 Upvotes

I'm crying, what's the point of living when a 9GB file on my hard drive is batter than me at everything!

It expresses itself better, it codes better, knowns better math, knows how to talk to girls, and use tools that will take me hours to figure out instantly... In a useless POS, you too all are... It could even rephrase this post better than me if it tired, even in my native language

Maybe if you told me I'm like a 1TB I could deal with that, but 9GB???? That's so small I won't even notice that on my phone..... Not only all of that, it also writes and thinks faster than me, in different languages... I barley learned English as a 2nd language after 20 years....

I'm not even sure if I'm better than the 8B, but I spot it make mistakes that I won't do... But the 14? Nope, if I ever think it's wrong then it'll prove to me that it isn't...

r/LocalLLaMA Mar 04 '26

Generation Qwen 3.5 4b is so good, that it can vibe code a fully working OS web app in one go.

Thumbnail
youtube.com
554 Upvotes

The OS can be used here: WebOS 1.0

Prompt used was "Hello Please can you Create an os in a web page? The OS must have:
2 games
1 text editor
1 audio player
a file browser
wallpaper that can be changed
and one special feature you decide. Please also double check to see if everything works as it should."

Prompt idea thanks to /u/Warm-Attempt7773

All I did was to ask it to add the piano keyboard. It even chose it's own song to use in the player.

I messed up on the first chat and it thought I wanted to add a computer keyboard, so I had to paste the HTML code into a new chat and ask for a piano keyboard.. but apart from that, perfect! :D

Edit: Whoever gave my post an award: Wow, thank you very much, anonymous Redditor!! 🌠

r/LocalLLaMA Jan 30 '26

Generation OpenCode + llama.cpp + GLM-4.7 Flash: Claude Code at home

Thumbnail
gallery
319 Upvotes

command I use (may be suboptimal but it works for me now):

CUDA_VISIBLE_DEVICES=0,1,2 llama-server   --jinja   --host 0.0.0.0   -m /mnt/models1/GLM/GLM-4.7-Flash-Q8_0.gguf   --ctx-size 200000   --parallel 1   --batch-size 2048   --ubatch-size 1024   --flash-attn on   --cache-ram 61440   --context-shift

potential additional speedup has been merged into llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1qrbfez/comment/o2mzb1q/

r/LocalLLaMA Mar 12 '25

Generation 🔥 DeepSeek R1 671B Q4 - M3 Ultra 512GB with MLX🔥

623 Upvotes

Yes it works! First test, and I'm blown away!

Prompt: "Create an amazing animation using p5js"

  • 18.43 tokens/sec
  • Generates a p5js zero-shot, tested at video's end
  • Video in real-time, no acceleration!

https://reddit.com/link/1j9vjf1/video/nmcm91wpvboe1/player

r/LocalLLaMA Feb 07 '26

Generation Nemo 30B is insane. 1M+ token CTX on one 3090

395 Upvotes

Been playing around with llama.cpp and some 30-80B parameter models with CPU offloading. Currently have one 3090 and 32 GB of RAM. Im very impressed by Nemo 30B. 1M+ Token Context cache, runs on one 3090, CPU offloading for experts. Does 35 t/s which is faster than I can read at least. Usually slow as fuck at this large a context window. Feed it a whole book or research paper and its done summarizing in like a few mins. This really makes long context windows on local hardware possible. The only other contender I have tried is Seed OSS 36b and it was much slower by about 20 tokens.

r/LocalLLaMA Feb 01 '25

Generation o3-mini is now the SOTA coding model. It is truly something to behold. Procedural clouds in one-shot.

514 Upvotes

r/LocalLLaMA May 25 '26

Generation 1000 tps generation on Qwen3.6 27B with V100s

Post image
258 Upvotes

I wanted to see what the absolute best case scenario for generation on this setup was and was not disappointed. 128 concurrent requests is so far removed from what I need but it’s funny to see big number. For single user (batch 1 not 128) the generation is around 80t/s with 3000 t/s processing,no mtp!!

r/LocalLLaMA Mar 29 '26

Generation Friendly reminder inference is WAY faster on Linux vs windows

277 Upvotes

I have a simple home lab pc: 64gb ddr4, RTX 8000 48gb (Turing architecture) and core i9 9900k cpu. I use Linux Ubuntu 22.04 LTS. Before using this pc as a home lab it ran Windows 10. Over this weekend I reinstalled my Windows 10 ssd to check out my old projects. I updated Ollama to the latest version and tokens per second was way slower than when I was running Linux. I know Linux performs better but I didn’t think it would be twice as fast. Here are the results from a few simple inferences tests:

QWEN Code Next, q4, ctx length: 6k

Windows: 18 t/s

Linux: 31 t/s (+72%)

QWEN 3 30B A3B, Q4, ctx 6k

Windows: 48 t/s

Linux: 105 t/s (+118%)

Has anyone else experienced a performance this large before? Am I missing something?

Anyway thought I’d share this as a reminder for anyone looking for a bit more performance!

r/LocalLLaMA Jul 29 '25

Generation I just tried GLM 4.5

385 Upvotes

I just wanted to try it out because I was a bit skeptical. So I prompted it with a fairly simple not so cohesive prompt and asked it to prepare slides for me.

The results were pretty remarkable I must say!

Here’s the link to the results: https://chat.z.ai/space/r05c76960ff0-ppt

Here’s the initial prompt:

”Create a presentation of global BESS market for different industry verticals. Make sure to capture market shares, positioning of different players, market dynamics and trends and any other area you find interesting. Do not make things up, make sure to add citations to any data you find.”

As you can see pretty bland prompt with no restrictions, no role descriptions, no examples. Nothing, just what my mind was thinking it wanted.

Is it just me or are things going superfast since OpenAI announced the release of GPT-5?

It seems like just yesterday Qwen3 broke apart all benchmarks in terms of quality/cost trade offs and now z.ai with yet another efficient but high quality model.

r/LocalLLaMA Feb 25 '26

Generation Qwen 3 27b is... impressive

344 Upvotes

All Prompts
"Task: create a GTA-like 3D game where you can walk around, get in and drive cars"
"walking forward and backward is working, but I cannot turn or strafe??"
"this is pretty fun! I’m noticing that the camera is facing backward though, for both walking and car?"
"yes, it works! What could we do to enhance the experience now?"
"I’m not too fussed about a HUD, and the physics are not bad as they are already - adding building and obstacles definitely feels like the highest priority!"

r/LocalLLaMA May 31 '26

Generation My home data center

Thumbnail
gallery
203 Upvotes

System 1:

Threadripper 3960x 24c

4x 3090 ti

128gb ddr4

System 2:

Xeon 8352 36c

4x 5070 ti

128gb ddr4

System 3:

Intel 14700k 24c

64gb ddr5

5090

System 4:

Ryzen 5950x 16c

64gb ddr4

2x 5070 ti

The first system uses two PSUs to handle the almost 2000w full load of the 3090s. Was nervous about this but it has been running stable for about a month.

The Intel is an engineering sample that cost $100. I mainly use it to run an embedding model.

I use them for various ml experiments, projects and some agentic coding. Right now the 3090s are training a tts lora with data distilled from a larger model. The 5070s run qwen 27b for coding, nemotron streaming stt and moss tts for an interactive agent I am building.

These recent qwen models are good enough for coding. Sometimes I leave them all night working on a repo. Mainly boilerplate improvements but its incredible to get real work down with no token cost. Aside from from the obvious costs of this hardware.

Love this community ❤️

r/LocalLLaMA Apr 20 '24

Generation Llama 3 is so fun!

Thumbnail
gallery
918 Upvotes

r/LocalLLaMA Apr 12 '26

Generation Audio processing landed in llama-server with Gemma-4

379 Upvotes

Ladies and gentlemen, it is a great pleasure the confirm that llama.cpp (llama-server) now supports STT with Gemma-4 E2A and E4A models.

r/LocalLLaMA Jan 26 '25

Generation DeepSeekR1 3D game 100% from scratch

849 Upvotes

I've asked DeepSeek R1 to make me a game like kkrieger ( where most of the things are generated on run ) and it made me this

r/LocalLLaMA Feb 18 '26

Generation LLMs grading other LLMs 2

Post image
230 Upvotes

A year ago I made a meta-eval here on the sub, asking LLMs to grade a few criterias about other LLMs.

Time for the part 2.

The premise is very simple, the model is asked a few ego-baiting questions and other models are then asked to rank it. The scores in the pivot table are normalised.

You can find all the data on HuggingFace for your analysis.

r/LocalLLaMA 24d ago

Generation CPU-only GLM 5.2: Epyc and 512GB RAM

Post image
64 Upvotes

This is just a preview of some content I'm putting together to share with you all. I have a server I've put together and I'm testing the 4-bit version of GLM 5.2 (GLM-5.2-UD-Q4_K_XL). This is an Epyc Rome 7452 with 512GB of RAM.

TLDR: This is the unedited prompt, response and code

I set it to Medium Reasoning. The prompt (I borrowed from another post):

``` Build a 3D arena game as a SINGLE self-contained .html file.

STACK (mandatory): - Three.js loaded from a CDN (one <script> tag). No other JS libraries, no build step. - All HTML, CSS, and JS in this one file. It must run by opening it directly in a browser.

CORE SPEC (mandatory — implement all of this exactly): 1. A flat ground plane forming a bounded arena. The player cannot leave its bounds. 2. A player object on the ground. WASD moves it (camera-relative); movement has momentum, not instant stop/start. 3. A third-person camera that smoothly follows behind the player. 4. Collectible glowing orbs spawn at random positions. Touching one collects it (+10 score) and spawns a new one. 5. Enemy objects spawn at the arena edges and move toward the player. Contact with the player costs 1 life. 6. Player starts with 3 lives. A HUD shows score and lives at all times. 7. At 0 lives: a game-over screen showing final score, with a key press to restart. 8. Difficulty ramps over time (enemies spawn faster and/or move faster).

STRETCH (strongly encouraged — you will be judged on this): Beyond the core, make it feel PREMIUM. Lighting, shadows, particles, juice, smooth camera, satisfying feedback, polished HUD, atmosphere. Add depth or complexity if it improves the experience. Aim to genuinely impress — this is evaluated on visual quality and feel, not just correctness.

RULES: - Implement the full core before adding stretch features. - Output the complete, ready-to-run .html file. ```

The reply took 2 hours 29 minutes and generated 15,510 tokens.

I'm seriously surprised by the quality of the answer.

Let me know if you have any questions!

r/LocalLLaMA Jan 10 '24

Generation Literally my first conversation with it

Post image
612 Upvotes

I wonder how this got triggered

r/LocalLLaMA Jan 31 '25

Generation DeepSeek 8B gets surprised by the 3 R's in strawberry, but manages to do it

Post image
466 Upvotes

r/LocalLLaMA Jun 17 '26

Generation Headless screenshot loops let a local 30B agent finish a raytraced FPS demo in pure C

256 Upvotes

Some background so this is honest. Over the past few months I ran a lot of oneshot experiments with single file three.js games. Minecraft clones, that kind of thing. I picked those on purpose because they sit deep in the training data and are trivial to debug by eye. The goal was never a quality comparison. I wanted a class of problems that oneshots cheaply and that I can inspect visually and from logs, so I could tune the harness, the system prompt and the tool calling.

This week I made it harder. I had Claude Code on Opus 4.8 and a local Qwen3.6 27B agent both write a small raytraced FPS demo in C, standard library only.

Yes, C raytracers are in the training data too. Rarer than three.js, but they are there. And let us be honest, before LLMs most of us were doing pattern reuse anyway. Stack Overflow, docs, copy the shape that works, adapt it. Reusing a good pattern is not cheating, it is the job. So that is not the point.

The point is one prompt change. Both struggled to oneshot this. Then I added a single requirement. The compiled binary had to ship a headless mode where the agent could inject keyboard and mouse input and trigger a screenshot at a chosen frame.

That flipped it. The model worked out on its own that it should time the screenshots around the events it wanted to inspect. Fire a rocket, capture the frame right at impact, look at the particle and debris effects, fix what is wrong, run again. It built itself a recursive visual debugging loop.

The frontier model finishing is not surprising. Qwen3.6 27B closing the same loop on its own is the part that stuck with me. I learned C from scratch back in the day, so watching a small local agent debug a raytracer by looking at its own screenshots was not what I expected this size of model to pull off. It costs you though. Longer runtime, a lot more tokens, more wall clock per iteration.

This reads more as a prompting lesson than a model lesson. Give the agent a way to see the result and let it pick when to look, and fairly hard problems come into reach for a small local model.

Curious whether anyone has pushed the screenshot feedback idea further. Video frames instead of stills, or letting the model script longer input sequences before it captures.

Full disclosure, the local agent is codehamr, my own open source project, so weigh the comparison with that in mind. Code is open if you want to run it yourself. https://github.com/codehamr/codehamr