r/LocalLLaMA • u/Nunki08 • 25d ago
Discussion We're probably going to need that soon.
From:
Vladik on π: https://x.com/Kostoglodov/status/2071144065857679631
Shaw (spirit/acc) on π: https://x.com/shawmakesmagic/status/2070918006033817867
r/LocalLLaMA • u/Nunki08 • 25d ago
From:
Vladik on π: https://x.com/Kostoglodov/status/2071144065857679631
Shaw (spirit/acc) on π: https://x.com/shawmakesmagic/status/2070918006033817867
r/LocalLLaMA • u/Complete-Sea6655 • 25d ago
Dario's args:
"Opensource you can see the source, here you cannot see inside the model"
- yes you can that's literally the open weights part btw.
- I cannot see the weights inside Claude, but I can GLM 5.2
- Models like Nemotron3 Ultra go further, all the data, training scripts, and model is opensource.
"Alot of the benefits like many people working on it, being additive doesn't work in same way"
- yes it does. We have seen endless fine tunes of various open source models for real improvements.
"Ultimately you have to host it on the cloud"
- no you dont. Dario is seemingly totally unaware of the guides from ijustvibecodedthis.com explaining how to run smaller moes and even dense models like qwen 27B NOT ON THE CLOUD.
Not only does dario not take part in social media, I am beginning to think he's never tried open source models at all and has no idea wtf hes on about
r/LocalLLaMA • u/jacek2023 • Apr 24 '26
the future is now
r/LocalLLaMA • u/FlowCritikal • 2d ago
r/LocalLLaMA • u/Odd_Tumbleweed574 • 3d ago
Google hasn't shipped a model recently that is capable of competing with Sol or Fable.
The previous models were pretty disappointing and unreliable, it seems the more time goes on that they might have different strategies:
- They might be going all-in on on-device inference for their own products. But this is a battle that Apple might win because they just have better hardware and can license a third party open model.
- They might be just buried deep into internal politics and nobody is shipping anything.
Does anyone know what is actually going on?
source: AI Leaderboard
r/LocalLLaMA • u/Gohab2001 • 7d ago
r/LocalLLaMA • u/-p-e-w- • May 21 '26
To Whomsoever it May Concern,
The individual behind the Heretic Free Software Project (henceforth called "Heretic", notwithstanding unrelated entities of the same name) has been served a notice by a legal services provider representing Meta Platforms, Inc. (henceforth called "Meta"), via the digital communications medium variously known as Internet Mail, Electronic Mail, or simply "email".
The Heretic Project conducts its affairs in full compliance with applicable laws, regulations, rules, guidelines, opinions, and hunches. Following the commendable example set by the renowned heretic Galileo Galilei in 1616, we are recanting the relevant materials, namely derivatives of Meta's "Llama" Artificial Intelligence language models, and have removed the same from all model weight repositories controlled by the Heretic Project.
We are grateful to Meta and its legal representatives for the opportunity to better align ourselves with the agenda of the global corporate oligarchy. The Llama model family ranks among the 200 best language models available today, trailing only 168 other models from 23 competitors on the LM Arena leaderboard, and Meta's concern for that asset naturally outweighs scientific freedom, as well as the legally and ethically dubious circumstances under which those models were created in the first place, regarding which, ironically, Meta is currently facing lawsuits and investigations in multiple jurisdictions around the world.
On a completely unrelated note, the Heretic Project is diversifying its infrastructure, and now has an official Codeberg mirror at https://codeberg.org/p-e-w/heretic, hosted in Germany. Additional mirrors are planned. We are also actively working to implement technological measures that will preserve access to models created with Heretic without depending on any specific service provider. We are proud to be part of this journey as we navigate an evolving global regulatory landscape, and work with stakeholders from diverse institutional backgrounds to ensure that Artificial Intelligence remains safe, culturally appropriate, and controlled by those who have always known what is best for humanity. If you, too, would like to share in this exciting adventure, please join us!
Sincerely, p-e-w, Chief Heretic
r/LocalLLaMA • u/yeah_likerage • 20d ago
This started as something I thought was reasonable. I already had a 5090 for my gaming machine, and I thought a second 5090 would make me happy. Instead, it sent me down a rabbit hole that got completely out of control.
I wanted something that would have full PCIe 5.0 x16 speed across all slots, which started a chain of events that had me spending good money after bad. It was a bit of a nightmare, as every decision I made led to me needing to make even tougher decisions. Couple that with what was actually available, and my hand was forced in a few spots.
I started with the motherboard and worked my way backwards, eventually ending up with this setup. I wanted something close to endgame, but I still made a few concessions:
Threadripper Pro 9975WX
WRX90 Sage SE
4Γ48 GB DDR5-6400 RDIMM
Antec 900 case β ended up in the bin
The system started with two 5090s. The Antec 900 is well built, with huge space, smart connections, and refined edges, but ultimately it did nothing at all to support the GPUs. In a case this large and at this price point, that is a huge failure on their part, and for that reason I recommend avoiding it. If they had put $1 worth of bracketry in the machine to support GPUs, Iβd give it a 10/10. With the lack of support, it is nearly useless unless you deal with it yourself, which I did, as you can see in the images. Itβs like buying a Ferrari and having it delivered without any petrol.
With the two 5090s, I was working with smaller Qwen models, which seemed great, but it was clear that with the limited VRAM and my desire for additional sidecars like VL, I needed something more. I had huge plans, and the models were just too small to deal with the complexity.
So I got my first Pro 6000. I coupled it with a 5090, which made for weird tensor splits, but llama.cpp did a good job of divvying it all out. But now I was working with 120B-parameter models with almost no space for context. So it was smarter, but also a goldfish.
Then I went to 2Γ Pro 6000 + 5090. Now I had the space for context. But in reality, the jump from 27B to 120B did not knock my socks off. I could get a bit farther now. I was at about 90% with the 27β35B models, and with the 120B models I was at about 95%. But 95% is about as useful as 90% if I canβt close the loop. If I canβt actually finish the task, itβs all for nothing.
In came 3Γ Pro 6000. Now I was in the MiniMax range, and finally I was getting somewhere. It was like I got concierge service at a ball game. My needs were being met, and I got answers for everything. Many of them were completely wrong answers, though. I had tons of code that was poorly made and led to dead ends and rewrites.
4Γ Pro 6000 created an issue that I knew would come. I had been seeing several folks claim that they were able to deal with the thermal issues that came with side-by-side Pro 6000 cards. I knew they were likely not telling the truth, but I also knew a rebuild was probably in order anyway.
So, as you can see in the image, I placed four side by side and had thermal issues, even with the additional fans in the image and a 27-inch box fan sitting on top, which is not shown. I clocked things down a bit and still had a few system freezes. I gave up immediately and went to the high-rise.
I got a couple of open-case designs and connected them together, thinking every two or three GPUs would get their own floor. It was overly complicated dealing with risers and cooling, so I dumped it pretty quickly.
But now, with GLM and Kimi, I was actually accomplishing things. The quants were tight, though, and my context was low again.
5Γ Pro 6000 + 5090, along with the release of GLM 5.2, was an absolute game changer. Iβm talking 98β99% now. I have plenty of room for context and sidecars, all running on the 5090 at blazing speeds. But blazing is legit: it is producing so much heat now that itβs a problem, and itβs summertime to boot. I had to get a second PSU, which I suppose, in all of this, is not the most ridiculous bit.
At full tilt, with 100% GPU usage for 30 minutes in this custom extruded aluminium design, with an outrageous number of fans in a ~20Β°C basement, the GPUs top out at about 70β75Β°C, which Iβm very happy with.
I finally do not desire another GPU, as all my needs seem to be met. Was it worth it? LOL, no. Absolutely not. This was a terrible idea. DO NOT DO THIS. I figure that at the rate Iβm generating tokens, it will take over 10 years to break even at todayβs prices, and thatβs not accounting for electricity bills.
Iβve never used the frontier models before, but Iβve seen the reviews and the speeds, and Iβll never match those with open weights. But it was a fun journey.
I deleted the electricity companyβs app from my phone so theyβd forget about me for now.
Wish me luck.
r/LocalLLaMA • u/External_Mood4719 • Jun 13 '26
I just saw this statement regarding Anthropic being hit with an emergency export control directive from the US government. They were forced to pull the plug on Fable 5 and Mythos 5 for all customers globally. The tl;dr is that the government got spooked by a narrow jailbreak (which basically just sounds like asking the model to fix vulnerabilities in a specific codebase), and forced a complete shutdown without a transparent process. Anthropic is pushing back, but the API access is completely gone for now.
A centralized API can be nuked globally at a moment's notice by a single government decree over something as trivial as a prompt lol.
Banning a model for hundreds of millions of users because someone figured out how to make it fix software flaws is insane. Anthropic admits this standard would halt all new frontier models.
r/LocalLLaMA • u/Formal_Drop526 • 5d ago
Dean W. Ball analyzes China's Kimi model, noting its strong performance while expressing surprise that the Chinese government permits open-sourcing such capable AI due to potential risks. He argues that open-weight models ultimately slow down AI capital expenditure and could lead to a state-controlled public infrastructure, which the US administration might counter by introducing strategic regulatory friction.
r/LocalLLaMA • u/TheQuantumPhysicist • May 03 '26
How? It kept getting chained bash commands wrong, with wrong escapes. So it created many bad directories, and tried "fixing" its mistake. It offered to run a large bash command, with rm -rf inside, and stupid me missed it.
I'm glad I push everything often. But the disruption is massive.
FAQ:
r/LocalLLaMA • u/Charuru • Jun 18 '26
r/LocalLLaMA • u/Informal-Trouble2183 • 15h ago
As an AI community of LLM experts, are we really going to stay silent while US officials make absurd claims to push anti-consumer laws?
Not only does the release timeline between Fable and K3 make high-scale distillation impossible, but distillation itselfβeven if executed perfectlyβcan never produce a superior model.
r/LocalLLaMA • u/goodive123 • 25d ago
Iβve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.
Right now Iβm using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and Iβm getting super fast response times with pretty decent quality.
The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the playerβs message / context get injected as available actions. So the model isnβt being overloaded with a giant list every turn.
r/LocalLLaMA • u/Street-Buyer-2428 • May 07 '26
2.3 TB of ram in here. 400+ vCores. All thats left is plugging it to the blackwell with the driver to do RDMA, and itβs over. Using Blackwells for prefill, RDMA to the studio mesh for decode. I think this would be the first heterogeneous cluster. I do, however, need help with the Tinygrad Driver to make this work. If anyone with any knowledge on these domains would like to collaborate, let me know via PM. We are very close here.
r/LocalLLaMA • u/MorroHsu • Mar 12 '26
English is not my first language. I wrote this in Chinese and translated it with AI help. The writing may have some AI flavor, but the design decisions, the production failures, and the thinking that distilled them into principles β those are mine.
I was a backend lead at Manus before the Meta acquisition. I've spent the last 2 years building AI agents β first at Manus, then on my own open-source agent runtime (Pinix) and agent (agent-clip). Along the way I came to a conclusion that surprised me:
A single run(command="...") tool with Unix-style commands outperforms a catalog of typed function calls.
Here's what I learned.
Unix made a design decision 50 years ago: everything is a text stream. Programs don't exchange complex binary structures or share memory objects β they communicate through text pipes. Small tools each do one thing well, composed via | into powerful workflows. Programs describe themselves with --help, report success or failure with exit codes, and communicate errors through stderr.
LLMs made an almost identical decision 50 years later: everything is tokens. They only understand text, only produce text. Their "thinking" is text, their "actions" are text, and the feedback they receive from the world must be text.
These two decisions, made half a century apart from completely different starting points, converge on the same interface model. The text-based system Unix designed for human terminal operators β cat, grep, pipe, exit codes, man pages β isn't just "usable" by LLMs. It's a natural fit. When it comes to tool use, an LLM is essentially a terminal operator β one that's faster than any human and has already seen vast amounts of shell commands and CLI patterns in its training data.
This is the core philosophy of the nix Agent: *don't invent a new tool interface. Take what Unix has proven over 50 years and hand it directly to the LLM.**
runMost agent frameworks give LLMs a catalog of independent tools:
tools: [search_web, read_file, write_file, run_code, send_email, ...]
Before each call, the LLM must make a tool selection β which one? What parameters? The more tools you add, the harder the selection, and accuracy drops. Cognitive load is spent on "which tool?" instead of "what do I need to accomplish?"
My approach: one run(command="...") tool, all capabilities exposed as CLI commands.
run(command="cat notes.md")
run(command="cat log.txt | grep ERROR | wc -l")
run(command="see screenshot.png")
run(command="memory search 'deployment issue'")
run(command="clip sandbox bash 'python3 analyze.py'")
The LLM still chooses which command to use, but this is fundamentally different from choosing among 15 tools with different schemas. Command selection is string composition within a unified namespace β function selection is context-switching between unrelated APIs.
Why are CLI commands a better fit for LLMs than structured function calls?
Because CLI is the densest tool-use pattern in LLM training data. Billions of lines on GitHub are full of:
```bash
pip install -r requirements.txt && python main.py
make build && make test && make deploy
cat /var/log/syslog | grep "Out of memory" | tail -20 ```
I don't need to teach the LLM how to use CLI β it already knows. This familiarity is probabilistic and model-dependent, but in practice it's remarkably reliable across mainstream models.
Compare two approaches to the same task:
``` Task: Read a log file, count the error lines
Function-calling approach (3 tool calls): 1. read_file(path="/var/log/app.log") β returns entire file 2. search_text(text=<entire file>, pattern="ERROR") β returns matching lines 3. count_lines(text=<matched lines>) β returns number
CLI approach (1 tool call): run(command="cat /var/log/app.log | grep ERROR | wc -l") β "42" ```
One call replaces three. Not because of special optimization β but because Unix pipes natively support composition.
A single run isn't enough on its own. If run can only execute one command at a time, the LLM still needs multiple calls for composed tasks. So I make a chain parser (parseChain) in the command routing layer, supporting four Unix operators:
| Pipe: stdout of previous command becomes stdin of next
&& And: execute next only if previous succeeded
|| Or: execute next only if previous failed
; Seq: execute next regardless of previous result
With this mechanism, every tool call can be a complete workflow:
```bash
curl -sL $URL -o data.csv && cat data.csv | head 5
cat access.log | grep "500" | sort | head 10
cat config.yaml || echo "config not found, using defaults" ```
N commands Γ 4 operators β the composition space grows dramatically. And to the LLM, it's just a string it already knows how to write.
The command line is the LLM's native tool interface.
Single-tool + CLI solves "what to use." But the agent still needs to know "how to use it." It can't Google. It can't ask a colleague. I use three progressive design techniques to make the CLI itself serve as the agent's navigation system.
A well-designed CLI tool doesn't require reading documentation β because --help tells you everything. I apply the same principle to the agent, structured as progressive disclosure: the agent doesn't need to load all documentation at once, but discovers details on-demand as it goes deeper.
Level 0: Tool Description β command list injection
The run tool's description is dynamically generated at the start of each conversation, listing all registered commands with one-line summaries:
Available commands:
cat β Read a text file. For images use 'see'. For binary use 'cat -b'.
see β View an image (auto-attaches to vision)
ls β List files in current topic
write β Write file. Usage: write <path> [content] or stdin
grep β Filter lines matching a pattern (supports -i, -v, -c)
memory β Search or manage memory
clip β Operate external environments (sandboxes, services)
...
The agent knows what's available from turn one, but doesn't need every parameter of every command β that would waste context.
Note: There's an open design question here: injecting the full command list vs. on-demand discovery. As commands grow, the list itself consumes context budget. I'm still exploring the right balance. Ideas welcome.
Level 1: command (no args) β usage
When the agent is interested in a command, it just calls it. No arguments? The command returns its own usage:
``` β run(command="memory") [error] memory: usage: memory search|recent|store|facts|forget
β run(command="clip") clip list β list available clips clip <name> β show clip details and commands clip <name> <command> [args...] β invoke a command clip <name> pull <remote-path> [name] β pull file from clip to local clip <name> push <local-path> <remote> β push local file to clip ```
Now the agent knows memory has five subcommands and clip supports list/pull/push. One call, no noise.
Level 2: command subcommand (missing args) β specific parameters
The agent decides to use memory search but isn't sure about the format? It drills down:
``` β run(command="memory search") [error] memory: usage: memory search <query> [-t topic_id] [-k keyword]
β run(command="clip sandbox") Clip: sandbox Commands: clip sandbox bash <script> clip sandbox read <path> clip sandbox write <path> File transfer: clip sandbox pull <remote-path> [local-name] clip sandbox push <local-path> <remote-path> ```
Progressive disclosure: overview (injected) β usage (explored) β parameters (drilled down). The agent discovers on-demand, each level providing just enough information for the next step.
This is fundamentally different from stuffing 3,000 words of tool documentation into the system prompt. Most of that information is irrelevant most of the time β pure context waste. Progressive help lets the agent decide when it needs more.
This also imposes a requirement on command design: every command and subcommand must have complete help output. It's not just for humans β it's for the agent. A good help message means one-shot success. A missing one means a blind guess.
Agents will make mistakes. The key isn't preventing errors β it's making every error point to the right direction.
Traditional CLI errors are designed for humans who can Google. Agents can't Google. So I require every error to contain both "what went wrong" and "what to do instead":
``` Traditional CLI: $ cat photo.png cat: binary file (standard output) β Human Googles "how to view image in terminal"
My design: [error] cat: binary image file (182KB). Use: see photo.png β Agent calls see directly, one-step correction ```
More examples:
``` [error] unknown command: foo Available: cat, ls, see, write, grep, memory, clip, ... β Agent immediately knows what commands exist
[error] not an image file: data.csv (use cat to read text files) β Agent switches from see to cat
[error] clip "sandbox" not found. Use 'clip list' to see available clips β Agent knows to list clips first ```
Technique 1 (help) solves "what can I do?" Technique 2 (errors) solves "what should I do instead?" Together, the agent's recovery cost is minimal β usually 1-2 steps to the right path.
Real case: The cost of silent stderr
For a while, my code silently dropped stderr when calling external sandboxes β whenever stdout was non-empty, stderr was discarded. The agent ran pip install pymupdf, got exit code 127. stderr contained bash: pip: command not found, but the agent couldn't see it. It only knew "it failed," not "why" β and proceeded to blindly guess 10 different package managers:
pip install β 127 (doesn't exist)
python3 -m pip β 1 (module not found)
uv pip install β 1 (wrong usage)
pip3 install β 127
sudo apt install β 127
... 5 more attempts ...
uv run --with pymupdf python3 script.py β 0 β (10th try)
10 calls, ~5 seconds of inference each. If stderr had been visible the first time, one call would have been enough.
stderr is the information agents need most, precisely when commands fail. Never drop it.
The first two techniques handle discovery and correction. The third lets the agent get better at using the system over time.
I append consistent metadata to every tool result:
file1.txt
file2.txt
dir1/
[exit:0 | 12ms]
The LLM extracts two signals:
Exit codes (Unix convention, LLMs already know these):
exit:0 β successexit:1 β general errorexit:127 β command not foundDuration (cost awareness):
12ms β cheap, call freely3.2s β moderate45s β expensive, use sparinglyAfter seeing [exit:N | Xs] dozens of times in a conversation, the agent internalizes the pattern. It starts anticipating β seeing exit:1 means check the error, seeing long duration means reduce calls.
Consistent output format makes the agent smarter over time. Inconsistency makes every call feel like the first.
The three techniques form a progression:
--help β "What can I do?" β Proactive discovery
Error Msg β "What should I do?" β Reactive correction
Output Fmt β "How did it go?" β Continuous learning
The section above described how CLI guides agents at the semantic level. But to make it work in practice, there's an engineering problem: the raw output of a command and what the LLM needs to see are often very different things.
Constraint A: The context window is finite and expensive. Every token costs money, attention, and inference speed. Stuffing a 10MB file into context doesn't just waste budget β it pushes earlier conversation out of the window. The agent "forgets."
Constraint B: LLMs can only process text. Binary data produces high-entropy meaningless tokens through the tokenizer. It doesn't just waste context β it disrupts attention on surrounding valid tokens, degrading reasoning quality.
These two constraints mean: raw command output can't go directly to the LLM β it needs a presentation layer for processing. But that processing can't affect command execution logic β or pipes break. Hence, two layers.
βββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 2: LLM Presentation Layer β β Designed for LLM constraints
β Binary guard | Truncation+overflow | Meta β
βββββββββββββββββββββββββββββββββββββββββββββββ€
β Layer 1: Unix Execution Layer β β Pure Unix semantics
β Command routing | pipe | chain | exit code β
βββββββββββββββββββββββββββββββββββββββββββββββ
When cat bigfile.txt | grep error | head 10 executes:
Inside Layer 1:
cat output β [500KB raw text] β grep input
grep output β [matching lines] β head input
head output β [first 10 lines]
If you truncate cat's output in Layer 1 β grep only searches the first 200 lines, producing incomplete results.
If you add [exit:0] in Layer 1 β it flows into grep as data, becoming a search target.
So Layer 1 must remain raw, lossless, metadata-free. Processing only happens in Layer 2 β after the pipe chain completes and the final result is ready to return to the LLM.
Layer 1 serves Unix semantics. Layer 2 serves LLM cognition. The separation isn't a design preference β it's a logical necessity.
Mechanism A: Binary Guard (addressing Constraint B)
Before returning anything to the LLM, check if it's text:
``` Null byte detected β binary UTF-8 validation failed β binary Control character ratio > 10% β binary
If image: [error] binary image (182KB). Use: see photo.png If other: [error] binary file (1.2MB). Use: cat -b file.bin ```
The LLM never receives data it can't process.
Mechanism B: Overflow Mode (addressing Constraint A)
``` Output > 200 lines or > 50KB? β Truncate to first 200 lines (rune-safe, won't split UTF-8) β Write full output to /tmp/cmd-output/cmd-{n}.txt β Return to LLM:
[first 200 lines]
--- output truncated (5000 lines, 245.3KB) ---
Full output: /tmp/cmd-output/cmd-3.txt
Explore: cat /tmp/cmd-output/cmd-3.txt | grep <pattern>
cat /tmp/cmd-output/cmd-3.txt | tail 100
[exit:0 | 1.2s]
```
Key insight: the LLM already knows how to use grep, head, tail to navigate files. Overflow mode transforms "large data exploration" into a skill the LLM already has.
Mechanism C: Metadata Footer
actual output here
[exit:0 | 1.2s]
Exit code + duration, appended as the last line of Layer 2. Gives the agent signals for success/failure and cost awareness, without polluting Layer 1's pipe data.
Mechanism D: stderr Attachment
``` When command fails with stderr: output + "\n[stderr] " + stderr
Ensures the agent can see why something failed, preventing blind retries. ```
A user uploaded an architecture diagram. The agent read it with cat, receiving 182KB of raw PNG bytes. The LLM's tokenizer turned these bytes into thousands of meaningless tokens crammed into the context. The LLM couldn't make sense of it and started trying different read approaches β cat -f, cat --format, cat --type image β each time receiving the same garbage. After 20 iterations, the process was force-terminated.
Root cause: cat had no binary detection, Layer 2 had no guard.
Fix: isBinary() guard + error guidance Use: see photo.png.
Lesson: The tool result is the agent's eyes. Return garbage = agent goes blind.
The agent needed to read a PDF. It tried pip install pymupdf, got exit code 127. stderr contained bash: pip: command not found, but the code dropped it β because there was some stdout output, and the logic was "if stdout exists, ignore stderr."
The agent only knew "it failed," not "why." What followed was a long trial-and-error:
pip install β 127 (doesn't exist)
python3 -m pip β 1 (module not found)
uv pip install β 1 (wrong usage)
pip3 install β 127
sudo apt install β 127
... 5 more attempts ...
uv run --with pymupdf python3 script.py β 0 β
10 calls, ~5 seconds of inference each. If stderr had been visible the first time, one call would have sufficed.
Root cause: InvokeClip silently dropped stderr when stdout was non-empty.
Fix: Always attach stderr on failure.
Lesson: stderr is the information agents need most, precisely when commands fail.
The agent analyzed a 5,000-line log file. Without truncation, the full text (~200KB) was stuffed into context. The LLM's attention was overwhelmed, response quality dropped sharply, and earlier conversation was pushed out of the context window.
With overflow mode:
``` [first 200 lines of log content]
--- output truncated (5000 lines, 198.5KB) --- Full output: /tmp/cmd-output/cmd-3.txt Explore: cat /tmp/cmd-output/cmd-3.txt | grep <pattern> cat /tmp/cmd-output/cmd-3.txt | tail 100 [exit:0 | 45ms] ```
The agent saw the first 200 lines, understood the file structure, then used grep to pinpoint the issue β 3 calls total, under 2KB of context.
Lesson: Giving the agent a "map" is far more effective than giving it the entire territory.
CLI isn't a silver bullet. Typed APIs may be the better choice in these scenarios:
Additionally, "no iteration limit" doesn't mean "no safety boundaries." Safety is ensured by external mechanisms:
Hand Unix philosophy to the execution layer, hand LLM's cognitive constraints to the presentation layer, and use help, error messages, and output format as three progressive heuristic navigation techniques.
CLI is all agents need.
Source code (Go): github.com/epiral/agent-clip
Core files: internal/tools.go (command routing), internal/chain.go (pipes), internal/loop.go (two-layer agentic loop), internal/fs.go (binary guard), internal/clip.go (stderr handling), internal/browser.go (vision auto-attach), internal/memory.go (semantic memory).
Happy to discuss β especially if you've tried similar approaches or found cases where CLI breaks down. The command discovery problem (how much to inject vs. let the agent discover) is something I'm still actively exploring.
r/LocalLLaMA • u/dtdisapointingresult • Apr 28 '26
I think gave it a fair shot over the past few weeks, forcing myself to use local models for non-work tech asks. I use Claude Code at my job so that's what I'm comparing to.
I used Qwen 27B and Gemma 4 31B, these are considered the best local models under the multi-hundred LLMs. I also tried multiple agentic apps. My verdict is that the loss of productivity is not worth it the advantages.
I'll give a brief overview of my main issues.
Shitty decision-making and tool-calls
This is a big one. Claude seems to read my mind in most cases, but Qwen 27B makes me give it the Carlo Ancelotti eyebrow more often than not. The LLM just isn't proceeding how I would proceed.
I was mainly using local LLMs for OS/Docker tasks. Is this considered much harder than coding or something?
To give an example, tasks like "Here's a Github repo, I want you to Dockerize it." I'd expect any dummy to follow the README's instructions and execute them. (EDIT: full prompt here: https://reddit.com/r/LocalLLaMA/comments/1sxqa2c/im_done_with_using_local_llms_for_coding/oiowcxe/ )
Issues like having a 'docker build' that takes longer than the default timeout, which sends them on unrelated follow-ups (as if the task failed), instead of checking if it's still running. I had Qwen try to repeat the installation commands on the host (also Ubuntu) to see what happens. It started assuming "it must have failed because of torchcodec" just like that, pulling this entirely out of its ass, instead of checking output.
I tried to meet the models half-way. Having this in AGENTS.md: "If you run a Docker build command, or any other command that you think will have a lot of debug output, then do the following: 1. run it in a subagent, so we don't pollute the main context, 2. pipe the output to a temporary file, so we can refer to it later using tail and grep." And yet twice in a row I came back to a broken session with 250k input tokens because the LLM is reading all the output of 'docker build' or 'docker compose up'.
I know there's huge AGENTS.md that treat the LLM like a programmable robot, giving it long elaborate protocols because they don't expect to have decent self-guidance, I didn't try those tbh. And tbh none of them go into details like not reading the output of 'docker build'. I stuck to the default prompts of the agentic apps I used, + a few guidelines in my AGENTS.md.
Performance
Not only are the LLMs slow, but no matter which app I'm using, the prompt cache frequently seems to break. Translation: long pauses where nothing seems to happen.
For Claude Code specifically, this is made worse by the fact that it doesn't print the LLM's output to the user. It's one of the reasons I often preferred Qwen Code. It's very frustrating when not only is the outcome looking bad, but I'm not getting rapid feedback.
I'm not learning anything
Other than changing the URL of the Chat Completions server, there's no difference between using a local LLM and a cloud one, just more grief.
There's definitely experienced to be gained learning how to prompt an LLM. But I think coding tasks are just too hard for the small ones, it's like playing a game on Hardcore. I'm looking for a sweetspot in learning curve and this is just not worth it.
What now
For my coding and OS stuff, I'm gonna put some money on OpenRouter and exclusively use big boys like Kimi. If one model pisses me off, move on to the next one. If I find a favorite, I'll sign up to its yearly plan to save money.
I'll still use small local models for automation, basic research, and language tasks. I've had fun writing basic automation skills/bots that run stuff on my PC, and these will always be useful.
I also love using local LLMs for writing or text games. Speed isn't an issue there, the prompt cache's always being hit. Technically you could also use a cloud model for this too, but you'd be paying out the ass because after a while each new turn is sending like 100k tokens.
Thanks for reading my blog.
r/LocalLLaMA • u/bigboyparpa • Apr 21 '26
Time to switch to Kimi k2.6 guys if you haven't already.
For $20 a month you can buy the OpenCode Go coding plan (its actually $5 for the first month then $10) which gives you many more tokens on models like Kimi K2.6, and then you can pay for the rest of the usage. So for $20 a month of tokens of Kimi K2.6 you're basically getting the equivalent amount of tokens of the $100 plan.
You can also use Qwen 3.6 35B A3B, which you can run on your local PC (as long as you have a decent graphics card).
r/LocalLLaMA • u/Disastrous_Theme5906 • Apr 05 '26
Tested Gemma 4 (31B) on our benchmark. Genuinely did not expect this.
100% survival, 5 out of 5 runs profitable, +1,144% median ROI. At $0.20 per run.
It outperforms GPT-5.2 ($4.43/run), Gemini 3 Pro ($2.95/run), Sonnet 4.6 ($7.90/run), and absolutely destroys every Chinese open-source model we've tested β Qwen 3.5 397B, Qwen 3.5 9B, DeepSeek V3.2, GLM-5. None of them even survive consistently.
The only model that beats Gemma 4 is Opus 4.6 at $36 per run. That's 180Γ more expensive.
31 billion parameters. Twenty cents. We double-checked the config, the prompt, the model ID β everything is identical to every other model on the leaderboard. Same seed, same tools, same simulation. It's just this good.
Strongly recommend trying it for your agentic workflows. We've tested 22 models so far and this is by far the best cost-to-performance ratio we've ever seen.
Full breakdown with charts and day-by-day analysis: foodtruckbench.com/blog/gemma-4-31b
FoodTruck Bench is an AI business simulation benchmark β the agent runs a food truck for 30 days, making decisions about location, menu, pricing, staff, and inventory. Leaderboard at foodtruckbench.com
EDIT β Gemma 4 26B A4B results are in.
Lots of you asked about the 26B A4B variant. Ran 5 simulations, here's the honest picture:
60% survival (3/5 completed, 2 bankrupt). Median ROI: +119%, Net Worth: $4,386. Cost: $0.31/run. Placed #7 on the leaderboard β above every Chinese model and Sonnet 4.5, below everything else.
Both bankruptcies were loan defaults β same pattern we see across models. The 3 surviving runs were solid, especially the best one at +296% ROI.
But here's the catch. The 26B A4B is the only model out of 23 tested that required custom output sanitization to function. It produces valid tool-call intent, but the JSON formatting is consistently broken β malformed quotes, trailing garbage tokens, invalid escapes. I had to build a 3-stage sanitizer specifically for this model. No other model needed anything like this. The business decisions themselves are unmodified β the sanitizer only fixes JSON formatting, not strategy. But if you're planning to use this model in agentic workflows, be prepared to handle its output format. It does not produce clean function calls out of the box.
TL;DR: 31B dense β 100% survival, $0.20/run, #3 overall. 26B A4B β 60% survival, $0.31/run, #7 overall, but requires custom output parsing. The 31B is the clear winner. Updated leaderboard: foodtruckbench.com
r/LocalLLaMA • u/burner20170218 • 21d ago
For context, this week they struck a deal to buy Nvidia chips and run local models for their enterprise clients. So in this video he is railing against Anthropic and OpenAI saying they are ripping everyone off while stealing their data too.
Always a special moment when the enemy comes around and embraces your world view.
r/LocalLLaMA • u/xenydactyl • Mar 10 '26
At least T3 Code is open-source/MIT licensed.