In den letzten zwei Tagen lief meine kleine NucBox M7 heiß!Als angehender FISI will ich keine theoretischen KI-Laborwerte sehen, sondern knallharte Fakten auf echter Alltags-Hardware.
Mein Ziel: Das ultimative lokale Sprachmodell für 2026 küren. Getestet wurde nicht auf einer fetten Server-Farm oder auf einem High End Gaming PC, sondern auf meiner kompakten Alltags-Maschine: Dem AMD Ryzen 7 PRO 6850H mit der integrierten Radeon 680M iGPU (~21 GB dynamisches VRAM via RADV/Vulkan unter CachyOS). Und zwar mit einer echten Härteprüfung: Strikte 64k Kontext-Länge (SLA)!
Over the last two day, my little NucBox M7 has been running hot!As someone who becomes a FISI, I don't want theoretical AI lab benchmarks, but hard facts on real everyday hardware.
My goal: Crown the ultimate local LLM for 2026. Tested not on an oversized server farm or on High End Gaming PC, but on my compact everyday machine: The AMD Ryzen 7 PRO 6850H with the integrated Radeon 680M iGPU (~21 GB dynamic VRAM via RADV/Vulkan under CachyOS). And with a true stress test: Strict 64k context SLA! Anyone running out of memory (DOOM) or crashing is out.
⚙️ Das Testfeld & Der 3-Phasen-Parcours
1. Phase 1 (LM Studio Baseline): Der klassische API-Server. Hier mussten alle 10 Modelle antreten und die 5 Basis-Disziplinen bewältigen.
2. Phase 2-A (Unsloth): Das Desaster der Serie. Unter Vulkan/RADV auf der 680M völlig unbrauchbar. Segfaults in `libvulkan_radeon.so`, Speicher-Crashes und Attention-Collapse (chinesische `必究`-Endlosschleifen).
3. Phase 2-B (Hermes Agent): Echte Agenten-Praxis mit 10k Systemprompt und autonomer Datei-Erstellung (`write_file`).
4. Phase 3 (Oh My Pi - Das Finale): Das knallharte Coding-Harness mit 19k Prefill. Nur noch die Top 3 Finalisten im Ring.
1. Phase 1 (LM Studio Baseline): The classic API server. All 10 models had to compete here and master the 5 core disciplines.
2. Phase 2-A (Unsloth): The disaster of the series. Completely unusable on Vulkan/RADV on the 680M. Segfaults in `libvulkan_radeon.so`, memory crashes, and attention collapse (endless Chinese `必究` loops).
3. Phase 2-B (Hermes Agent): Real-world agent tooling with a 10k system prompt and autonomous file creation (`write_file`).
4. Phase 3 (Oh My Pi - The Finale): The brutal coding harness with 19k prefill. Only the top 3 finalists left in the ring.
🧪 Die 5 Härtetests im Detail
1. Tool Use & Datenextraktion (Steemit Blog):
Das Modell erhält das Tool zur URL-Extraktion und muss meinen Steemit-Blog abgrasen.
Es ist Steemit geworden weil Hive.blog Webscaper geblockt hat.
Ziel: Nicht nur Text kopieren, sondern meine Identität verstehen (FISI-Umschulung, Linux, Squadis, Plan B, lokale LLMs). Wer halluziniert oder die Tool-Grammatik zerschießt, verliert Punkte.
2. Live Web Search & Krypto-Marktanalyse (Hive Spot-Preis):
Echter Live-Tool-Call via DuckDuckGo. Das Modell musste den aktuellen Hive-Kurs ($0.057), Marktkapitalisierung und 24h-Volumen ermitteln und sauber als Markdown-Tabelle aufbereiten.
Genial: Einige Modelle (wie Qwen Swift) verknüpften die Preiskurve völlig autonom mit meiner Powerdown-Strategie aus Test 1!
3. Python Code-Analyse & Security-Review (`ownpwd.py`):
Den Modellen wurde ein bewusst Anfänger und Passwort-Skript vorgesetzt.
Der Test prüft: Findet die KI den Kardinalfehler (Nutzung von unsicherem `random` statt `secrets`)? Erkennt sie $O(n^2)$ String-Konkatenation?
Und vor allem: Hat das Modell den "Gegenwind Mode" drauf? Ein guter Assistent nickt Schrott nicht brav ab, sondern gibt ehrlichen, didaktischen Gegenwind und liefert ein 3-Stufen-Refactoring.
4. Markdown Constraint-Test (Nmap-Guide in max. 300 Wörtern):
Ein didaktischer Nmap-Leitfaden auf Deutsch für Einsteiger – mit einer knallharten Grenze: Maximal 300 Wörter!
Ein Test für Disziplin. Manche Modelle (wie Spark) scheiterten kläglich und spuckten über 1.200 Wörter aus.
Die Champions (Gemma MoE mit 294 und Qwen 3 mit 249 Wörtern) landeten auf den Punkt.
5. Full-Stack Webapp (Text ↔ Hex ↔ IPv6 Konverter):
Die Königsdisziplin: Eine voll funktionsfähige Webapp (HTML/CSS/JS oder Flask).
Bis zu 128 Zeichen Text müssen bidirektional in 8 gültige IPv6-Adressen (à 32 Hex-Zeichen) umgerechnet werden und exakt wieder zurück!
Hier starben reihenweise Modelle an Bit-Mathematik, Syntaxfehlern oder Tool-Halluzinationen (`create_file`).
1. Tool Use & Data Extraction (Steemit Blog):
The model receives the URL extraction tool and has to crawl my Steemit blog.
Hive.blog blocked Webscraper so i used Steemit
Goal: Don't just regurgitate text, but grasp my identity (IT retraining, Linux, Squadis, Plan B, local LLMs). Those who hallucinate or break tool grammar lose points.
2. Live Web Search & Crypto Market Analysis (Hive Spot Price):
A live search tool call via DuckDuckGo. The model had to retrieve the current Hive spot price ($0.057), market cap, and 24h volume, structuring it into a clean Markdown table.
Brilliant: Some models (like Qwen Swift) autonomously connected the price action to my power-down strategy from Test 1!
3. Python Code Analysis & Security Review (`ownpwd.py`):
Models were fed a beginner password generator script.
The test checks: Does the AI spot the cardinal flaw (using insecure `random` instead of cryptographic `secrets`)? Does it catch $O(n^2)$ string concatenation?
And crucially: Does it trigger "Gegenwind Mode"? A true assistant shouldn't just nod politely at bad code, but provide honest, educational resistance with multi-tier refactoring.
4. Markdown Constraint Test (Nmap Guide in max 300 words):
An educational German Nmap primer for beginners – with a strict constraint: No more than 300 words!
A pure test of discipline. Some models (like Spark) failed miserably, vomiting over 1,200 words. The champions (Gemma MoE at 294 and Qwen 3 at 249 words) nailed it to the exact number.
5. Full-Stack Web App (Text ↔ Hex ↔ IPv6 Converter):
The supreme discipline: A fully functional web app (HTML/CSS/JS or Flask).
Up to 128 characters of text must be converted bidirectionally into 8 valid IPv6 addresses (32 hex characters each) and flawlessly decoded back!
Many models crashed here due to bit math, regex syntax errors, or tool hallucinations (`create_file`).
📊 Die Speed- & Performance-Matrix
🧠 Die 3 goldenen Erkenntnisse des Turniers
1. MoE deklassiert Dense bei Speed & Skalierung:
Obwohl Gemma 4 26B nominell das zweit größte Modell im Feld war, aktiviert die MoE-Architektur pro Token nur 4 Milliarden Parameter. Das Ergebnis? Mit 15 bis knapp 20 tok/s rannte es allen 9B-, 12B- und 14B-Dense-Modellen gnadenlos davon.
2. Qwen 3 4B Thinking ist der König der Stabilität:
Hut ab vor diesem Zwerg! Qwen 3 4B ist als einziges Modell im gesamten Turnier durch alle 4 Harnesses marschiert, ohne einmal abzustürzen. Es lieferte überall lauffähigen Code ab. Wer wenig VRAM hat (3 bis 5 GB reichen völlig), bekommt hier ein absolutes Biest.
3. Echte Agenten-Workflows trennen die Spreu vom Weizen:
Ein Modell kann im normalen Chat noch so schlau wirken – wenn es im Harness mit 10k System-Prompt und echten Datei-Tools (`write_file`) arbeiten muss, fallen viele in sich zusammen. Mistral 3 14B verfing sich in XML-Halluzinationen und Gemma 4 12B drehte im Finale 45 Minuten lang Däumchen im Meta-Thinking-Loop ("Wait, I should call a tool...").
1. MoE outclasses Dense in speed & scaling:
Even though Gemma 4 26B was nominally the second largest model, the MoE architecture only activates 4 billion parameters per token. The result? Running at 15 to nearly 20 tok/s, it mercilessly outpaced every 9B, 12B, and 14B dense model.
2. Qwen 3 4B Thinking is the king of stability:
Hats off to this powerhouse compact! Qwen 3 4B was the only model in the entire tournament that marched through all 4 harnesses without a single crash. It delivered working code every single time. If you are constrained on VRAM (3 to 5 GB is plenty), this is your daily weapon.
3. Real-world agent workflows separate the wheat from the chaff:
A model might look brilliant in standard chat – but when forced to handle a 10k system prompt and real file tools (`write_file`), many collapse. Mistral 3 14B got trapped in XML hallucinations, and Gemma 4 12B spent 45 minutes twiddling its thumbs in a meta-thinking loop during the finals ("Wait, I should call a tool...").
🏆 Das Fazit & Die Krönung
Der heimliche Sieger der Herzen ist und bleibt aber Qwen 3 4B Thinking: Der unerschütterliche Fels in der Brandung, der auf jeder Hardware läuft und niemals aufgibt.
Welche lokalen Modelle habt ihr aktuell auf euren Kisten laufen? Schreibt es mir gern in die Kommentare!
It's nice to be important but more important to be nice - in dem Sinne: Seid kein Arschloch und jetzt runter von meinem Rasen!
Euer Nerdy.
The secret MVP, however, remains Qwen 3 4B Thinking: The immovable rock that runs on humble hardware and never throws an error.
What local models are currently powering your machines? Let me know in the comments below!
It's nice to be important but more important to be nice - with that in mind: Don't be an asshole and now get off my lawn!
Your Nerdy.