14 language models, 14 real tasks, one AMD Ryzen AI Max+ 395: This is the Praxis-Arena. Here you'll find the complete ranking — and can view every single model result directly. From bug reports to animated web apps.
Fourteen models. Fourteen tasks. One chip. The AMD Ryzen AI Max+ 395 (Strix Halo) with 96 GB VRAM isn't a server rack — it sits under my desk. But it runs models that otherwise only exist in the cloud. The question isn't whether, but which one. Here's my benchmark: no synthetic tests, but real tasks — bug hunting, research chains, architecture diagrams, interactive web apps. And the best part: you can view every single result yourself.
Most LLM benchmarks measure how well a model answers multiple-choice questions. That says little about how a model performs in actual use — when it needs to find a bug, build a research chain, or program a working web app.
The Praxis-Arena does it differently. Each model gets 14 real tasks — and must solve them independently. No hints, no revisions. One run, one result.
The task types:
The hardware: GMKtec EVO-X2 with AMD Ryzen AI Max+ 395 (Strix Halo), 128 GB RAM (of which 96 GB usable as VRAM). All local models run via llama.cpp or LM Studio.
The cloud references: Three cloud models (GLM 5.2, Qwen3.8 Max, Mistral Small 4) are included as reference — clearly marked as ☁️.
Sorted by average score across all completed tasks. Click a model to open the Result Explorer below.
Pick a model and a task — and see the actual result. For HTML tasks, the generated page is rendered directly. For text tasks, you see the full report. For the Excalidraw task, the architecture diagram.
Select a model and a task above.
There’s no perfect model, but there are specialists. Here’s the quick overview — based on the actual arena scores.
Gemma 4 31B delivers the best weather dashboard (10/10), Qwen 3.6 27B is the most consistent all-rounder for HTML tasks (avg 5.9, all 6 tasks). Cloud reference: GLM 5.2 (avg 7.6).
Qwen 3.6 27B and Ornith 35B share the top spot for research tasks (both avg 7.7). Detective chains, geo-localization, data archaeology — the Qwen family shines here.
Qwen 3.6 35B is locally unbeatable for project management and citation networks (avg 8.3). Great for complex, structured documents of any kind.
KAT Coder v2.5 is the specialist model for code (avg 8.0 in bug hunting + PokeJson). Closely followed by Qwen 3.6 35B and Ornith (both 7.5).
Ornith 35B and Qwen 3.6 35B — both A3B MoE with only 3B active parameters. That means: maximum speed at full 35B quality. Ideal for agent loops with many tool calls.
The entire A3B model family (Qwen 35B A3B, Ornith, KAT Coder) is fastest — only 3 billion active parameters at 35B total. Qwen 3.6 27B offers the best bang for the buck: 27B params, avg 7.0.
The following is my personal, practical assessment after months of daily use with these models. Not a benchmark result — but real experience.
After all the numbers: what do I actually use? I have two AI subscriptions (GLM/Z.AI and Ollama Cloud) that I use for most tasks. But local models have their own place — and sometimes they’re even the better choice.
The model I run locally most often. Ornith is a fine-tune of the Qwen 3.6 35B A3B base and offers me the best compromise between speed and quality. The A3B MoE architecture (only 3 billion active parameters at 35B total) makes it extremely fast — fast enough for fluid work, qualitatively on par with the big models.
Exact name: ornith-1.0-35b-mtp-apex (Quality variant)
The model for tasks where no guardrails should get in the way. It offers me possibilities that I don't have with cloud models — simply because I can trust it with everything. Coding, research, creative tasks, things where other models hold back.
Exact name: qwen3.6-35b-a3b-uncensored-heretic-native-mtp-preserved-apex
Both models (Ornith and Qwen 35B) derive from the same Qwen 3.6 35B A3B base, just like KAT Coder. Choosing between them is ultimately also a matter of taste — the differences are subtle.
27 billion parameters — and yet it keeps up with models five times its size. I'm consistently surprised by what this little model can do when compared to GLM 5.2 or other cloud models. For chatting, it's absolutely sufficient.
For coding or agentic work it's too slow for me — but that's not a quality problem, it's a speed problem. The answers are good, they just take longer.
Exact name: qwen3.6-27b-mtp
Comments