Home / Agents / Homelab

Praxis-Arena: 14 Local LLMs ed on the Strix HaloBenchmark

14 language models, 14 real tasks, one AMD Ryzen AI Max+ 395: This is the Praxis-Arena. Here you'll find the complete ranking — and can view every single model result directly. From bug reports to animated web apps.

Fourteen models. Fourteen tasks. One chip. The AMD Ryzen AI Max+ 395 (Strix Halo) with 96 GB VRAM isn't a server rack — it sits under my desk. But it runs models that otherwise only exist in the cloud. The question isn't whether, but which one. Here's my benchmark: no synthetic tests, but real tasks — bug hunting, research chains, architecture diagrams, interactive web apps. And the best part: you can view every single result yourself.

1.The Praxis-Arena

Most LLM benchmarks measure how well a model answers multiple-choice questions. That says little about how a model performs in actual use — when it needs to find a bug, build a research chain, or program a working web app.

The Praxis-Arena does it differently. Each model gets 14 real tasks — and must solve them independently. No hints, no revisions. One run, one result.

The task types:

  • Agent Tasks (8): Bug hunting in open-source projects, multi-step research, API integration, data archaeology, project management documents, citation networks, architecture diagrams. Scored with absolute ratings (1–10).
  • Arena Tasks (6): Interactive web apps — a penguin powerlifting animation, a sourdough bread recipe, an hourglass with physics, a supply chain game, a deep-sea scrollytelling, a weather dashboard. Scored via arena comparison.

The hardware: GMKtec EVO-X2 with AMD Ryzen AI Max+ 395 (Strix Halo), 128 GB RAM (of which 96 GB usable as VRAM). All local models run via llama.cpp or LM Studio.

The cloud references: Three cloud models (GLM 5.2, Qwen3.8 Max, Mistral Small 4) are included as reference — clearly marked as ☁️.

14Models
14Tasks
170Results
96GB VRAM

2.The Leaderboard

Sorted by average score across all completed tasks. Click a model to open the Result Explorer below.

3.Result Explorer

Pick a model and a task — and see the actual result. For HTML tasks, the generated page is rendered directly. For text tasks, you see the full report. For the Excalidraw task, the architecture diagram.

Select a model and a task above.

4.Which Model for What?

There’s no perfect model, but there are specialists. Here’s the quick overview — based on the actual arena scores.

🎨 Web Apps & HTML

Gemma 4 31B delivers the best weather dashboard (10/10), Qwen 3.6 27B is the most consistent all-rounder for HTML tasks (avg 5.9, all 6 tasks). Cloud reference: GLM 5.2 (avg 7.6).

🔍 Research & Facts

Qwen 3.6 27B and Ornith 35B share the top spot for research tasks (both avg 7.7). Detective chains, geo-localization, data archaeology — the Qwen family shines here.

📄 Long & Structured Reports

Qwen 3.6 35B is locally unbeatable for project management and citation networks (avg 8.3). Great for complex, structured documents of any kind.

🐛 Code Analysis & Bug Hunting

KAT Coder v2.5 is the specialist model for code (avg 8.0 in bug hunting + PokeJson). Closely followed by Qwen 3.6 35B and Ornith (both 7.5).

💻 Coding & Agentic Work

Ornith 35B and Qwen 3.6 35B — both A3B MoE with only 3B active parameters. That means: maximum speed at full 35B quality. Ideal for agent loops with many tool calls.

⚡ Speed & Efficiency

The entire A3B model family (Qwen 35B A3B, Ornith, KAT Coder) is fastest — only 3 billion active parameters at 35B total. Qwen 3.6 27B offers the best bang for the buck: 27B params, avg 7.0.

5.My Personal Recommendations

The following is my personal, practical assessment after months of daily use with these models. Not a benchmark result — but real experience.

After all the numbers: what do I actually use? I have two AI subscriptions (GLM/Z.AI and Ollama Cloud) that I use for most tasks. But local models have their own place — and sometimes they’re even the better choice.

#1

Ornith 1.0 35B — my daily workhorse

The model I run locally most often. Ornith is a fine-tune of the Qwen 3.6 35B A3B base and offers me the best compromise between speed and quality. The A3B MoE architecture (only 3 billion active parameters at 35B total) makes it extremely fast — fast enough for fluid work, qualitatively on par with the big models.

Exact name: ornith-1.0-35b-mtp-apex (Quality variant)

#2

Qwen 3.6 35B A3B — when I need everything

The model for tasks where no guardrails should get in the way. It offers me possibilities that I don't have with cloud models — simply because I can trust it with everything. Coding, research, creative tasks, things where other models hold back.

Exact name: qwen3.6-35b-a3b-uncensored-heretic-native-mtp-preserved-apex

Both models (Ornith and Qwen 35B) derive from the same Qwen 3.6 35B A3B base, just like KAT Coder. Choosing between them is ultimately also a matter of taste — the differences are subtle.

#3

Qwen 3.6 27B — the little wonder

27 billion parameters — and yet it keeps up with models five times its size. I'm consistently surprised by what this little model can do when compared to GLM 5.2 or other cloud models. For chatting, it's absolutely sufficient.

For coding or agentic work it's too slow for me — but that's not a quality problem, it's a speed problem. The answers are good, they just take longer.

Exact name: qwen3.6-27b-mtp

#llm#benchmark#strix-halo#amd#local#qwen#glm#arena#leaderboard#open-source

Comments

Formatting: **bold**, ==highlight==, `code`