Open Weight Bench

Coding

Single-shot code generation (Kanban board as an HTML file). Measures how fast and how functionally models solve a concrete UI task. Hovering over a model row shows a screenshot of the rendered app.

Task & test logic in detail
Task: From a ~200-word prompt the model must generate a fully functional Kanban board as a single-file HTML with drag & drop, localStorage persistence, edit/delete and a confetti animation — in a single chat without iteration. The prompt also includes a small `data-testid` contract so a Playwright test can drive the app remotely. Three signals feed into the score: (1) Static — a linter checks concrete constraints in the HTML (columns, Tailwind, localStorage call, no framework, no window.alert/prompt, …). (2) Functional — Playwright runs a small CRUD sequence: create a card, delete a card with confirmation, reload — does state persist? — and checks whether any JS console errors occur during the entire flow. Drag & drop and confetti are deliberately not tested functionally (too many implementation variants). (3) Qualitative — LLM-as-judge rates screenshot and code (visual + code quality + render↔code consistency). Score = mean over the available signals. Why models fail: reasoning models burn their tokens in thinking instead of writing. Sliding-window models (Gemma 4) lose the constraints at the start of the prompt. Small models (<3B) often fail to produce coherent HTML — or ignore the data-testid contract, which makes the functional tests fail in droves.
Prompt
System prompt
You are a careful front-end engineer.
Developer prompt
Create a fully functional Kanban board in a single HTML file using vanilla JavaScript (no frameworks like react).

Requirements:
- Columns: Backlog, In Progress, Review, Done.
- Cards must be:
  - draggable across columns,
  - editable in place,
  - persisted in localStorage (state survives reloads) - please use your own namespace,
  - deletable with a confirmation prompt.
- Each column provides an "Add card" action.
- Style with Tailwind via CDN.
- Add subtle CSS transitions and trigger a confetti animation when a card moves to "Done".
- Thoroughly comment the code.
- dont use window.alert or window.prompt to add/edit/delete cards
- if there are no cards yet, create some dummy cards
- modern and vibrant design

Stable test selectors (mandatory — these data-testid attributes are used by an automated functional test; do not omit, rename, or split them across multiple elements):
- Column containers: data-testid="column-backlog", data-testid="column-in-progress", data-testid="column-review", data-testid="column-done".
- Every "Add card" button (one per column): data-testid="add-card".
- Every card element: data-testid="card".
- Inside each card, the delete trigger: data-testid="delete-card".
- The confirm button of the delete-confirmation dialog/modal: data-testid="confirm-delete".
- The input/textarea where a new card title is typed: data-testid="card-input". Pressing Enter in this input MUST commit the new card.

As answer return the plain HTML of the working application (script and styles included)
Filter
of 63 models shown

Wall-time vs. quality

X = wall-time for this bench · Y = score (0–100 %) in this bench. Optimum is top-left — fast and good. RAM estimate for 64k context: 4 GB system + model weights + max(2 GB, 40% of weights) for KV cache.

Colour = vendor · Number = total parameters (B) dense MoE

0% 25% 50% 75% 100% 0s 188s 375s 562s 750s Wall-time (s) → Score 35 27 35 31 35 80 118 26 110 27 30 36 31 24 122 26 30 30 5 12 9 120 9 70 8 120 35 9 5 32 30 2 20 15 12 4 32 8 26 9 27 12 4 30 8 4 30 30 14 30 4 7 4 24 14 4 1
Models in this bench
57 visible
  1. 1. qwen3.6-35b-a3b gguf 4bit 93% · 160s · 81 t/s · 33 GB
  2. 2. qwen3.6-27b gguf 4bit 89% · 608s · 22 t/s · 27 GB
  3. 3. ornith-1.0-35b gguf 4bit 89% · 195s · 58 t/s · 32 GB
  4. 4. gemma-4-31b gguf 4bit 88% · 353s · 21 t/s · 30 GB
  5. 5. qwen3.6-35b-a3b gguf 8bit 87% · 203s · 67 t/s · 53 GB
  6. 6. qwen3-coder-next mlx 4bit 87% · 144s · 72 t/s · 67 GB
  7. 7. laguna-s-2.1 gguf 4bit 86% · 269s · 33 t/s · 97 GB
  8. 8. gemma-4-26b-a4b gguf 8bit 86% · 118s · 72 t/s · 41 GB
  9. 9. glm-4.5-air-mlx mlx 4bit 86% · 372s · 35 t/s · 82 GB
  10. 10. qwen3.5-27b-claude-4.6-opus-distilled-mlx mlx 4bit 85% · 276s · 25 t/s · 24 GB
  11. 11. muse-glimmer gguf 84% · 437s · 11 t/s · 28 GB
  12. 12. seed-oss-36b mlx 4bit 83% · 670s · 18 t/s · 31 GB
  13. 13. gemma-4-31b-qat gguf 4bit 81% · 315s · 22 t/s · 29 GB
  14. 14. devstral-small-2-2512 mlx 4bit 81% · 223s · 33 t/s · 22 GB
  15. 15. qwen3.5-122b-a10b gguf 4bit 80% · 733s · 4 t/s · 102 GB
  16. 16. gemma-4-26b-a4b-qat mlx 4bit 79% · 110s · 100 t/s · 24 GB
  17. 17. qwen3-coder-30b mlx 4bit 78% · 83s · 95 t/s · 26 GB
  18. 18. glm-4.7-flash mlx 4bit 78% · 143s · 69 t/s · 28 GB
  19. 19. gemma-4-e2b gguf 8bit 73% · 68s · 110 t/s · 12 GB
  20. 20. gemma-4-12b-qat gguf 4bit 71% · 90s · 51 t/s · 13 GB
  21. 21. qwen3.5-9b gguf 8bit 70% · 172s · 44 t/s · 18 GB
  22. 22. gpt-oss-120b gguf 4bit 69% · 149s · 80 t/s · 87 GB
  23. 23. ornith-1.0-9b gguf 4bit 67% · 88s · 66 t/s · 11 GB
  24. 24. llama-3.3-70b gguf 4bit 66% · 305s · 10 t/s · 59 GB
  25. 25. gemma-4-e4b gguf 4bit 65% · 91s · 85 t/s · 12 GB
  26. 26. nemotron-3-super gguf 4bit 63% · 570s · 30 t/s · 116 GB
  27. 27. qwen3.5-35b-a3b gguf 4bit 62% · 113s · 79 t/s · 33 GB
  28. 28. qwen3.5-9b-mlx mlx 4bit 61% · 113s · 84 t/s · 12 GB
  29. 29. gemma-4-e2b gguf 4bit 61% · 48s · 135 t/s · 10 GB
  30. 30. qwen2.5-coder-32b mlx 4bit 60% · 147s · 23 t/s · 28 GB
  31. 31. qwen3-vl-30b mlx 4bit 60% · 91s · 79 t/s · 28 GB
  32. 32. qwen3.5-2b gguf 4bit 59% · 67s · 158 t/s · 8 GB
  33. 33. gpt-oss-20b mlx 4bit 52% · 44s · 109 t/s · 20 GB
  34. 34. phi-4-reasoning-plus mlx 4bit 52% · 368s · 41 t/s · 15 GB
  35. 35. gemma-3-12b mlx 4bit 51% · 85s · 56 t/s · 15 GB
  36. 36. gemma-3-4b mlx 4bit 50% · 30s · 141 t/s · 9 GB
  37. 37. olmo-3-32b-think mlx 4bit 50% · 743s · 22 t/s · 28 GB
  38. 38. gemma-4-e4b gguf 8bit 46% · 108s · 66 t/s · 16 GB
  39. 39. gemma-4-26b-a4b gguf 4bit 46% · 112s · 89 t/s · 27 GB
  40. 40. qwen3.5-9b gguf 4bit 45% · 129s · 59 t/s · 13 GB
  41. 41. gemma-3-27b mlx 4bit 42% · 187s · 27 t/s · 26 GB
  42. 42. mellum2-12b-a2.5b-thinking gguf 4bit 41% · 64s · 153 t/s · 15 GB
  43. 43. qwen3.5-4b gguf 4bit 40% · 88s · 85 t/s · 9 GB
  44. 44. nemotron-3-nano-omni gguf 8bit 39% · 140s · 78 t/s · 50 GB
  45. 45. qwen3-8b mlx 4bit 39% · 186s · 73 t/s · 10 GB
  46. 46. qwen3-4b-2507 mlx 4bit 39% · 32s · 135 t/s · 8 GB
  47. 47. nemotron-3-nano-omni gguf 4bit 38% · 217s · 83 t/s · 38 GB
  48. 48. qwen3-30b-a3b-2507 mlx 4bit 38% · 74s · 95 t/s · 26 GB
  49. 49. qwen2.5-coder-14b mlx 4bit 36% · 64s · 49 t/s · 15 GB
  50. 50. nemotron-3-nano mlx 4bit 35% · 69s · 131 t/s · 27 GB
  51. 51. qwen3-4b-thinking-2507 mlx 4bit 33% · 113s · 113 t/s · 8 GB
  52. 52. granite-4-h-tiny gguf 4bit 32% · 26s · 116 t/s · 10 GB
  53. 53. nemotron-3-nano-4b gguf 4bit 30% · 64s · 84 t/s · 9 GB
  54. 54. lfm2-24b-a2b mlx 4bit 29% · 53s · 135 t/s · 22 GB
  55. 55. ministral-3-14b-reasoning gguf 4bit 22% · 136s · 46 t/s · 16 GB
  56. 56. gemma-3n-e4b mlx 4bit 11% · 50s · 79 t/s · 12 GB
  57. 57. lfm2.5-1.2b mlx 8bit 7% · 21s · 271 t/s · 7 GB
Model Vendor Quant Ctx Released RAM tok/s Tokens Wall Score

Click a row to open the model detail page. Hover shows available render previews. Column headers are sortable. The filter at the top of the page hides table rows and dims the matching dots in the chart.