MLX Model Test Report

Generated 2026-03-06 13:43 · AFM MLX Backend · v0.9.6

mlx-model-test.sh --models

Test Runs
37
Passed
34
Failed
3
Best tok/s
369.1
Fastest
mlx-community/lille-130m-instruct-8bit

Performance Ranking (by tokens/sec)

Click a row to jump to its full response below.

# Model / Config Status Temp Load (s) Tokens Gen (s) Tokens/sec Prompt
1 mlx-community/lille-130m-instruct-8bit OK 0.7 1.0 1928 5.22
369.1
Explain calculus concepts from limits through multivariable ...
2 Qwen/Qwen3-0.6B-MLX-4bit OK 0.7 1.0 2418 7.6
318.2
Explain calculus concepts from limits through multivariable ...
3 mlx-community/exaone-4.0-1.2b-4bit OK 0.7 1.0 1046 3.99
262.3
Explain calculus concepts from limits through multivariable ...
4 mlx-community/LFM2-2.6B-4bit OK 0.7 1.0 2406 10.32
233.2
Explain calculus concepts from limits through multivariable ...
5 mlx-community/LFM2-VL-3B-4bit OK 0.7 2.0 775 3.37
229.6
Explain calculus concepts from limits through multivariable ...
6 mlx-community/Ling-mini-2.0-4bit OK 0.7 3.0 3259 15.09
215.9
Explain calculus concepts from limits through multivariable ...
7 mlx-community/SmolLM3-3B-4bit OK 0.7 1.0 2887 16.45
175.6
Explain calculus concepts from limits through multivariable ...
8 mlx-community/Qwen3.5-2B-bf16 OK 0.7 2.0 2654 20.11
132.0
Explain calculus concepts from limits through multivariable ...
9 mlx-community/Qwen3-VL-4B-Instruct-4bit OK 0.7 2.0 4761 37.06
128.4
Explain calculus concepts from limits through multivariable ...
10 mlx-community/gpt-oss-20b-MXFP4-Q4 OK 0.7 4.0 2980 23.63
126.1
Explain calculus concepts from limits through multivariable ...
11 mlx-community/Qwen3.5-35B-A3B-4bit OK 0.7 2.0 4450 37.3
119.3
Explain calculus concepts from limits through multivariable ...
12 mlx-community/granite-4.0-h-tiny-4bit OK 0.7 2.0 1183 10.63
111.3
Explain calculus concepts from limits through multivariable ...
13 mlx-community/Qwen3.5-9B-MLX-4bit OK 0.7 2.0 5000 45.2
110.6
Explain calculus concepts from limits through multivariable ...
14 mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit OK 0.7 6.0 3629 34.15
106.3
Explain calculus concepts from limits through multivariable ...
15 mlx-community/Qwen3-30B-A3B-4bit OK 0.7 6.0 3639 34.41
105.8
Explain calculus concepts from limits through multivariable ...
16 mlx-community/gemma-3-4b-it-8bit OK 0.7 3.0 1643 16.53
99.4
Explain calculus concepts from limits through multivariable ...
17 mlx-community/Qwen3-VL-4B-Instruct-8bit OK 0.7 2.0 3488 36.62
95.3
Explain calculus concepts from limits through multivariable ...
18 mlx-community/gemma-3n-E2B-it-lm-4bit OK 0.7 2.0 2570 29.52
87.1
Explain calculus concepts from limits through multivariable ...
19 mlx-community/Qwen3.5-35B-A3B-8bit OK 0.7 12.0 4681 53.79
87.0
Explain calculus concepts from limits through multivariable ...
20 mlx-community/mistralai_Ministral-3-14B-Instruct-2512-MLX-MXFP4 OK 0.7 3.0 2883 39.7
72.6
Explain calculus concepts from limits through multivariable ...
21 mlx-community/Qwen3-Coder-Next-4bit OK 0.7 14.0 4428 65.36
67.8
Explain calculus concepts from limits through multivariable ...
22 mlx-community/Apertus-8B-Instruct-2509-4bit OK 0.7 3.0 46 0.7
65.7
Explain calculus concepts from limits through multivariable ...
23 mlx-community/GLM-4.7-Flash-4bit OK 0.7 6.0 3482 53.08
65.6
Explain calculus concepts from limits through multivariable ...
24 mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit OK 0.7 6.0 3530 57.35
61.5
Explain calculus concepts from limits through multivariable ...
25 mlx-community/Qwen3.5-122B-A10B-4bit OK 0.7 23.0 4302 70.84
60.7
Explain calculus concepts from limits through multivariable ...
26 mlx-community/JoyAI-LLM-Flash-4bit-DWQ OK 0.7 9.0 2701 53.99
50.0
Explain calculus concepts from limits through multivariable ...
27 mlx-community/MiniMax-M2.5-5bit OK 0.7 52.0 1799 42.24
42.6
Explain calculus concepts from limits through multivariable ...
28 mlx-community/MiniMax-M2.5-6bit OK 0.7 77.0 5000 125.98
39.7
Explain calculus concepts from limits through multivariable ...
29 mlx-community/Qwen3.5-397B-A17B-4bit OK 0.7 66.0 4491 113.99
39.4
Explain calculus concepts from limits through multivariable ...
30 mlx-community/functiongemma-270m-it-bf16 OK 0.7 2.0 42 1.32
31.9
Explain calculus concepts from limits through multivariable ...
31 mlx-community/mistralai_Devstral-Small-2-24B-Instruct-2512-MLX-8Bit OK 0.7 9.0 3613 140.37
25.7
Explain calculus concepts from limits through multivariable ...
32 mlx-community/Qwen3.5-27B-8bit OK 0.7 9.0 4622 209.46
22.1
Explain calculus concepts from limits through multivariable ...
33 mlx-community/Llama-3.3-70B-Instruct-4bit-DWQ OK 0.7 17.0 1666 105.83
15.7
Explain calculus concepts from limits through multivariable ...
34 mlx-community/Kimi-K2.5-3bit OK 0.7 155.0 157 192.85
0.8
Why is the sky blue? Be brief

Failed Runs

ModelErrorConfigLoad (s)
lmstudio-community/gemma-3n-E4B-it-MLX-4bit [- ] --.-% loading model | mem 0.19 GB | 0.2s [\ ] --.-% loading model | mem 0.25 GB | 0.4s [| ] --.-% loading model | mem 0.32 t=0.7 2
microsoft/Phi-4-reasoning-vision-15B microsoft/Phi-4-reasoning-vision-15B: Unsupported model type: phi4-siglip t=0.7 1
mlx-community/GLM-5-4bit Server process died (check /tmp/mlx-server-36.log) t=0.7 149

Full Responses

mlx-community/lille-130m-instruct-8bit 1928 tokens · 369.1 tok/s

Qwen/Qwen3-0.6B-MLX-4bit 2418 tokens · 318.2 tok/s

mlx-community/exaone-4.0-1.2b-4bit 1046 tokens · 262.3 tok/s

mlx-community/LFM2-2.6B-4bit 2406 tokens · 233.2 tok/s

mlx-community/LFM2-VL-3B-4bit 775 tokens · 229.6 tok/s

mlx-community/Ling-mini-2.0-4bit 3259 tokens · 215.9 tok/s

mlx-community/SmolLM3-3B-4bit 2887 tokens · 175.6 tok/s

mlx-community/Qwen3.5-2B-bf16 2654 tokens · 132.0 tok/s

mlx-community/Qwen3-VL-4B-Instruct-4bit 4761 tokens · 128.4 tok/s

mlx-community/gpt-oss-20b-MXFP4-Q4 2980 tokens · 126.1 tok/s

mlx-community/Qwen3.5-35B-A3B-4bit 4450 tokens · 119.3 tok/s

mlx-community/granite-4.0-h-tiny-4bit 1183 tokens · 111.3 tok/s

mlx-community/Qwen3.5-9B-MLX-4bit 5000 tokens · 110.6 tok/s

mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit 3629 tokens · 106.3 tok/s

mlx-community/Qwen3-30B-A3B-4bit 3639 tokens · 105.8 tok/s

mlx-community/gemma-3-4b-it-8bit 1643 tokens · 99.4 tok/s

mlx-community/Qwen3-VL-4B-Instruct-8bit 3488 tokens · 95.3 tok/s

mlx-community/gemma-3n-E2B-it-lm-4bit 2570 tokens · 87.1 tok/s

mlx-community/Qwen3.5-35B-A3B-8bit 4681 tokens · 87.0 tok/s

mlx-community/mistralai_Ministral-3-14B-Instruct-2512-MLX-MXFP4 2883 tokens · 72.6 tok/s

mlx-community/Qwen3-Coder-Next-4bit 4428 tokens · 67.8 tok/s

mlx-community/Apertus-8B-Instruct-2509-4bit 46 tokens · 65.7 tok/s

mlx-community/GLM-4.7-Flash-4bit 3482 tokens · 65.6 tok/s

mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit 3530 tokens · 61.5 tok/s

mlx-community/Qwen3.5-122B-A10B-4bit 4302 tokens · 60.7 tok/s

mlx-community/JoyAI-LLM-Flash-4bit-DWQ 2701 tokens · 50.0 tok/s

mlx-community/MiniMax-M2.5-5bit 1799 tokens · 42.6 tok/s

mlx-community/MiniMax-M2.5-6bit 5000 tokens · 39.7 tok/s

mlx-community/Qwen3.5-397B-A17B-4bit 4491 tokens · 39.4 tok/s

mlx-community/functiongemma-270m-it-bf16 42 tokens · 31.9 tok/s

mlx-community/mistralai_Devstral-Small-2-24B-Instruct-2512-MLX-8Bit 3613 tokens · 25.7 tok/s

mlx-community/Qwen3.5-27B-8bit 4622 tokens · 22.1 tok/s

mlx-community/Llama-3.3-70B-Instruct-4bit-DWQ 1666 tokens · 15.7 tok/s

mlx-community/Kimi-K2.5-3bit 157 tokens · 0.8 tok/s