MLX Model Test Report

Generated 2026-02-26 21:22 · AFM MLX Backend · v0.9.5

mlx-model-test.sh --models

Test Runs
30
Passed
28
Failed
2
Best tok/s
426.3
Fastest
mlx-community/lille-130m-instruct-8bit

Performance Ranking (by tokens/sec)

Click a row to jump to its full response below.

# Model / Config Status Temp Load (s) Tokens Gen (s) Tokens/sec Prompt
1 mlx-community/lille-130m-instruct-8bit OK 0.7 1.0 769 1.8
426.3
Explain calculus concepts from limits through multivariable ...
2 mlx-community/exaone-4.0-1.2b-4bit OK 0.7 1.0 1070 3.82
280.3
Explain calculus concepts from limits through multivariable ...
3 mlx-community/LFM2-2.6B-4bit OK 0.7 1.0 2319 9.84
235.7
Explain calculus concepts from limits through multivariable ...
4 mlx-community/LFM2-VL-3B-4bit OK 0.7 1.0 880 3.83
229.6
Explain calculus concepts from limits through multivariable ...
5 mlx-community/Ling-mini-2.0-4bit OK 0.7 3.0 3157 14.38
219.6
Explain calculus concepts from limits through multivariable ...
6 mlx-community/SmolLM3-3B-4bit OK 0.7 1.0 1778 10.01
177.7
Explain calculus concepts from limits through multivariable ...
7 mlx-community/granite-4.0-h-tiny-4bit OK 0.7 2.0 1399 9.64
145.2
Explain calculus concepts from limits through multivariable ...
8 mlx-community/gpt-oss-20b-MXFP4-Q4 OK 0.7 4.0 3525 26.82
131.4
Explain calculus concepts from limits through multivariable ...
9 mlx-community/Qwen3-VL-4B-Instruct-4bit OK 0.7 2.0 4717 36.31
129.9
Explain calculus concepts from limits through multivariable ...
10 mlx-community/Qwen3.5-35B-A3B-4bit OK 0.7 1.0 4449 36.86
120.7
Explain calculus concepts from limits through multivariable ...
11 mlx-community/Qwen3-30B-A3B-4bit OK 0.7 6.0 4089 38.6
105.9
Explain calculus concepts from limits through multivariable ...
12 mlx-community/Qwen3-VL-4B-Instruct-8bit OK 0.7 2.0 3755 39.57
94.9
Explain calculus concepts from limits through multivariable ...
13 mlx-community/gemma-3n-E2B-it-lm-4bit OK 0.7 2.0 2339 25.86
90.4
Explain calculus concepts from limits through multivariable ...
14 mlx-community/Qwen3.5-35B-A3B-8bit OK 0.7 11.0 4825 60.6
79.6
Explain calculus concepts from limits through multivariable ...
15 mlx-community/Apertus-8B-Instruct-2509-4bit OK 0.7 2.0 43 0.58
74.7
Explain calculus concepts from limits through multivariable ...
16 mlx-community/mistralai_Ministral-3-14B-Instruct-2512-MLX-MXFP4 OK 0.7 3.0 4411 61.15
72.1
Explain calculus concepts from limits through multivariable ...
17 mlx-community/functiongemma-270m-it-bf16 OK 0.7 2.0 27 0.39
69.8
Explain calculus concepts from limits through multivariable ...
18 mlx-community/Qwen3-Coder-Next-4bit OK 0.7 14.0 5000 73.03
68.5
Explain calculus concepts from limits through multivariable ...
19 mlx-community/GLM-4.7-Flash-4bit OK 0.7 6.0 3623 53.43
67.8
Explain calculus concepts from limits through multivariable ...
20 mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit OK 0.7 6.0 3476 55.48
62.6
Explain calculus concepts from limits through multivariable ...
21 mlx-community/JoyAI-LLM-Flash-4bit-DWQ OK 0.7 9.0 1943 35.7
54.4
Explain calculus concepts from limits through multivariable ...
22 mlx-community/MiniMax-M2.5-5bit OK 0.7 56.0 3531 81.16
43.5
Explain calculus concepts from limits through multivariable ...
23 mlx-community/MiniMax-M2.5-6bit OK 0.7 66.0 5000 125.65
39.8
Explain calculus concepts from limits through multivariable ...
24 mlx-community/Qwen3.5-397B-A17B-4bit OK 0.7 66.0 4995 125.69
39.7
Explain calculus concepts from limits through multivariable ...
25 mlx-community/mistralai_Devstral-Small-2-24B-Instruct-2512-MLX-8Bit OK 0.7 8.0 3366 130.7
25.8
Explain calculus concepts from limits through multivariable ...
26 mlx-community/Llama-3.3-70B-Instruct-4bit-DWQ OK 0.7 13.0 1726 113.05
15.3
Explain calculus concepts from limits through multivariable ...
27 mlx-community/GLM-5-4bit OK 0.7 125.0 4676 376.56
12.4
Explain calculus concepts from limits through multivariable ...
28 mlx-community/Kimi-K2.5-3bit OK 0.7 142.0 279 112.73
2.5
Why is the sky blue? Be brief

Failed Runs

ModelErrorConfigLoad (s)
mlx-community/gemma-3-4b-it-8bit Connection error. t=0.7 2.0
lmstudio-community/gemma-3n-E4B-it-MLX-4bit [- ] --.-% loading model | mem 0.19 GB | 0.2s [\ ] --.-% loading model | mem 0.25 GB | 0.4s [| ] --.-% loading model | mem 0.32 t=0.7 2

Full Responses

mlx-community/lille-130m-instruct-8bit 769 tokens · 426.3 tok/s

mlx-community/exaone-4.0-1.2b-4bit 1070 tokens · 280.3 tok/s

mlx-community/LFM2-2.6B-4bit 2319 tokens · 235.7 tok/s

mlx-community/LFM2-VL-3B-4bit 880 tokens · 229.6 tok/s

mlx-community/Ling-mini-2.0-4bit 3157 tokens · 219.6 tok/s

mlx-community/SmolLM3-3B-4bit 1778 tokens · 177.7 tok/s

mlx-community/granite-4.0-h-tiny-4bit 1399 tokens · 145.2 tok/s

mlx-community/gpt-oss-20b-MXFP4-Q4 3525 tokens · 131.4 tok/s

mlx-community/Qwen3-VL-4B-Instruct-4bit 4717 tokens · 129.9 tok/s

mlx-community/Qwen3.5-35B-A3B-4bit 4449 tokens · 120.7 tok/s

mlx-community/Qwen3-30B-A3B-4bit 4089 tokens · 105.9 tok/s

mlx-community/Qwen3-VL-4B-Instruct-8bit 3755 tokens · 94.9 tok/s

mlx-community/gemma-3n-E2B-it-lm-4bit 2339 tokens · 90.4 tok/s

mlx-community/Qwen3.5-35B-A3B-8bit 4825 tokens · 79.6 tok/s

mlx-community/Apertus-8B-Instruct-2509-4bit 43 tokens · 74.7 tok/s

mlx-community/mistralai_Ministral-3-14B-Instruct-2512-MLX-MXFP4 4411 tokens · 72.1 tok/s

mlx-community/functiongemma-270m-it-bf16 27 tokens · 69.8 tok/s

mlx-community/Qwen3-Coder-Next-4bit 5000 tokens · 68.5 tok/s

mlx-community/GLM-4.7-Flash-4bit 3623 tokens · 67.8 tok/s

mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit 3476 tokens · 62.6 tok/s

mlx-community/JoyAI-LLM-Flash-4bit-DWQ 1943 tokens · 54.4 tok/s

mlx-community/MiniMax-M2.5-5bit 3531 tokens · 43.5 tok/s

mlx-community/MiniMax-M2.5-6bit 5000 tokens · 39.8 tok/s

mlx-community/Qwen3.5-397B-A17B-4bit 4995 tokens · 39.7 tok/s

mlx-community/mistralai_Devstral-Small-2-24B-Instruct-2512-MLX-8Bit 3366 tokens · 25.8 tok/s

mlx-community/Llama-3.3-70B-Instruct-4bit-DWQ 1726 tokens · 15.3 tok/s

mlx-community/GLM-5-4bit 4676 tokens · 12.4 tok/s

mlx-community/Kimi-K2.5-3bit 279 tokens · 2.5 tok/s