r/LocalLLaMA • u/perelmanych • 10h ago
Resources All currently popular local models in one table + Opus 4.8 results
If you are thinking what model will fit best your HW specs and tasks you are doing here is one table with all currently popular models that still can be considered as local.
LLM Test Scores
| Feature | DeepSeek-V4-Flash-Vision-Exp | DeepSeek-V4-Flash-0731 | Qwen3.8-Flash-Next | GLM-5.3-Flash | Qwen3.8-27B | Opus-4.8 |
|---|---|---|---|---|---|---|
| Total parameters | ≈285B | 284B | 125B | 320B | 27B | not published |
| Active parameters | 13B | 13B | 6B | 18B | 27B | not published |
Agentic benchmarks
| Benchmark | DeepSeek-V4-Flash-Vision-Exp | DeepSeek-V4-Flash-0731 | Qwen3.8-Flash-Next | GLM-5.3-Flash | Qwen3.8-27B | Opus-4.8 |
|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 83.9 | 82.7 | – | 82.6 | 73.0 | 85.0 |
| NL2Repo | 57.7 | 54.2 | 48.1 | 52.1 | 42.3 | 69.7 |
| DeepSWE | 59.3 | 54.4 | 58.7 | 61.1 | 42.2 | 58.0 |
| Toolathlon-Verified | 75.9 | 70.3 | 73.5 | 72.1 | – | 76.2 |
| Agents' Last Exam | 27.3 | 25.2⁷ | 24.3 | 28.1 | 20.4 | 25.7 |
| AutomationBench (Public) | 25.7 | 25.1 | – | 25.3 | – | 27.2 |
| GDPval-AA v2 | – | 68.1 | – | 72.3 | – | 75.1 |
| Cybergym | 75.3 | 76.7 | – | – | – | 78.3 |
| DSBench-Hard | 63.6 | 59.6 | – | – | – | 71.7 |
| DSBench-FullStack | – | 68.7 | – | – | – | 71.6 |
| ApexBench (Pass@1) | 36.5 | 26.2⁷ | – | – | – | 39.4 |
| HLE with tools (full set) | – | 16.8 | – | 22.9 | – | 25.4 |
Coding benchmarks
| Benchmark | DeepSeek-V4-Flash-Vision-Exp | DeepSeek-V4-Flash-0731 | Qwen3.8-Flash-Next | GLM-5.3-Flash | Qwen3.8-27B | Opus-4.8 |
|---|---|---|---|---|---|---|
| SWE-bench Pro | – | 56.0 | 62.5 | – | 61.7 | 69.2 |
| SWE-bench Multilingual | – | – | 81.0 | – | 73.8 | 84.4 |
| CoWorkBench | – | 45.1 | 73.9 | – | 70.7 | – |
| JobBench | – | 41.3 | 55.7 | – | 33.4 | – |
General benchmarks
| Benchmark | DeepSeek-V4-Flash-Vision-Exp | DeepSeek-V4-Flash-0731 | Qwen3.8-Flash-Next | GLM-5.3-Flash | Qwen3.8-27B | Opus-4.8 |
|---|---|---|---|---|---|---|
| GPQA Diamond | – | 90.8 | 91.7 | – | 89.2 | 93.6 |
| HLE (without tools) | – | 33.8 | 35.9 | – | 30.8 | 49.8 |
| LiveCodeBench v6 | – | 90.6 | 91.9 | – | 90.3 | – |
| IFBench | – | 79.2 | 81.3 | – | 79.5 | – |
Multimodal benchmarks
| Benchmark | DeepSeek-V4-Flash-Vision-Exp | DeepSeek-V4-Flash-0731 | Qwen3.8-Flash-Next | GLM-5.3-Flash | Qwen3.8-27B | Opus-4.8 |
|---|---|---|---|---|---|---|
| Chartography | 64.3 | – | – | – | – | 65.0 |
| ZeroBench (Pass@5) | 35.0 | – | – | – | – | 34.0 |
| BabyVision | – | – | – | 73.0 | 65.7 / 85.6 | 34.1 |
| MathVision | – | – | 90.6 / 95.7 | – | 90.0 / 94.6 | – |
| RealWorldQA | – | – | 88.5 | – | 85.9 | – |
| AndroidWorld | – | – | 84.5 | – | 81.9 | – |
| OSWorld 2.0 (partial credit) | – | – | 52.3 | – | 48.0 | – |
| Vision2Web | – | – | 64.0 | – | 62.9 | – |
| ClawEval-MM (Pass@3) | – | – | 64.4 | – | 57.4 | – |
| RecreationBench | – | – | 49.9 | – | 47.1 | – |
| ERQA | – | – | 72.3 | – | 65.5 | – |
Note: I used GLM-5.3 to compose the table from official HF pages of the models.
Note2: Opus-4.8 results are presented only for illustration and are omitted from selecting the best model in a row.
Upd: Added SWE-bench Pro, SWE-bench Multilingual, GPQA Diamond and HLE (without tools) scores for Opus 4.8 from its System Card.
4
u/leocus4 7h ago
It looks there are a bit too many missing results in these tables to do a proper comparison
1
u/perelmanych 7h ago
These all what was at HF pages of the models. As you understand I am not going to run missing tests myself. If you find somewhere additional results write here I will add them to the table. I still think there are enough results to make a comparison.
3
9h ago
[removed] — view removed comment
-1
u/perelmanych 9h ago
Totally agree, but I still find it useful. If a model's score you are interested in is in bold, then you are Ok if not you can immediately see how far it is from the best.
5
u/my_name_isnt_clever 6h ago
Q3.8FN is a monster for only 6b active, and there were so many comments dismissing it before release because of that alone. I can't wait to try the fully trained version.
3
u/perelmanych 6h ago
Yes, the model looks very good particularly because of only 6B of active parameters, but I don't understand what do you mean by fully trained version? This is a preview of their Qwen 4.0 series and as was the case with Qwen3-Next-80B-A3B there most probably won't be another better trained model on base of this one, only new Qwen-4.0 models.
1
u/my_name_isnt_clever 3h ago
That's what I mean, this is a preview of the architecture. There will be a similar size model that's proper qwen 4.
2
u/wapxmas 9h ago
Coding benchmarks
no Opus scores, that means what exactly? no coding task for opus?
0
u/perelmanych 9h ago
With bold I highlighted the best score for a bench. For this I used only local models and Opus result is there just to assess how close local models are to yesterday's SOTA model.
1
u/wapxmas 9h ago
Didnt get you. In opus column there are only dashes as scores, how SOTA's dash could be compared to llm in coding benchmark.
2
u/perelmanych 8h ago edited 8h ago
There are very few benchmark results in Opus 4.8 announcement blogpost. I had to go to Opus 4.8 System Card pdf and added for 4 additional results from there, but that is it.
1
u/EvolvingDior 3h ago
I can do 500/18 with q38f at iq4. ds4f requires iq2 on the same system and nets 200/12.
1
u/Due-Competition4564 1h ago
What context window size did you set? What was the peak memory utilisation during these runs?
19
u/reto-wyss 10h ago
I'm sticking with DSV4 Flash for now.