147 points volf_ 50 minutes ago 47 comments
algoth1 40 minutes ago | parent
vatsachak 37 minutes ago | parent
Some features of the release I like:
- Demonstration of diverse tasks, such as using a DAW
- Graphs from various benchmarks and price ranges
- Real world use of the model in scientific environments
unpopularopp 33 minutes ago | parent
platinumrad 30 minutes ago | parent
verdverm 11 minutes ago | parent
Curious if Verizon / ATT still force apps on your phone, eg. NFL and Amazon apps, Fi service is subpar
InsideOutSanta 15 minutes ago | parent
I'm sure they're doing all kinds of terrible things, like all major companies. I just can't help but like them. Also, this model looks great, and I'll give their subscription a shot next month.
algoth1 14 minutes ago | parent
DanMcInerney 33 minutes ago | parent
ddxv 33 minutes ago | parent
omani 33 minutes ago | parent
but now I got my "proof".
nemothekid 32 minutes ago | parent
sandblast 24 minutes ago | parent
danvayn 20 minutes ago | parent
rao-v 30 minutes ago | parent
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
stymaar 26 minutes ago | parent
Pro [2]:, 1.02T total / 42B activated parameters
verdverm 10 minutes ago | parent
syntaxing 25 minutes ago | parent
brcmthrowaway 22 minutes ago | parent
stymaar 17 minutes ago | parent
[1] https://sebastianraschka.com/llm-architecture-gallery/per-la...
[2]: See DS 4.1-Flash and Qwen-3.8-Next.
verdverm 9 minutes ago | parent
bertili 20 minutes ago | parent
spwa4 17 minutes ago | parent
MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks
Perhaps with IQ2 flash will run on 128G M5?
MisterMunchkin 15 minutes ago | parent
Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.
jwpapi 8 minutes ago | parent
It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.
I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.
hashmush 3 minutes ago | parent
lwansbrough 7 minutes ago | parent
user43928 6 minutes ago | parent
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
GPT 6 Astra 59.6 Claude Fable 5.1 55.1 Claude Opus 5 49.0 MiMo-V2.6-Pro 34.9 MiMo-V2.6-Flash 28.8 DeepSeek V4.1 Flash 26.8 MiMo-V2.5-Pro 1.5
ExploitGym
GPT 6 Astra 42.4 Claude Fable 5.1 30.4 Claude Opus 5 22.1 MiMo-V2.6-Pro 17.8 MiMo-V2.6-Flash 6.0 MiMo-V2.5-Pro 0.1
DeepSWE v1.1
DeepSeek V4.1 Flash 74.2 Claude Opus 5 74.0 GPT 6 Astra 74.0 MiMo-V2.6-Pro 71.9 Claude Fable 5 70.0 MiMo-V2.6-Flash 67.9 MiMo-V2.5-Pro
mokre 4 minutes ago | parent