115 points snehesht 1 hour ago 43 comments

snehesht 1 hour ago | parent

I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

proc0 1 hour ago | parent

Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.

incognito124 1 hour ago | parent

Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore

snehesht 1 hour ago | parent

Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.

nicce 43 minutes ago | parent

I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.

mickeyp 57 minutes ago | parent

I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

snehesht 44 minutes ago | parent

This is interesting, thanks. - https://github.com/Neroued/ninfer

thatsabadlook 39 minutes ago | parent

Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.

geye1234 33 minutes ago | parent

I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.

PcChip 9 minutes ago | parent

Spelling mistakes?

What inference engine are you using for flash next?

thatsabadlook 41 minutes ago | parent

Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me

roscas 23 minutes ago | parent

Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

esafak 1 hour ago | parent

Has anyone calculated the effective intelligence of these quantized models?

mkl 1 hour ago | parent

There's some info in the README, including:

> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

https://github.com/Niko1221/Strata#which-model-should-i-pick

javier2 59 minutes ago | parent

ok that is getting interesting!

nicce 46 minutes ago | parent

I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?

nisarg2 19 minutes ago | parent

92% is halfway to 99%

Holds up pretty well

nsagent 32 minutes ago | parent

See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

  We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

[1]: https://arxiv.org/abs/2608.08188

merbanan 25 minutes ago | parent

I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.

quietFalcon 1 hour ago | parent

Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?

snehesht 58 minutes ago | parent

They have some community benchmarks published https://github.com/Niko1221/Strata/tree/main/bench/results

merbanan 32 minutes ago | parent

Q2_0 does 33 tok/s decode and ~600t/s prompt processing at 128k context on RTX2060 8GB VRAM.

ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing

gdevenyi 1 hour ago | parent

I had this working with the FreeToken inference engine a month ago when they launched.

https://github.com/FlashML-org/FreeToken

deadbunny 53 minutes ago | parent

> Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

And I thought piping to bash was bad

snehesht 52 minutes ago | parent

Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.

gchamonlive 40 minutes ago | parent

Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.

prettyblocks 47 minutes ago | parent

I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).

hypfer 47 minutes ago | parent

Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

The Readme doesn't say, but it's all AI generated, so..

panny 37 minutes ago | parent

I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.

MrDrMcCoy 34 minutes ago | parent

Ternary Bonsai 2 might be for you.

somenameforme 15 minutes ago | parent

The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

MaxikCZ 5 minutes ago | parent

The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it

0xbadcafebee 35 minutes ago | parent

Lol, sure, if you quant it to hell (Q2) it'll go real fast...

They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

snehesht 8 minutes ago | parent

You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.

sigbottle 6 minutes ago | parent

It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?

Tepix 31 minutes ago | parent

Q2 quantization. Not interested.

tcdent 7 minutes ago | parent

All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.

kamranjon 24 minutes ago | parent

Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

ryan_glass 21 minutes ago | parent

Anyone know how it compares to GLM 5.3 for real world use?

nialv7 15 minutes ago | parent

There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...

Neywiny 11 minutes ago | parent

I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.