115 points snehesht 1 hour ago 43 comments
snehesht 1 hour ago | parent
proc0 1 hour ago | parent
incognito124 1 hour ago | parent
snehesht 1 hour ago | parent
nicce 43 minutes ago | parent
mickeyp 57 minutes ago | parent
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
snehesht 44 minutes ago | parent
thatsabadlook 39 minutes ago | parent
geye1234 33 minutes ago | parent
PcChip 9 minutes ago | parent
What inference engine are you using for flash next?
thatsabadlook 41 minutes ago | parent
roscas 23 minutes ago | parent
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
esafak 1 hour ago | parent
mkl 1 hour ago | parent
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
https://github.com/Niko1221/Strata#which-model-should-i-pick
nsagent 32 minutes ago | parent
We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.merbanan 25 minutes ago | parent
quietFalcon 1 hour ago | parent
snehesht 58 minutes ago | parent
merbanan 32 minutes ago | parent
ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing
gdevenyi 1 hour ago | parent
deadbunny 53 minutes ago | parent
And I thought piping to bash was bad
snehesht 52 minutes ago | parent
gchamonlive 40 minutes ago | parent
prettyblocks 47 minutes ago | parent
hypfer 47 minutes ago | parent
The Readme doesn't say, but it's all AI generated, so..
panny 37 minutes ago | parent
MrDrMcCoy 34 minutes ago | parent
somenameforme 15 minutes ago | parent
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
MaxikCZ 5 minutes ago | parent
0xbadcafebee 35 minutes ago | parent
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
Tepix 31 minutes ago | parent
tcdent 7 minutes ago | parent
kamranjon 24 minutes ago | parent
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
ryan_glass 21 minutes ago | parent
nialv7 15 minutes ago | parent
Neywiny 11 minutes ago | parent