118 points jonotime 20 hours ago 96 comments
verdverm 18 hours ago | parent
giancarlostoro 1 hour ago | parent
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).
• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
kristopolous 56 minutes ago | parent
Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.
Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.
The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...
There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.
If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.
The market is legally locked down and we're in hostage pricing mode.
And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...
Nobody is coming to save us. That's our job.
m463 37 minutes ago | parent
wonder what voting would be like?
gamer vote ++
datacenter hater vote --
datacenter lobby ++
micron lobby --
mwambua 32 minutes ago | parent
TeMPOraL 14 minutes ago | parent
phil21 34 minutes ago | parent
Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.
Plus expanding other existing facilities.
These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.
Samsung and HK Hynix also have fabs under construction and planned.
CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.
Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.
Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.
> The Micron CEO just recently said this is the exact plan
CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.
ttul 7 minutes ago | parent
Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.
BizarroLand 7 minutes ago | parent
xyzsparetimexyz 31 minutes ago | parent
kristopolous 24 minutes ago | parent
mh- 13 minutes ago | parent
(Disclosure: I work in MarTech and have access to non-public data; I wouldn't actually do this.)
ashdksnndck 27 minutes ago | parent
ByteAtATime 53 minutes ago | parent
holoduke 52 minutes ago | parent
CorrectHorseBat 44 minutes ago | parent
functionmouse 42 minutes ago | parent
1660 ti, 4790k, 16gb ddr3
petu 40 minutes ago | parent
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
liuliu 35 minutes ago | parent
anvuong 32 minutes ago | parent
nullc 30 minutes ago | parent
girvo 19 minutes ago | parent
It has a set of n-gram tables which you can stream from system RAM or even NVMe
That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?
cookiengineer 14 minutes ago | parent
I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)
I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.
Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.
apitman 12 minutes ago | parent
Because it's an open model so providers compete on price.
keammo1 8 minutes ago | parent
CamperBob2 5 minutes ago | parent
No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that. For tasks that don't require vision to support, I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better, but they are both well beyond awesome.
booi 1 hour ago | parent
ActionHank 53 minutes ago | parent
jacquesm 51 minutes ago | parent
UncleOxidant 20 minutes ago | parent
sampullman 38 minutes ago | parent
f311a 28 minutes ago | parent
tengbretson 1 hour ago | parent
smallmancontrov 1 hour ago | parent
browningstreet 1 hour ago | parent
Is OpenAI coming in $20B under a sign of "freaking out"?
efficax 1 hour ago | parent
jerf 46 minutes ago | parent
People tend to conflate the question "is AI a useful technology?" with "are the AI companies going to do well?" but they're surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon's paradox, and bearing in mind there's no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.
The markets need a very particular rate of progress. It isn't entirely clear to me that it's even a possible rate of progress, it may be overconstrained, but they certainly don't have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs doesn't work if you can't economically "kill all your competition" because the economics favor them in the spending spree.
And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.
hirako2000 26 minutes ago | parent
pessimizer 15 minutes ago | parent
This also assumes heavy utilization, though. If there's heavy utilization, it might mean they're doing well. If they're all spinning, it's time to raise prices.
anguralbanish2 1 hour ago | parent
wg0 59 minutes ago | parent
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
sampullman 39 minutes ago | parent
hirako2000 31 minutes ago | parent
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at end.
wg0 21 minutes ago | parent
simpaticoder 58 minutes ago | parent
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
agoodusername63 43 minutes ago | parent
Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.
And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king
thefourthchime 57 minutes ago | parent
Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
AIblemblio 56 minutes ago | parent
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
hmontazeri 56 minutes ago | parent
jacquesm 53 minutes ago | parent
ctolsen 45 minutes ago | parent
On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
pdhborges 28 minutes ago | parent
swiftcoder 54 minutes ago | parent
pianopatrick 48 minutes ago | parent
Would be cool if they added it.
hirako2000 29 minutes ago | parent
There are some quirks if your harness use unsupported features of course.
lmf4lol 46 minutes ago | parent
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
sroussey 44 minutes ago | parent
gsky 41 minutes ago | parent
kristianp 40 minutes ago | parent
Can't you just say "shrank to 1/437th the size"? It's not that hard.
aguilaair 39 minutes ago | parent
see https://artificialanalysis.ai/models/releases/comparisons?co...
patresh 13 minutes ago | parent
MisterMunchkin 37 minutes ago | parent
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
Octoth0rpe 17 minutes ago | parent
I do wonder how long it'll be before a us-hosted offering is available via bedrock, copilot, etc.
doctorpangloss 35 minutes ago | parent
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
distantsounds 32 minutes ago | parent
wewewedxfgdf 30 minutes ago | parent
BarryMilo 23 minutes ago | parent
elmer2 30 minutes ago | parent
RGS1811 27 minutes ago | parent
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
sergiotapia 27 minutes ago | parent
xyzsparetimexyz 26 minutes ago | parent
mlinsey 26 minutes ago | parent
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
alex-moon 24 minutes ago | parent
pessimizer 17 minutes ago | parent
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
pizza234 16 minutes ago | parent
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
p1necone 13 minutes ago | parent
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
aszen 7 minutes ago | parent