27 points ronfriedhaber 1 day ago 6 comments

simonw 47 minutes ago | parent

> We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200.

If this holds up that's a really big deal.

wayfwdmachine 32 minutes ago | parent

Huge if true. As it were.

ismael_rr 23 minutes ago | parent

Super awesome. Wish they would release the paper about what they did to achieve this. I remember nous released the token superposition paper which improved pretraining FLOPs some, but not 50x: https://nousresearch.com/token-superposition. Wondering if they also found some cool tokenization strategiesa

monneyboi 17 minutes ago | parent

Imagine the sheer amount of power you could save by releasing the paper.

speedgoose 12 minutes ago | parent

But thanks to the Jevon Paradox, the global power consumption would probably increase.

https://en.wikipedia.org/wiki/Jevons_paradox

vatsachak 8 minutes ago | parent

Cool story. If it's true the company will be bought by open AI/Anthropic and Chinese labs will discover the trick and open source it by next quarter.