134 points fsbonetto 2 hours ago 117 comments
fsbonetto 2 hours ago | parent
vatsachak 2 hours ago | parent
skybrian 1 hour ago | parent
fsbonetto 1 hour ago | parent
For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.
pcarolan 1 hour ago | parent
hehimself 1 hour ago | parent
traverseda 1 hour ago | parent
Also can't keep them closed source if you do that.
skeskinen 1 hour ago | parent
Also, it's hard to get fab capacity for any project. Let alone something so experimental.
zitterbewegung 1 hour ago | parent
ohazi 1 hour ago | parent
slowin 1 hour ago | parent
yorwba 1 hour ago | parent
Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.
birdatlaw 1 hour ago | parent
But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.
I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.
I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.
I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.
zdragnar 1 hour ago | parent
It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.
I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware
fhdkweig 1 hour ago | parent
fsbonetto 1 hour ago | parent
LoganDark 1 hour ago | parent
2. They don't have enough capacity either
The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).
You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.
CamperBob2 34 minutes ago | parent
It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.
monocasa 1 hour ago | parent
zdragnar 1 hour ago | parent
Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.
If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?
There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.
My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.
jerf 1 hour ago | parent
rjh29 1 hour ago | parent
Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)
exmadscientist 55 minutes ago | parent
You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing.
(Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)
dist1ll 15 minutes ago | parent
It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs a couple hundred bucks at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.
LoganDark 53 minutes ago | parent
That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.
HoldOnAMinute 57 minutes ago | parent
All aboard! We're racing to the bottom now.
sanderjd 51 minutes ago | parent
bluefirebrand 47 minutes ago | parent
I don't want to live at the bottom
sanderjd 43 minutes ago | parent
bluefirebrand 19 minutes ago | parent
sanderjd 4 minutes ago | parent
jb1991 53 minutes ago | parent
zdragnar 48 minutes ago | parent
_puk 34 minutes ago | parent
Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.
zdragnar 22 minutes ago | parent
fsbonetto 1 hour ago | parent
So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.
fnordpiglet 1 hour ago | parent
The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.
I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.
Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.
But such is the cycle
cestith 1 hour ago | parent
warkdarrior 14 minutes ago | parent
Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.
schleck8 1 hour ago | parent
pmarreck 1 hour ago | parent
fsbonetto 1 hour ago | parent
dmitrygr 1 hour ago | parent
Much of a model are weights, and high-density ROMs are very very very hard.
__MatrixMan__ 1 hour ago | parent
sanderjd 46 minutes ago | parent
jolt42 1 hour ago | parent
fsbonetto 1 hour ago | parent
It's upcoming second generation could run the inference of the models that are being used to improve it...
samuelknight 1 hour ago | parent
AIblemblio 1 hour ago | parent
Your optimized hardware chip might be obsolete before its back from the fab.
SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.
Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.
buriram 1 hour ago | parent
jjcm 53 minutes ago | parent
I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.
sanderjd 49 minutes ago | parent
bluGill 20 minutes ago | parent
sanderjd 7 minutes ago | parent
But I think it's a good bet. I think that in two years, if I can get opus/sonnet 5.5 or the gpt-6 models for much cheaper and faster than whatever the "frontier" is at that point, that this will probably be a great trade for most of my work. I certainly don't know that for sure, that's why it's a bet, but it's what I think right now.
I wouldn't quite say that about any of the open weight models at this point. But I'm hopeful that will change in the next generation or two of those models.
casta 6 minutes ago | parent
root_axis 5 minutes ago | parent
rfgplk 1 hour ago | parent
jetemple 1 hour ago | parent
xg15 1 hour ago | parent
Also: Here is our recursive self-improvement hard at work...
lelanthran 1 hour ago | parent
> Also: Here is our recursive self-improvement hard at work...
Soon we will see
token-providers: "The torment nexus is a cautionary tale"
Also token-providers: "Finally, we have created the torment nexus that we first told you about!"
nialse 1 hour ago | parent
mrob 1 hour ago | parent
Jtsummers 1 hour ago | parent
What's your basis for thinking ASI will kill all biological life, and how do you think it's going to happen?
mrob 36 minutes ago | parent
I think it's likely to do that because any unbounded goal that doesn't explicitly protect biological life (and we have no idea how to actually define such a stipulation) is best solved by killing all biological life. This is an obvious consequence of unbounded goals consuming unbounded resources, conflicting with biological life needing resources to sustain itself.
>how do you think it's going to happen?
I can speculate (e.g. we're nowhere close to the maximum killing power of drones), but I don't know because I only have human intelligence. An ASI is by definition smarter than me and surely capable of coming up with better ideas. But I do know that it's not going to do anything that would make a good sci-fi plot, because those always give the humans a chance to win, which would be stupid. Everything will seem to be going great and then everybody suddenly and unexpectedly dies.
Jtsummers 34 minutes ago | parent
What resources are unbounded? There are limits to growth in the real world, how are these ASIs going to escape physical reality?
mrob 22 minutes ago | parent
Jtsummers 3 minutes ago | parent
athrowaway3z 1 hour ago | parent
The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.
But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.
felixgallo 1 hour ago | parent
fsbonetto 1 hour ago | parent
chris_money202 1 hour ago | parent
Companies typically combined multiple platforms together such as HAPs, Zebu, Palladium, fleets of FPGAs, and Virtual Platforms in order to design and verify ASICS. So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design.
bitwize 1 hour ago | parent
rcarmo 56 minutes ago | parent
ASalazarMX 17 minutes ago | parent
It is still a very good read, but the machines are much better written than the people.
srameshc 1 hour ago | parent
amelius 1 hour ago | parent
ASalazarMX 13 minutes ago | parent
The elite gurus will get paid handsomely, while promptgrammers will be paid less since they've become a less-skilled commodity, and the company has to pay for the expensive tokens they'll avidly consume.
I've seen someone jump from Wordpress to deploying internet-facing APIs because 'they have PHP experience', and the holes in their knowledge were filled blindly by an LLM. I have also argued with a seasoned developer about how their code didn't need linting because LLMs 'already follow best practices'.
The future doesn't look bright when LLMs allow future generations to feign required knowledge.
AnimalMuppet 1 hour ago | parent
chris_money202 40 minutes ago | parent
In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU.
So its missing ALOT
gfalcao 57 minutes ago | parent
rcarmo 57 minutes ago | parent
QuantumNomad_ 54 minutes ago | parent
figassis 45 minutes ago | parent
_diyar 39 minutes ago | parent
altmanaltman 37 minutes ago | parent
mbgerring 42 minutes ago | parent
No, it isn't.
A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.
holmesworcester 35 minutes ago | parent
Are we confident that no existing LLM is capable of similarly effective prompts to those this author used? (I agree it's a stretch, but would not reject it out of hand.)
Even if not yet, will the existence of this repo soon change that, because LLMs will soon ingest it?
dpoloncsak 23 minutes ago | parent
I think OP is trying to convey the idea that LLMs do not take initiative to do anything, and these are not 'beings' capable of doing things. These are tools being used by humans.
fragmede 19 minutes ago | parent
mitxela 9 minutes ago | parent
complex_fir_rea 2 minutes ago | parent
besterman23 35 minutes ago | parent
Like sure it didn’t have the inclination to make the sim and hardware designs, but it did make them though yes?
sailingparrot 26 minutes ago | parent
Getting an LLM to design something in its own simulator that is not accurate w.r.t reality is not useful nor terribly impressive.
besterman23 16 minutes ago | parent
jmoggr 30 minutes ago | parent
Getting LLMs to prompt other LLMs in a loop is not hard, it doesn't produce great results most of the time, but that is changing.
knicholes 30 minutes ago | parent
scarmig 24 minutes ago | parent
It's not clear what this fad of attributing everything an AI does to the human prompting it is supposed to accomplish.
mbgerring 21 minutes ago | parent
It's meant to assign agency and accountability where it actually lies instead of mystifying it with anthropomorphic language.
Failing to do so has real and harmful consequences, such as enabling OpenAI to escape accountability for clearly criminal behavior.
ben_w 16 minutes ago | parent
Questions of agency are for lawyers, questions of personhood for philosophers, we're engineers and our question is capability.
Does it really have the capability? By default I'm sceptical for the same reasons given by sailingparrot: https://news.ycombinator.com/item?id=49982068
Tanjreeve 6 minutes ago | parent
> Questions of agency are for lawyers, questions of personhood for philosophers, we're engineers and our question is capability.
But it's objectively not capable without a human telling specifying things. Same as an oven can't cook a three course meal without a chef. We get around that with training data but there will always be things with no/less data or outdated knowledge.
jibalt 2 minutes ago | parent
jibalt 4 minutes ago | parent
deepsun 35 minutes ago | parent
random__duck 31 minutes ago | parent
fsbonetto 9 minutes ago | parent
mitxela 6 minutes ago | parent
jonahss 2 minutes ago | parent