455 points bradleyg223 58 minutes ago 235 comments
babelfish 55 minutes ago | parent
Gemini not beating the "can't release a model" allegations
modeless 52 minutes ago | parent
ionwake 39 minutes ago | parent
ok bro thx
Androider 37 minutes ago | parent
vlyan 31 minutes ago | parent
AuthAuth 19 minutes ago | parent
XzAeRosho 19 minutes ago | parent
bakugo 34 minutes ago | parent
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
A_D_E_P_T 19 minutes ago | parent
jstummbillig 14 minutes ago | parent
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
kevinh 2 minutes ago | parent
iamronaldo 54 minutes ago | parent
bottlepalm 53 minutes ago | parent
colordrops 51 minutes ago | parent
Scrapemist 44 minutes ago | parent
fer 29 minutes ago | parent
NiloCK 36 minutes ago | parent
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
bottlepalm 34 minutes ago | parent
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
yacthing 4 minutes ago | parent
rsstack 48 minutes ago | parent
Rzor 14 minutes ago | parent
polotics 45 minutes ago | parent
eamsen 42 minutes ago | parent
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
mattkevan 25 minutes ago | parent
Many other models get things wrong, but Gemini is the only one to go on the defensive.
RachelF 27 minutes ago | parent
Hamuko 21 minutes ago | parent
abixb 4 minutes ago | parent
SwellJoe 53 minutes ago | parent
jastanton 46 minutes ago | parent
wasting_time 7 minutes ago | parent
SwellJoe 4 minutes ago | parent
https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...
It means I am saying something that is not very believable.
TeMPOraL 3 minutes ago | parent
blueaquilae 40 minutes ago | parent
hn_acc1 38 minutes ago | parent
thefourthchime 16 minutes ago | parent
kccqzy 53 minutes ago | parent
pliiight 53 minutes ago | parent
linksbro 52 minutes ago | parent
Jokes aside, looks like an impressive model!
jjcm 52 minutes ago | parent
nurettin 41 minutes ago | parent
mydreamof 6 minutes ago | parent
gopalv 52 minutes ago | parent
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
polotics 46 minutes ago | parent
janustimes 40 minutes ago | parent
So no, Google is not being punished, nor are they the people behind this technique.
tazjin 51 minutes ago | parent
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
ChickeNES 49 minutes ago | parent
baq 49 minutes ago | parent
timmg 49 minutes ago | parent
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
qalmakka 44 minutes ago | parent
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
YuechenLi 36 minutes ago | parent
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
vovavili 49 minutes ago | parent
Maxatar 43 minutes ago | parent
gorbot 43 minutes ago | parent
boshalfoshal 43 minutes ago | parent
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
computerdork 28 minutes ago | parent
boshalfoshal 11 minutes ago | parent
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
mike_hearn 10 minutes ago | parent
They did that for Go and it seems to have worked out for them though.
lesuorac 7 minutes ago | parent
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
fg137 42 minutes ago | parent
bvinc 41 minutes ago | parent
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
minimaxir 31 minutes ago | parent
culi 28 minutes ago | parent
LarsDu88 15 minutes ago | parent
bitexploder 3 minutes ago | parent
adamrezich 14 minutes ago | parent
wewewedxfgdf 51 minutes ago | parent
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
VirusNewbie 49 minutes ago | parent
jjice 47 minutes ago | parent
LoganDark 45 minutes ago | parent
bel8 45 minutes ago | parent
And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.
krat0sprakhar 36 minutes ago | parent
lunarboy 34 minutes ago | parent
mattlondon 42 minutes ago | parent
Behind how?
wewewedxfgdf 39 minutes ago | parent
I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.
I have no interest in benchmarks.
mattlondon 32 minutes ago | parent
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
wewewedxfgdf 26 minutes ago | parent
mattlondon 18 minutes ago | parent
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
dhdjcjcjnd 38 minutes ago | parent
gniv 36 minutes ago | parent
georgemcbay 34 minutes ago | parent
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.
singingtoday 19 minutes ago | parent
ASalazarMX 27 minutes ago | parent
- Person 1: X is garbage compared to Y!
- Person 2: Why?
- Person 1: Because I like Y.
TacticalCoder 50 minutes ago | parent
So Google is migrating codebases from C to Rust? That is interesting...
jasonjmcghee 50 minutes ago | parent
what about input?
(Maybe I missed it)
murkt 38 minutes ago | parent
jasonjmcghee 21 minutes ago | parent
And was 2M tokens IIRC after release.
There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.
nikope 49 minutes ago | parent
lanthissa 47 minutes ago | parent
I think that should be a really bad sign, but hope its great.
tamimio 47 minutes ago | parent
LoganDark 46 minutes ago | parent
alehlopeh 40 minutes ago | parent
w4yai 8 minutes ago | parent
LoganDark 5 minutes ago | parent
scirob 45 minutes ago | parent
taylorfinley 45 minutes ago | parent
spankalee 39 minutes ago | parent
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
mapontosevenths 35 minutes ago | parent
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
drusepth 30 minutes ago | parent
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
esafak 27 minutes ago | parent
IndeanCondor 38 minutes ago | parent
mapontosevenths 37 minutes ago | parent
alightsoul 35 minutes ago | parent
warkdarrior 24 minutes ago | parent
aspect0545 20 minutes ago | parent
FranzFerdiNaN 12 minutes ago | parent
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
luckydata 20 minutes ago | parent
otabdeveloper4 19 minutes ago | parent
bel8 35 minutes ago | parent
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
amanguliani 20 minutes ago | parent
gottorf 18 minutes ago | parent
yegle 17 minutes ago | parent
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
netdur 44 minutes ago | parent
tom1337 41 minutes ago | parent
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
gravisultra 44 minutes ago | parent
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
bananaflag 44 minutes ago | parent
osiris970 44 minutes ago | parent
retropragma 43 minutes ago | parent
helsinkiandrew 42 minutes ago | parent
https://www.bloomberg.com/news/articles/2026-09-30/google-gr...
nickysielicki 42 minutes ago | parent
Nobody has a moat.
aleph_minus_one 36 minutes ago | parent
This is the kind of story that you tell to investors to justify the huge amount of cash burn. :-)
mapontosevenths 33 minutes ago | parent
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
jaggederest 16 minutes ago | parent
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
ehsankia 23 minutes ago | parent
aleph_minus_one 1 minute ago | parent
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of investments and cash burn. :-)
hirako2000 34 minutes ago | parent
altruios 33 minutes ago | parent
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
culi 30 minutes ago | parent
SwellJoe 32 minutes ago | parent
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
scottyah 13 minutes ago | parent
TeMPOraL 5 minutes ago | parent
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
culi 31 minutes ago | parent
funnym0nk3y 13 minutes ago | parent
Aboutplants 30 minutes ago | parent
RachelF 30 minutes ago | parent
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
handfuloflight 25 minutes ago | parent
Not if the genius level IQs take the market share.
fumar 24 minutes ago | parent
vb-8448 29 minutes ago | parent
verdverm 26 minutes ago | parent
---
maybe it's this Anthropic post on GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
Rzor 18 minutes ago | parent
Iolaum 5 minutes ago | parent
esafak 4 minutes ago | parent
xnx 24 minutes ago | parent
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
zem 21 minutes ago | parent
nylonstrung 19 minutes ago | parent
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
torginus 7 minutes ago | parent
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
arizen 17 minutes ago | parent
LarsDu88 17 minutes ago | parent
SecretDreams 11 minutes ago | parent
esafak 8 minutes ago | parent
qgin 2 minutes ago | parent
Heidaradar 1 minute ago | parent
FranzFerdiNaN 42 minutes ago | parent
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
Razengan 42 minutes ago | parent
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
elAhmo 42 minutes ago | parent
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
arjunchint 42 minutes ago | parent
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
KeplerBoy 1 minute ago | parent
darksaints 40 minutes ago | parent
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
deanc 36 minutes ago | parent
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
ThaFresh 34 minutes ago | parent
skavi 34 minutes ago | parent
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
jeffbee 22 minutes ago | parent
computerdork 17 minutes ago | parent
mrshadowgoose 34 minutes ago | parent
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
uvdn7 34 minutes ago | parent
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
SwellJoe 20 minutes ago | parent
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
tonyhart7 16 minutes ago | parent
mattlondon 7 minutes ago | parent
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
paul7986 34 minutes ago | parent
hypfer 31 minutes ago | parent
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
nonethewiser 31 minutes ago | parent
sebzim4500 6 minutes ago | parent
sergiotapia 30 minutes ago | parent
davmar 2 minutes ago | parent
sandos 29 minutes ago | parent
This would explain why benchmarks are seemingly meaningless.
localhoster 27 minutes ago | parent
xnx 27 minutes ago | parent
yzydserd 25 minutes ago | parent
bobkb 24 minutes ago | parent
dlahoda 23 minutes ago | parent
thefourthchime 23 minutes ago | parent
tomjen3 22 minutes ago | parent
NiloCK 21 minutes ago | parent
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
ariwilson 20 minutes ago | parent
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
zem 11 minutes ago | parent
ariwilson 6 minutes ago | parent
Just compare how much better presented the Astra announcement was compared to this one: https://openai.com/index/gpt-6-astra/
dcchambers 20 minutes ago | parent
trentor 19 minutes ago | parent
jeffbee 19 minutes ago | parent
sarjann 17 minutes ago | parent
waldrews 15 minutes ago | parent
lifty 13 minutes ago | parent
tinco 10 minutes ago | parent
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
xnx 10 minutes ago | parent
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
AM1010101 7 minutes ago | parent
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
Heidaradar 3 minutes ago | parent
rao-v 1 minute ago | parent