19 points vinhnx 2 days ago 60 comments

someonebaggy 10 hours ago | parent

Everything coming from GNOME about software quality should be taken with a Strategic Petroleum Reserve of salt.

While other projects rewrite things in memory-safe ways, GNOME's response is to ask a chatbox if there are any memory vulnerabilities.

tom_ 9 hours ago | parent

Say what you like about the random word sequence generation machine, but it does actually seem to be usefully good at generating sequences of random words that correspond to problems in your software. And if you're inclined to write the code by hand, it's probably going to be easier to fix up your existing shit than rewrite it all. (And if you're going to use AI, then you're hardly going to listen to me.)

jeremyjh 9 hours ago | parent

The purpose of this is to address a practice common in certain open source ideologies of banning all AI contributions - including vulnerability reports. Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.

lelanthran 9 hours ago | parent

> Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.

That's not only an uncharitable take, it's also wrong.

If 999 out of every 1000 "reports" from a specific source is wrong, then it is not irrational to disregard all 1000, especially when they can be generated faster than you can read.

I mean, it's just probabilities, right? If you're okay trusting output from an LLM, you should be okay with using statistics in general as a source for informing decision-making.

jeremyjh 9 hours ago | parent

The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?

lelanthran 7 hours ago | parent

> The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?

Right, but that assertion does not contradict what I said: there's a difference between the articles premise (AI reports are mostly valid) and what I said (3rd-party submitted AI-reports are mostly invalid).

See my reply to a sibling poster who also implies I did not read the article.

boxed 8 hours ago | parent

You are out of touch unfortunately. Your logic was correct last year, but things have changed radically and super fast. You need to re-evaluate, and this article by a core GNOME developer specifically doing security work should have made you do that re-evaluation!

juped 8 hours ago | parent

The last 1000 times someone has said "nah bro AI was bad last year but this year it's good trust" have been wrong, I am also comfortable being informed by evidence.

qarl 8 hours ago | parent

> I am also comfortable being informed by evidence.

Then look. If you can't judge, then trust the experts. This article is written by experts.

someonebaggy 7 hours ago | parent

They did this the last 999 times and always came to the same conclusion: that the AI promoters were talking nonsense. At what point should you stop listening to people who only talk nonsense so far, to avoid getting DoS attacked? Must the villagers look for the wolf every time the boy cries?

qarl 7 hours ago | parent

No. But when the wolf experts announce that there are a dangerous number of wolves - and you ignore it - that's a problem.

The people writing this article are experts. They cite other experts.

If you can cite real data from the last few months that still claims there are no wolves - and it's not just insane anti-wolf propaganda - I'd love for you to show me.

someonebaggy 7 hours ago | parent

What if the last 999 times wolf experts announced there were a dangerous number of wolves, no wolves were found?

qarl 7 hours ago | parent

You're going to need to leave the hyperbole behind and use actual facts if you want to continue.

Cite data from the last few months.

someonebaggy 6 hours ago | parent

But you get to use hyperbole to defend your side? No, I'm not interested in fighting Brandolini's Law.

qarl 5 hours ago | parent

Friend, I only used hyperbole to answer yours.

Let's return to actual data: the article we are discussing is experts claiming that AI should be used to report errors in software.

Do you have data to support your side?

EDIT: I did a little research. These projects accept AI contributions: Python, NumPy, SciPy, pandas, scikit-learn, Django, Kubernetes, the Linux kernel, Firefox, Flutter, Homebrew, curl and PyTorch.

Why do you think they do, if it's such a bad idea? Are they all idiots? They have no idea what they're doing? Linus? Really?

lelanthran 7 hours ago | parent

> The people writing this article are experts. They cite other experts.

"Economists have predicted 18 of the last 2 recessions".

I mean, c'mon! You have never read that?

Besides, when "expert in $FOO" means "familiar with $FOO that's only 6 months old", then it's not unreasonable to be skeptical.

In other fields, an expert is someone who's studied the specific field $FOO for decades. Here we're talking about a skill level that is not distinguishable between "1 weeks experience" and "two years experience".

qarl 6 hours ago | parent

Your argument boils down to "I'm not trusting that bridge! Have you seen how many mistakes astrologers make!!"

I've found that software engineers are good at analyzing the public reports they receive for their own projects.

Let's not over generalize, shall we?

ethersteeds 8 hours ago | parent

If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:

> Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)

lelanthran 7 hours ago | parent

> If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:

I read the article very carefully, including the bit that you quoted. Here's what I read:

> Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.

And that's with them running the scanner, not with submitted reports by 3rd-parties! When you welcome AI reports, everybody is going to submit the same report, just differently ordered and differently worded.

I mean, he even said:

> I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:

Sure, he attributes it to a financial incentive, but it's clear that submitted AI reports will overwhelm, and the only way they got to a measly 40% real-bugs was by do the scanning themselves.

(Also, I wish all these sibling posters implying that I did not very carefully and thoroughly read the article would, themselves, read the article!)

qarl 6 hours ago | parent

Yes. They will need to use AI to process the increased load of AI reports.

I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers. Sure, some people refuse... but...

bigstrat2003 6 hours ago | parent

"use the tool to fix issues created by using the tool" is not a valid solution. The solution is to stop using bad tools.

> I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers.

When LLMs actually can reliably do their jobs (which compilers do), then they might be an essential tool. Not before. For now, they are slop machines used by people who care more about going fast than getting things correct.

qarl 6 hours ago | parent

> When LLMs actually can reliably do their jobs

LLM agents do a fantastic job of finding exploitable bugs in code. MUCH better than humans.

They would also do a fantastic job isolating the duplicate reports as described above.

So what's your issue?

saghm 7 hours ago | parent

I think the disagreement here is whether this constitutes a "specific source" or not. I've seen some people produce things with incredibly quality and others produce useless slop all with the same AI tools and models, so defining that as one single source doesn't seem like a very good way of viewing things.

lelanthran 7 hours ago | parent

Lets assume your argument is valid, and further assume hat only the high-skill people submit AI-reports.

Out of, say, 40 bugs that SOTA models can find, you're still going to have to sift through all the hopeful wannabes who each submit that same list of 40, but differently worded, differently explained and with different PoC code.

The problem still remains when welcoming AI-reports from the world: you could potentially spend all your time on examining and then discarding these reports without even getting to any new bugs in those reports.

saghm 6 hours ago | parent

You could make the same argument for rejecting all reports from third-parties in a pre-AI world though. It just seems like you're picking an arbitrary property to extrapolate a trend from when there are plenty of other similarly arbitrary properties you could extrapolate similar trends from. If I noticed that most bug reports that come on Tuesdays are low quality, I don't think it would make sense to have any bugs that get reported on Tuesdays get auto-closed, regardless of the statistical trend.

wonnage 6 hours ago | parent

Projects were already doing this by making you jump through hoops to submit an issue. It filters out the low effort chaff. I think asking for people (instead of their agents) to submit reports is on the same level.

lelanthran 4 hours ago | parent

> You could make the same argument for rejecting all reports from third-parties in a pre-AI world though.

No one did though, because the friction involved in finding and submitting bug reports meant that a single individual could not overwhelm a project in spurious reports.

saghm 1 hour ago | parent

If a single individual overwhelms a project now, it's possible to just reject their reports without mandating no AI help at all. I can't tell if you're claiming that all LLMs (or specific vendors/models) should be treated as a single individual or if you're missing the point I've been trying to make.

troyvit 6 hours ago | parent

I'm a KDE guy through and through, don't get me wrong, but statements like this deserve attribution, otherwise it's just FUD.

nottorp 8 hours ago | parent

Hey does that mean they'll use "AI" to allow users to customize the desktop environment again?

someonebaggy 8 hours ago | parent

No - non-customizability is a design goal of the GNOME desktop, since they want the experience to be identical for everyone.

WD-42 7 hours ago | parent

You can do this by vibe coding extensions, and it appears many people are doing this. Ironically this probably makes GNOME the easiest DE to customize at this point.

type0 6 hours ago | parent

> GNOME the easiest DE to customize at this point.

It's not even false, you do know that KDE exists

WD-42 4 hours ago | parent

KDE cannot be modified by extensions in the same way GNOME can. This is a fact. The problem is the extension needed to be written. They are trivial to vibe now. You can make them do whatever you want.

skydhash 3 hours ago | parent

> KDE cannot be modified by extensions in the same way GNOME can.

I was pretty sure that was false. I did a quick check and found

https://develop.kde.org/docs/plasma/scripting/

https://develop.kde.org/docs/plasma/kwin/

https://develop.kde.org/docs/plasma/krunner/

WD-42 3 hours ago | parent

Yes, you can do some light scripting or create your own plugins for specific apps (krunner). This is not near what you get with gnome extensions, which allow you to essentially execute arbitrary code in the shell’s main process.

The shell got a lot of hate for being written in gjs, but you won’t get that kind of functionality from a shell written in a compiled language.

Sevii 8 hours ago | parent

They should be using an automated AI agent to validate vulnerability reports. Have it pull up the code base, confirm the bug exists, try to reproduce, then update the ticket. You might even run another AI agent pass to clean up the text before humans look at it. AI writing quality improves a lot with multiple passes.

saghm 7 hours ago | parent

"Vibe reviewing" is an underrated way of dealing with vibe coded PRs. There's going to be a lot of low-quality noise that's probably better to just close without a human looking at it, but it would be unfortunate to ignore the stuff that might be worthwhile because of that.

tonyedgecombe 7 hours ago | parent

> AI writing quality improves a lot with multiple passes.

Or you end up with an AI version of Chinese whispers.

jayd16 7 hours ago | parent

Sounds really expensive as far as the work you're willing to let the public invoke on your project.

asjq178 6 hours ago | parent

[flagged]

HPsquared 6 hours ago | parent

AI is well-suited to the tedious work of testing and QA.

hungryhobbit 6 hours ago | parent

"of testing and QA" ... for a certain percentage of "testing and QA". The rest still needs humans.

itishappy 6 hours ago | parent

The article makes a strong point that the slice that needs humans is dwindling quickly. The article seems to suggest our ability to write prose for other humans is our main differentiator.

convolvatron 6 hours ago | parent

just from first principles its really not. given its propensity to just make shit up and cheat, ai is really useful when there is some kind of formal or exhaustive checking as a wall for it to throw crap at. so if you rigorously define success and spend a bunch of tokens, you could easily save time and money. but if you ask the ai 'is this correct' and it says 'yes!', or even 'no!', then you've really learned nothing. this isn't a can you can kick arbitrarily far down the road.

I think writing tests is a great use of ai, but only if the tests themselves are throughly reviewed or are themselves validated by statements in a formal system.

itishappy 6 hours ago | parent

> given its propensity to just make shit up and cheat

That's more-or-less how I define QA work. The goal is not proving overall correctness, it's surfacing individual issues.

saltcured 5 hours ago | parent

I think you could use AI for this if you do it properly as a sort of adversarial coding project. Have it build tests to break the code. Don't give it the job of making a test suite that passes.

I think people higher up are warning against a naive mistake, which is asking one agent to write both the product and the test suite. Here, making things up and cheating becomes a problem of quality theater...

itishappy 4 hours ago | parent

Totally agree. My point is simply that this aligns with existing roles.

The QA role (regardless of whether it's performed by an AI or a human) typically involves breaking tests, not writing them. (This applies more to unit tests, integration tests blur this line.)

convolvatron 4 hours ago | parent

For me QA is tasked with making sure that the product functions as intended. that means coming up with clever test suites to explore the spaces that the developers didn't think to, making sure the coverage for operational problems is adequate, and also long term tracking of performance and memory utilization.

this doesn't sound like your model, and I'm unclear what it means to break a test. maybe test automation, in which case, sure that seems fair game for AI, but that not where the real meat is.

skydhash 3 hours ago | parent

> For me QA is tasked with making sure that the product functions as intended. that means coming up with clever test suites to explore the spaces that the developers didn't think to, making sure the coverage for operational problems is adequate, and also long term tracking of performance and memory utilization.

This pretty much. My last job didn't really have QA so I did the next best thing which is writing a bunch of integration tests for the use cases that matters for the product. They were not an indication for correctness, but more like a canary to warn me if I break something while developing. Bugs reported by consumers usually warn me of area not well covered.

QA would play the same whole. They shouldn't need to check for code correctness, their most useful task is to surface bugs that breaks the product requirements (performance, security, business logic,...). And for that, having knowledge of the implementation is unnecessary. In the above example of writing integration tests, I took care of only using the public interface of the modules.

itishappy 2 hours ago | parent

That largely aligns with my view. I think in a large enough codebase (roughly the size QA becomes necessary/relevant) it becomes impossible/infeasible to guarantee correctness. I think the line "clever test suites" highlights the distinction.

In my mind, tests are everyone's responsibility, but the goal differs. A regular developer adds features and should be writing tests to prove their code functions as intended. QA does not add features, so their goal is finding holes in code added by others.

It's the adversarial relationship mentioned by a parent:

> Have [QA] build tests to break the code. Don't give [QA] the job of making a test suite that passes.

centuryfall 6 hours ago | parent

This same statement has been said countless times, replacing “AI” with whatever trend is big at the time. It has also been wrong in every case where that thing is said to replace QA.

ciupicri 2 hours ago | parent

Is this some kind of joke? What more features could Gnome reduce?

tonymet 5 hours ago | parent

why not have AI do more quality control?

KronisLV 44 minutes ago | parent

I’d posit that well over half of tokens spent should go not just to write new code/software alone, but in review, testing, as well as tooling. Every bit of code an agent writes should at the very least have an adversarial review loop.

tonymet 29 minutes ago | parent

I totally agree, and the differential token cost is minimal if they are included during product code generation.

greatgib 3 hours ago | parent

> GNOME is primarily written using unsafe programming languages where simple mistakes in our code lead to devastating consequences for our users, and we make these mistakes all the time. No matter how much we try, GNOME developers will fail write secure code when using unsafe languages like C, C++, or Vala: it’s just too hard for even experienced developers to do properly.

Once I saw that I knew that this article was a ridiculous stack of bullshit. Gnome might have some bugs and shortcomings but now millions desktops are using Gnome based window manager for decade and the world didn't collapse yet.

And let's not forget the ridiculous assumption that you can't handle a software quality without using AI...

skydhash 3 hours ago | parent

This remind me of OpenBSD's code. Not saying there isn't any bugs there, but the risk is mitigated by writing simple and clean code. They do not rush to add the latest ideas and wishes to the codebase.

bunderbunder 52 minutes ago | parent

But if I may steelman the article a bit:

What was good enough in the past may not be good enough now. In the past these overflow defects were as hard for attackers to find as they were for developers, because they had the same tools available.

Now we have LLMs, and they can apparently find all sorts of problems that, for whatever reason, weren’t being found with manual review, static analysis and fuzzing. We have to assume that hackers will use the technology to find vulnerabilities. If maintainers don’t do the same, then they are ceding an advantage and leaving their users unnecessarily exposed.

I don’t know that I completely agree with the above. (For example, I don’t know how scrupulous GNOME has been in the past about non-AI tools for automated defect discovery, or how well they compare to AI.) But it at least feels like a much more charitable interpretation of the article’s main thrust.

jongjong 2 minutes ago | parent

Producing correct code is insanely difficult. It's hard to convey this to junior or even mid-level engineers.

In fact, for many complex projects, it's essentially impossible, even with AI.

Eventually, you get to a point when you have every feature you could possibly want but the list of tradeoffs is long and yet not worth trimming.