111 points firstSpeaker 1 hour ago 102 comments
hanifbbz 1 hour ago | parent
If anyone has counter-arguments or cares to make me smarter, I'm all ears.
boxed 54 minutes ago | parent
Insanity 40 minutes ago | parent
But maybe I'm misremembering how fragile GH was in the 2010s.
boxed 6 minutes ago | parent
automatic6131 52 minutes ago | parent
N_Lens 59 minutes ago | parent
xnorswap 54 minutes ago | parent
rfgplk 32 minutes ago | parent
federicobrancas 52 minutes ago | parent
flohofwoe 45 minutes ago | parent
ben_w 36 minutes ago | parent
That's not a boast, I don't think I was particularly good at that back then, e.g. I didn't really get how to think about automated tests until much later.
It's just to say that no, coding and software engineering are not the same thing. "Code Monkey" is a dead (or perhaps "undead") role now, but it wasn't always so.
vatsachak 50 minutes ago | parent
Long term planning in LLMs has not been solved.
hanifbbz 45 minutes ago | parent
antonmks 49 minutes ago | parent
_fzslm 47 minutes ago | parent
GitHub's Copilot cloud agent offering is suffering with a case of some of the worst corporate ADHD I've seen. We built a cloud agentic development pipeline on it, and it seems like almost every other week they silently change something with zero public announcement or documentation that creates real disruption for our team.
That's real, breaking changes to the platform that clearly aren't being tested/reviewed before being pushed to prod. Again with zero public announcement or documentation.
Support is useless – we're paying customers in the 4-5 figures and our tickets go unanswered.
wannabe44 44 minutes ago | parent
Shank 43 minutes ago | parent
I have no doubt that if you provide any AI system with an oracle with expected behavior that it can match that oracle with some amount of $ and tokens. I haven't seen any demonstration of anything else. Rewriting a codebase was always a challenge for humans not because of complexity, but because of the time and effort involved in matching the old version's prior behavior. It doesn't have anything to do with the serious level of work required to build something truly new from scratch in a performant way.
rfgplk 28 minutes ago | parent
Shank 11 minutes ago | parent
For example, any amount of software development involves fixing bugs, getting feedback from users on ideal workflows, an iteration loop of performance and bug tuning, etc. AI cannot simply create, from scratch, perfect software. Even using the SOTA models on max effort does not produce bug free software of any meaningful complexity or innovation out of the box. All that has changed is that the act of physically writing code and implementing existing patterns is now effectively a marginal cost.
Most line of business software is not e.g., delivering a company's income. Most software is in back-of-the-house internal products that do various internal tasks. I have no doubt that these processes are now far easier to build.
If the new Copilot is so great, why is it completely out of the current zeitgeist when compared to Codex and Claude Code?
spaqin 41 minutes ago | parent
OtherShrezzing 37 minutes ago | parent
Especially expensive when you take into account the amount of that code which must have been boilerplate & meta-code in nature, meaning it should have been straightforward to move.
verdverm 16 minutes ago | parent
When the Go team ported the original compiler from C to Go, they wrote a program that did ~99% of the work
flohofwoe 35 minutes ago | parent
verdverm 17 minutes ago | parent
would be curious to know how many times "unsafe" appears in there, have seen rust devs comment on how the ais like to use unsafe to work around difficulties with memory management, like how they will sometimes subvert tests
j45 49 minutes ago | parent
askonomm 48 minutes ago | parent
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
whatever1 44 minutes ago | parent
rfgplk 36 minutes ago | parent
whatever1 25 minutes ago | parent
It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.
bunderbunder 4 minutes ago | parent
For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:
1. The promised scalability improvements did not materialize. Instead, it got worse.
2. Observability had been lost. The telemetry was no longer trustworthy.
3. Users stopped trusting it because it was producing incorrect outputs.
What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.
But the real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.
But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.
0c3ca83 19 minutes ago | parent
Ambolia 12 minutes ago | parent
rgoulter 43 minutes ago | parent
Brings to mind this classification https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...
"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""
banannaise 34 minutes ago | parent
Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).
automatic6131 20 minutes ago | parent
Melkazt 3 minutes ago | parent
ben_w 43 minutes ago | parent
> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
Thanemate 40 minutes ago | parent
ben_w 26 minutes ago | parent
And?
The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: https://en.wiktionary.org/wiki/Lizardman%27s_Constant)
It's a continuum for both them and us, even if the mechanism is wildly different.
> Making mistakes is not the same as non-deterministic.
i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.
p-e-w 24 minutes ago | parent
hanifbbz 42 minutes ago | parent
rfgplk 38 minutes ago | parent
This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"?
swiftcoder 34 minutes ago | parent
Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate.
dnikolovv 11 minutes ago | parent
askonomm 2 minutes ago | parent
Zardoz84 39 minutes ago | parent
You will be surprised how many times, catches errores made by the AI coding agent. However,as you point, isn't deterministic. And you can guarantee the end results is 100% fine code
huijzer 37 minutes ago | parent
What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.
empath75 26 minutes ago | parent
mjr00 14 minutes ago | parent
Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.
Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.
bushido 47 minutes ago | parent
Instead, I think what's closer to solved and what we're in the process of solving is product development.
Story: A while ago, I had a few programmers who were really, really fast almost always missed the mark on the assignment wrong. I loved having them on projects because in the time my senior precise engineers could deliver a MVP, the fast engineers would build the wrong thing, collect feedback, reiterate, build the wrong thing, collect feedback, eventually inching closer and closer to a product people would pay for, and it would almost always get delivered faster than my seniors.
I feel AI does the same thing.
username_my1 39 minutes ago | parent
I got lazy around claude fable and astra, and asked them to work in loop (pick specified issue, develop it, qa it ...) have a separate CTO checking on arch.
at the end both models swore that the code is perfect and well designed and nothing is lacking.
I ran the software and it suddenly started writing large amount of data to CSV files instead of the typical DB usage.
AI decided to use csv for testing, and just drifted away. 0 regards to the actual project, 0 regards to common sense.
anecdotal but really weird, the project category is rather standard, I wouldn't accept such a mistake from a junior developer.
fingerlocks 12 minutes ago | parent
It compiled and ran just fine. If you weren’t reviewing the code holistically or keeping tight book keeping of your allocations you would not have noticed. Every single commit in isolation looks perfect. Very eye-opening
rfgplk 30 minutes ago | parent
Isn't it the opposite? How to build something is rather solved, but what to build isn't?
bushido 22 minutes ago | parent
But that's not solved in traditional product development either.
Product development an iterative process to get a product fully functional. In 2021, if you ask me what the timeline for a small product/substantial feature, I'd say a few weeks to a month to get a basic MVP, and then another 12 to 18 months to get a feature polished and in a good shape to be stable.
When people put it in the coding frame, what they do it as is saying we've gone from 18 months to minutes or days. That's just not true.
We have gone from eighteen months to depending on the complexity, a 1-4 months.
aside: To be candid though, the compressed time also means the frustrations people experience with a product in 18 months have also been compressed. They still exist, they're all there, they're now just non-stop.
grim_io 46 minutes ago | parent
I don't like the feeling being judged and tested by the author (missing number 5 point in the list).
gyesxnuibh 39 minutes ago | parent
hibikir 43 minutes ago | parent
This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.
There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.
So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.
hanifbbz 39 minutes ago | parent
The difference seems to be that some companies are above the law apparently.
jstummbillig 43 minutes ago | parent
Dear lord. Is that supposed to reflect the average thoughts and motivation of a person you want to hire? Or that of their employer?
IshKebab 42 minutes ago | parent
Doesn't matter what you think about AI, "it isn't perfect" is clearly a nonsense reason not to object to it.
vincent-uden 36 minutes ago | parent
It's for example impossible to have a discussion with an LLM where you both learn something which you can apply tomorrow. The LLM doesn't learn until the next model is released and by then your discussion is just a tiny fraction of the training data (if present at all). AGENTS.md, skills and so on are just a proxy for what we actually want, an agent that listens and understands. A proxy mind you, that requires constant tweaking with no sign of generalisation in sight.
gyesxnuibh 36 minutes ago | parent
I'm also not sure what humans being non-deterministic even means here. The point is if you're comparing results with NFR, pure agentic coding falls short.
idz 33 minutes ago | parent
gradus_ad 41 minutes ago | parent
hanifbbz 35 minutes ago | parent
vmg12 40 minutes ago | parent
AI can write CRUD API endpoints almost perfectly now. It can also write quicksort, a heap, whatever much quicker than I can.
It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).
I suspect this will remain a problem for the models for a long time. All the things that the models are currently good at are the low hanging fruit of reinforcement learning for coding.
Think about the kind of reinforcement learning environment that needs to be created to train a model to become good at building and designing large scale software end to end. It would be a slog because you need to build the large scale software up front and then break it down to train the model to construct it in a systematic manner that allows for the software to evolve. And then you need enough of these training environments for it to generalize. I think they will eventually figure it out though but it may take a while.
nemo44x 30 minutes ago | parent
Does that really matter? Those are things so that humans can better understand and extend a code base. That mattered when writing code was expensive and took time.
Now if it can pass all the tests it’s fine. If there’s an issue just have it rewrite things immediately. New bug? Generate a new test and rewrite code.
All, or many, of the old things that mattered just sort of don’t anymore.
mxey 22 minutes ago | parent
efficax 40 minutes ago | parent
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
12ag5a 27 minutes ago | parent
This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.
bananaflag 18 minutes ago | parent
MattDamonSpace 17 minutes ago | parent
rowanG077 2 minutes ago | parent
grumple 26 minutes ago | parent
flatline 16 minutes ago | parent
Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.
pu_pe 4 minutes ago | parent
Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.
TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.
Sevii 38 minutes ago | parent
rkozik1989 33 minutes ago | parent
That concept might work a lot of the time but you will definitely run into situations where that'll never produce a correct or working response. To actually learn something you need an environment/playground to apply what you think you know and observe the results. Without that you're not really learning, you're jus regurgitating what people want to hear.
giovannibonetti 33 minutes ago | parent
I think a few of the industries listed like defense and aviation have low risk tolerance. However, from my (somewhat brief) experience of working in two health techs for a couple of years, I strongly disagree that healthcare has low risk tolerance for tech. Granted, they make run-of-the-mill CRMs, but I was baffled at how tolerable it is to have egregious user experience that makes users waste multiple hours per month with clerical work that is very painful because the UIs are very slow and buggy.
HotHotLava 21 minutes ago | parent
It means risk that the software stops working after an update. Which usually trades off iteration speed and best practices (i'm pretty sure the average startup has way better security practices by just delegating to google/aws than the average manufacturing software business) in exchange for a rigorous testing and rollout schedule.
So I'm also not sure that the article has a point at all, the human writing the code was never relevant to avoiding the "risk" in these industries in the first place.
rgoulter 30 minutes ago | parent
It's not clear to me if the claim is:
(1) "If you used an LLM to generate code, and the code works, you're wrong if you think the code is okay"
or
(2) "If you used an LLM to generate code, you reviewed the code and found it to be of decent quality, then you're wrong".
> If you’re toying around, LLMs do a great job. That’s why some of the most aggressive proponents of the “coding is solved” narrative have nothing to show for it.
I also don't get the "LLM proponents have nothing to show for it" statement.
It's really quite common now to see on HN all sorts of LLM-assisted programming projects. The quality varies from slop where little thought was put into it, to high quality results where LLM coding assistance was able to let talented developers produce things they otherwise wouldn't have time to do.
I'd say it's obvious that LLM coding agents can be very useful for a lot of programming related tasks.
EDIT: That is to say, LLMs are obviously useful for use cases above/beyond toying around. It's not a dichotomy between "I'm never touching an AI" and "thoughtlessly accepting everything the LLM outputs".
metalspot 30 minutes ago | parent
But to anyone even vaguely thinking of taking this seriously, go look at what antirez, dhh, jared sumner, mark brooker, and many other real engineers who have ship real things are doing and saying.
Most of these people have spent their entire lives contributing to open source, and they have proved their skill shipping working software and scale for decades. They are really trying to help people by showing and telling them exactly how AI works and how to use it to make better software.
lordnacho 28 minutes ago | parent
- Coding in the small is solved. I have a current state, I want to change it, and I know how I want to change it. Eg, I have a blocking TCP handler for some reason, and I want to make it async. I can either fiddle with it or just let LLM make the changes for me.
- Coding in the larger sense is never solved. You need judgement to decide what you want made. No matter what you're building, there will be decisions to make (Who/what is it for?) and those decisions change over time. LLMs can take some default decisions for you, and if you're fine with those, you get the default (great for POCs). However you might not even realize what it decided to do for you. At some scale, you will be spending a lot of time going over those decisions. But what we have now is that the friction of changing the decisions is quite a lot lower. You can now test a lot of things that previously were very time consuming.
- The point that LLMs are probabilistic is not as important as it's made out to be. If I ask a junior dev to code up something, I also don't know what he'll make. Heck, you can be sure that you are able to solve something, yet you yourself don't know what the solution will look like. Maybe it turns out the library you were going to use isn't appropriate after all. You don't know what you will use in the end, but you do know that something will fix the issue. There can be more than one solution to a problem, and it doesn't always matter which one you find.
- I STILL think that LLMs are at their best mostly as advanced predictive text. In the sense that it's mostly good at implementing things that you've decided are needed. This can mean a heck of a lot of code, but you have to know the tradeoffs. What was decided, what were the costs of those decisions in terms of maintainability, money, time to change it, and so on.
Marha01 28 minutes ago | parent
"Don't confuse coding with software engineering" is a valid point, the rest seems like ranting.
verdverm 21 minutes ago | parent
lr4444lr 28 minutes ago | parent
As for accountability, it always laid with the employer. You think those nameless contractors whom Boeing hired suffered any consequences for that 737 Max glitch? Using AI won't change that.
AI doesn't have to solve all these coding problems to be worth handing the reins to it: it just has to substantially better on average than humans over the long haul, which it already is, especially if you have good verification of "done" and "working" in place through automated testing mechanisms. Perhaps we might say that QA is having its moment.
It doesn't mean humans aren't needed, but they aren't writing much if any code anymore.
somewhereoutth 27 minutes ago | parent
So - prose, code, or image, it appears that some work has been done, but in fact the [actually needed] work has likely not been done.
tmarice 26 minutes ago | parent
If you were a professional software developer, you a) learned to touch type, b) started using vim/emacs keybindings to navigate around the project, and c) used a framework which already abstracted away a large part of the menial work.
And going all-in on the loop and no-code-review nonsense in a project someone is actually paying you for, I can only assume means you're hoping not to be around when the slop tower collapses.
lilerjee 25 minutes ago | parent
They are thinking: Please input everything you know, or just use it and it will collect everything in your PC or server automatically.
Stop lazy, stupid and dangerous behaviors.
samayashar 23 minutes ago | parent
All this doesn't change the fact that software engineers are going nowhere because nobody trusts AI. If a model can escape highly secured sandboxes, then we're definitely not running these agents overnight on our systems. I am sure the next-gen of models will focus more on security and the trust factor will start developing, but that's a long way down the road.
People trust people, not systems.
jbs789 22 minutes ago | parent
What’s your number?
mglvsky 20 minutes ago | parent
"coding is solved" == "gastown-like systems give a brand-new and useful software"
I don't recall whether GasTown succeeded...
AnotherGoodName 6 minutes ago | parent
Over this weekend in chat with the games discord watching as it iterated a harness built an entire implementation of the board game terraforming mars https://tfmbot.com using agents and harnesses for them.
I think if you can implement a board game end to end by feeding in the rulebooks and having a harness spawn agents to validate it’s reasonably solved.
Trusteando 15 minutes ago | parent
armchairhacker 10 minutes ago | parent
Coding is not solved because you can’t simply prompt an LLM to make an AAA game or enterprise tool.
wg0 4 minutes ago | parent
drwallace 2 minutes ago | parent