44 points swolpers 1 hour ago 58 comments
FinnLobsien 1 hour ago | parent
I believe that we're in a scenario where usage is unlikely to go down and neither are frontier AI costs.
I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
A marketer doesn't need to default to Opus 5.5 to upload a blog article with MCP, which could be done by a model 10% of the price.
awesan 50 minutes ago | parent
After we started excessively documenting and extracting skills everyone has stopped complaining about running out of tokens, because agents stopped having to reconstruct the full context each time from scratch. Harnesses like claude code also push the model to aggressively keep this documentation in sync so there's little concern about drift.
The worst thing to do from a token usage pov is to give a model a vague open ended prompt because they are so scared to be wrong that they'll waste a ton of tokens "thinking" through the issue and verifying everything. Whereas they almost trust skills blindly and skip all this unnecessary work.
skeptic_ai 49 minutes ago | parent
StilesCrisis 10 minutes ago | parent
Sevii 30 minutes ago | parent
mrweasel 23 minutes ago | parent
bluGill 17 minutes ago | parent
Tanjreeve 29 minutes ago | parent
> I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
bluGill 14 minutes ago | parent
FinnLobsien 9 minutes ago | parent
Also this feels like a good way to detect runaway jobs. I’ve seen something similar implemented for data plans where you had unlimited data, but every 50 GB had to request another free 50 GB to prevent people from turning their phone plan into their home WiFi
bunderbunder 14 minutes ago | parent
This summer my team basically shut down when we hit a usage cap because we had recently pushed a decent portion of our devops workflow into agent skills. It had been done in such a way that it was difficult for humans to navigate. Relevant scripts turned out to be buggy and poorly documented, and nobody realized because these harnesses that are tuned to be absurdly tenacious about searching for workarounds had been quietly burning heaps of tokens on muddling through instead of raising any alerts about the horribly broken state of the system.
I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
FinnLobsien 4 minutes ago | parent
> I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
It's true that moderate usage harms current harness providers, but I think that's because we currently conceptualize them as AI services. I think in the near future, companies will have harnesses that centrally configure MCPs, CLIs, model routing, etc.
and they'll be much closer to an auth/permission service than to an AI service.
simianwords 1 hour ago | parent
1. choose a good model
2. choose the appropriate reasoning effort
3. choose a prompt to nudge it even further
Then it comes down to understanding the intuition of what kind of task deserves what effort?
CharlieDigital 55 minutes ago | parent
(Couldn't even get this to happen in a 30 person team...)
epistasis 49 minutes ago | parent
I use these models all day long, experiment, and have no clue how to choose that.
It just showed up one day in the interface with no explanation or guidance. Its use is mysterious, its effects unclear except through intensive experimentation, and to this day it mostly seems "how many bad decisions will Claude go forward with when it finally dumps out screenfuls of text instead of getting better guidance early on" though it's certainly not a guarantee on anything.
These models are being released at breakneck speed even before their creators know how to use them. It's a big project of collective discovery to figure out what they are doing and how to use them.
simianwords 41 minutes ago | parent
hilariously 59 minutes ago | parent
Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
This is getting stupid folks.
nathanaldensr 57 minutes ago | parent
2OEH8eoCRo0 56 minutes ago | parent
bluGill 19 minutes ago | parent
ndriscoll 49 minutes ago | parent
Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
rolosa 39 minutes ago | parent
delecti 25 minutes ago | parent
js8 38 minutes ago | parent
On these things, yes. But in the case of LLMs, this is impossible. It's not possible to understand what the models are doing, and for several different reasons. You can at best evaluate them, like you do with your fellow engineers when hiring. But that's not a guarantee of anything.
And that's why I don't think this is an engineering renaissance. Engineering progress goes with better understanding of our tools, and adoption of more rigorous practices. LLMs go in the opposite direction.
sanderjd 34 minutes ago | parent
In practice, I think that kind of effort is mostly a hindrance at the moment, because of how fast things are moving, and because so much about this is subjective and everyone has different preferences.
analognoise 31 minutes ago | parent
bunderbunder 25 minutes ago | parent
Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o's capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn't be replicated on any of the newer "long" term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.
Combing back to coding agents, where they're releasing new models and harness tweaks multiple times per month, and vibes are the only - let's not say sensible, maybe realistic - thing a developer reasonably has to go on.
bluGill 22 minutes ago | parent
If you enjoy being an expert in which LLM is best, then I'm all for it. However that isn't an interesting problem for me or my boss. I'm happy someone else figured it out so I can work on interesting problems.
The above likely scares all the model providers: they are a commodity with easially substitute competition. There is a minimum quality standard, but once you meet that there is nothing to differentiate you and so price matters and it becomes a race to the bottom.
themgt 46 minutes ago | parent
Right, the problem is frontier models can do ~all of that pretty well, a lot better than they could 9 months ago. Luna 6 can do most of it practically free and instantaneously. So, your quote is likely what's increasingly being said in c-suite meetings as they decide to start mass layoffs.
SolarNet 45 minutes ago | parent
The level of appropriate LLM usage in my opinion? Ask Gemini some questions when you need to search the web then read the sources it gives you.
LLM written code still has the problem IBM identified. A computer cannot be heald responsible, and so it cannot be allowed to make decisions. That applies to executive management AND software engineering.
Of course that requires working in an industry, like aerospace, where engineers are (usually, Boeing not counting) heald accountable.
StilesCrisis 16 minutes ago | parent
ndriscoll 11 minutes ago | parent
You can ask it to implement the pattern you want. Like we had a bunch of gnarly intertwined logic inside of looping constructs that I could say, hey, create iterators for this, extract out the filtering logic, extract out the transformation logic, etc. I'm making the decision, but it can do the mechanical work.
You can also use them to quickly build out prototypes for different approaches so that you can make better informed decisions at an architectural level.
You can give them a description of a bug and it will read through your code and figure out where the problem lies, even across multiple repos with complex interactions. You can of course still confirm it, but they've been superhuman at this since at least last December.
LLMs are absolutely a game changer for software development in the hands of somebody who knows what they want. For low-level things, they're an extremely solid junior (they don't ever really make mistakes, they just write tasteless crap). For high-level things, they're an extremely solid discussion partner with wide knowledge and great reasoning skills.
palmotea 35 minutes ago | parent
All while suffering from the increased the mental load/lack of focus from even more workload. "You've got AI, you should be able to get it done today, right?"
ericd 27 minutes ago | parent
Everyone has a great researcher piped to their desk now. You can have it follow the Twitter zeitgeist for you, you don't have to do it yourself.
neom 57 minutes ago | parent
rglover 55 minutes ago | parent
IMO, human in the loop is the only serious usage of AI (I know, I know, "software factories bro"). Everything else is a hope and a prayer and a big bill.
mglvsky 18 minutes ago | parent
edit: grammar
bentt 53 minutes ago | parent
cyanydeez 49 minutes ago | parent
We're sorta in the age of alchemy. Lots of cranks out there, but there are real recipes.
macNchz 40 minutes ago | parent
gorjusborg 40 minutes ago | parent
- full autonomous camp
- developer augmentation camp
- no-ai camp
The trouble I see with all of it is that the future seems unpredictable at the moment. The costs related to AI are low enough at the moment that full autonomous seems to be possible, but we have reasons to believe that costs will rise significantly, which may change that calculus. The no-AI camp is in ostrich mode, and is betting on this all going away once the bubble pops. The developer augmentation camp treats it like just another tool, which is somewhere in the middle.
The trouble is that even if there is a clear advantage today, the ground truth of costs built in is probably not stable.
1223975 33 minutes ago | parent
https://www.reuters.com/business/media-telecom/musk-says-he-...
That will fix adoption of a broken technology and all debt issues!
sanderjd 29 minutes ago | parent
fidotron 28 minutes ago | parent
If you don't have at least some workloads doing that you're missing out on the biggest wins from the current phase of the technology.
winwang 14 minutes ago | parent
sajithdilshan 48 minutes ago | parent
dominotw 44 minutes ago | parent
If i had to guess upwards of 80 percent in corporate is "useless stuff".
sanderjd 24 minutes ago | parent
JMKH42 16 minutes ago | parent
dominotw 46 minutes ago | parent
Most of the work i've done in my career has been some random shit no one cared about.
sreekanth850 42 minutes ago | parent
Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
combobyte 29 minutes ago | parent
The whole point of AI is for companies to invest less resources into employees. There's no way they start spending money on training people now, when that was already something that made executives roll their eyes before AI.
tkdb 4 minutes ago | parent
Alternate take: it takes a lot of up front effort to ensure when you "max out tokens" the work is valuable.
The same thing goes for token use operating business processes (vs. building software that runs business processes): Let's imagine buying $250k tokens per year to operating some business processes - how much human effort is needed to ensure that level of spend is valuable? I'm using a number that could be "we could hire a skilled human at that level" as that changes the feel of the question.
thadt 41 minutes ago | parent
At this point, local models have become feasible, and people using them are beginning to get a feel for the tradeoffs vs the frontier models. As the frontier advances, the question becomes “how much will I spend for a given quantity and quality of AI work?” with local hardware providing a pricing anchor point.
When I can price hardware and ops for a given capability level - I have a budget again. From there it’s a question of how much faster/capable/cheaper is a given provider (and, you know, how much do I trust sending them all my IP?).
bravetraveler 37 minutes ago | parent
01284a7e 28 minutes ago | parent
Yeah, that guy you hired to write the prototype went on a 2 month bender and created 0 usable code. Did you budget for that? Oh, okay. The strategies to deal with spending on tokens are nothing compared to the overhead of managing actual people and their outputs.
tkdb 20 minutes ago | parent
I'm glad humans are freaking out when they realize they don't know if spending $10k/month on AI tokens is good or bad.
I'm glad humans are freaking out when an AI pumps out vapid presentations and other humans go ahead and present it to clients.
But.
Too many businesses have tolerated the same mindlessness when humans were in the place of LLMs.
Too many companies telling themselves and investors headcount growth is good without knowing that the new hires are actually doing.
Too many human-slop presentations float around, with the authors and audience just going through the motions.
Did digital photography raise the bar for what constitutes commercially valuable photos? I think so.
I hope AI will similarly raise the bar across all the industries it is touching.
randusername 11 minutes ago | parent
It continues to be an uphill battle to help people at $DAYJOB understand that the relationship between turns and cost is nonlinear. And many still don't understand the idea of a system prompt, that they can control how chatty all responses are.