457 points OfficialTurkey 1 hour ago 256 comments
hehimself 1 hour ago | parent
beardsciences 1 hour ago | parent
Cu3PO42 1 hour ago | parent
EDIT: this doesn't say anything about availability on either Azure or AWS. I'm assuming it will show up later, but it would be interesting if it didn't.
potwinkle 1 hour ago | parent
Readerium 1 hour ago | parent
droidjj 1 hour ago | parent
nickandbro 1 hour ago | parent
physicallyIllfr 43 minutes ago | parent
Does anyone care about code quality anymore?
blovescoffee 39 minutes ago | parent
minimaxir 14 minutes ago | parent
pookieinc 1 hour ago | parent
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
ModelInput
Output
Price reduction
GPT‑6 Sol vs. GPT‑5.6 Sol
$4 → $2
$20 → $10
50% cheaper
GPT‑6 Luna vs. GPT‑5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
50% cheaper
thereitgoes456 1 hour ago | parent
giancarlostoro 1 hour ago | parent
Readerium 1 hour ago | parent
Should be B vs A correct?
Else it's confusing
bitmasher9 1 hour ago | parent
wyre 47 minutes ago | parent
atq2119 36 minutes ago | parent
The "paradox" is when an increase in efficiency which would decrease the use of a resource all else equal, instead indirectly causes more use.
wyre 11 minutes ago | parent
Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.
blovescoffee 47 minutes ago | parent
Shekelphile 21 minutes ago | parent
When they cut prices on luna the first time around they took (literally) millions of users from anthropic.
mchusma 59 minutes ago | parent
Its a great release, I will use both heavily.
etothet 58 minutes ago | parent
mfiguiere 57 minutes ago | parent
https://developers.openai.com/api/docs/pricing?latest-pricin...
onlyrealcuzzo 56 minutes ago | parent
This is great, but practically, I'm not going to start working on more side projects.
Perhaps in another 6-12 months I'll be fine to drop down to $20/m instead of $200.
charliegoforit 50 minutes ago | parent
wyre 46 minutes ago | parent
onlyrealcuzzo 37 minutes ago | parent
A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.
I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.
I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.
adam_arthur 14 minutes ago | parent
There are a ton of use cases that open up with cheaper models.
E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc
shmoil 52 minutes ago | parent
>> $4 → $2
>> $20 → $10
Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.
blovescoffee 46 minutes ago | parent
yzydserd 42 minutes ago | parent
jameshart 25 minutes ago | parent
edf13 50 minutes ago | parent
LZ_Khan 49 minutes ago | parent
AustinDev 47 minutes ago | parent
nradov 46 minutes ago | parent
OutOfHere 34 minutes ago | parent
trentor 15 minutes ago | parent
an0malous 46 minutes ago | parent
solenoid0937 45 minutes ago | parent
blovescoffee 36 minutes ago | parent
selectodude 33 minutes ago | parent
If they’re subsidizing my usage, that’s great.
infinitezest 19 minutes ago | parent
persedes 42 minutes ago | parent
joshstrange 40 minutes ago | parent
If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.
hombre_fatal 36 minutes ago | parent
The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.
My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.
ignoramous 33 minutes ago | parent
Did cache read/write also decrease by 50% or similar? That's where most (95%+) of the cost is for agentic coding workloads.
minimaxir 24 minutes ago | parent
mchusma 1 hour ago | parent
sfkgtbor 1 hour ago | parent
Readerium 1 hour ago | parent
Wierd!!
samuelknight 1 hour ago | parent
MrBuddyCasino 57 minutes ago | parent
o_m 46 minutes ago | parent
minimaxir 43 minutes ago | parent
cesarvarela 1 hour ago | parent
petesergeant 53 minutes ago | parent
afro88 22 minutes ago | parent
recitedropper 1 hour ago | parent
wahnfrieden 59 minutes ago | parent
dominotw 58 minutes ago | parent
rtaylorgarlock 57 minutes ago | parent
ronsor 57 minutes ago | parent
But the reason people say "Claude can't compete" is because Claude Opus has been going downhill since 4.7, and many have found Opus 5 intolerable. Fable is much better, but also much more expensive than OpenAI's offerings.
CaptWorld 54 minutes ago | parent
droidjj 58 minutes ago | parent
mydreamof 56 minutes ago | parent
saadn92 56 minutes ago | parent
qoez 52 minutes ago | parent
LeBit 47 minutes ago | parent
monkeydust 50 minutes ago | parent
Madmallard 49 minutes ago | parent
minimaxir 36 minutes ago | parent
woah 47 minutes ago | parent
sidrag22 46 minutes ago | parent
cmrdporcupine 46 minutes ago | parent
Asking cuz I don't think I'm a bot [pats self], I legitimately prefer the GPT models to Anthropic's, don't like Anthropic's customer service/reliability story at all, and I welcome a massive price reduction. Seems like something I should be happy to get.
If you'd told me I'd be typing this a year ago I'd be skeptical though.
theanonymousone 1 hour ago | parent
badatnames 1 hour ago | parent
chaos_emergent 54 minutes ago | parent
cmrdporcupine 49 minutes ago | parent
Then they do a new model launch, issue quota resets all around, and it's a party for 2-3 weeks before things return to normal.
badatnames 46 minutes ago | parent
wyre 45 minutes ago | parent
badatnames 40 minutes ago | parent
OpenAI/Anthropic meanwhile feel a bit like they're hoping to sell iPhones in a market about to be flooded by $20 flip phones, with almost no channel of their own to do it. And for whatever mad reason OpenAI are now signalling they will attempt to compete on price with flip phones despite their cost of labour, energy, and just about everything else being far higher
cmrdporcupine 58 minutes ago | parent
Which... fine, I'll take that.
jrflo 52 minutes ago | parent
m_fayer 51 minutes ago | parent
redox99 48 minutes ago | parent
cmrdporcupine 38 minutes ago | parent
Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").
And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.
But it also feels sloppier? Somehow. And too expensive to use.
We'll see how Sol 6 is.
jeffnash 30 minutes ago | parent
Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.
I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].
mavsman 28 minutes ago | parent
NorthSouthNorth 45 minutes ago | parent
nickreese 45 minutes ago | parent
mcast 43 minutes ago | parent
jdw64 40 minutes ago | parent
BowBun 38 minutes ago | parent
bradly 37 minutes ago | parent
Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.
AaronAPU 34 minutes ago | parent
apitman 9 minutes ago | parent
meerita 49 minutes ago | parent
simianparrot 48 minutes ago | parent
sehw 47 minutes ago | parent
msh 47 minutes ago | parent
blovescoffee 44 minutes ago | parent
OutOfHere 40 minutes ago | parent
Tadpole9181 43 minutes ago | parent
Terra ended up just being an awkward middle ground that was not particularly suited for any workload.
smith7018 41 minutes ago | parent
sandos 13 minutes ago | parent
Funny thing is they very recently also set a real limit per-user/month, so why even limit the models because theyre "too expensive".
miohtama 4 minutes ago | parent
devinprater 46 minutes ago | parent
markerbrod 45 minutes ago | parent
Edit: Yes, it applies also to subscriptions, source https://x.com/thsottiaux/status/2102463847714247142
fHr 41 minutes ago | parent
yipinwong 41 minutes ago | parent
Now GPT 6 Luna is even cheaper, and more intelligent, there is no going back... to SOL 5.6 for intelligent layer.
dmazin 38 minutes ago | parent
I was hoping for a serious Luna upgrade. It was already cheap enough. This feels more like a price reduction than an upgrade.
That said, if the new Luna is able to handle ultra mode and subagents v2 in codex cli, then at least that’s a win.
recitedropper 41 minutes ago | parent
My previous comment--which suggested astroturfing--was the highest upvoted comment here until it got flagged. Which implies to me that atleast the other remaining humans on this forum see it as well.
26 minutes, 89 comments, upvoted instantly to the top, posted within one hour of the Opus 5.5 announcement. You tell me.
john_strinlai 32 minutes ago | parent
fyi, i flagged it because it is boring reading and against the rules.
if you suspect astroturfing, flag the comments and contact the mods.
(complaining that your complaint got flagged is also tiresome. contact the mods. "@dang" doesnt work, use the email.)
recitedropper 29 minutes ago | parent
I respect you for replying here though, and yes I get that HN forum standards would suggest flagging my previous comment. But it is just sad to see a place used to be so vibrant get manipulated because of how much weight it holds for us in the industry.
And yea sure, I could go and flag all the bots and message Dang. But probably time to stop shouting into the void. :)
seizethecheese 19 minutes ago | parent
Also, I don’t see that much astroturfing here? (And I tend to see it a lot on HN.)
minimaxir 16 minutes ago | parent
recitedropper 9 minutes ago | parent
But do you really think people were so excited about cheaper versions of Astra that they were just waiting around to comment the instant this was posted? More than two comments per minute? All the initial comments were really similar too: brief one liners celebrating the cheap prices.
I think AI right now is a sort of Rorschach test. What it is clearly revealing to me is that I don't trust organizations with enormous financial incentives to manipulate public opinion. So I see bots everywhere. :)
minimaxir 7 minutes ago | parent
> what the fuck
john_strinlai 16 minutes ago | parent
yes, i think so.
because, unfortunately, complaining about bots (or astroturfing, or whatever) doesn't stop them. so we end up with threads that have both the potential bot/astroturfing/whatever activity and complaints, which further drowns out any interesting comments.
recitedropper 6 minutes ago | parent
BenzeneDream 4 minutes ago | parent
tomhow 16 minutes ago | parent
On the other hand, you have previously written: I'll gladly admit I think what these companies are doing is unethical, and I'm sure that biases my thinking toward skepticism. [1]
You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence. This is in breach of the guidelines, because comments like this poison discussions far more than the comments they're complaining about.
We – of course – want all comments and posts on HN to be authentic. HN is only a place where anyone wants to participate because since the beginning, we've had software mechanisms and moderation practices that detect and weed out inauthentic commenting and voting. We're identifying and dealing with it every day, continually improving the software to detect and remove it. Most of that happens quietly and efficiently in the background without anyone having to see it. When users see evidence of manipulation and report it to us via email, we happily and thoroughly investigate it.
Most of the time, what we find is simply that people are authentically excited and passionate about the topic, which is what is happening here. I understand it can be hard to accept that if you're skeptical about the topic.
It's fine to be skeptical about the topic and you're welcome to express your skeptical views on the topic. People do that every day on HN, about AI-related topics and countless others. Healthy debate is what we're here for.
But you can't keep poisoning HN, by (1) continually posting these unfounded claims, then (2) when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted. This is not what people do when they care about a forum's health.
jeffnash 40 minutes ago | parent
1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.
2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.
3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.
I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.
ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.
spijdar 35 minutes ago | parent
sidrag22 21 minutes ago | parent
And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I'm surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan/relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster/context poison.
noname120 35 minutes ago | parent
As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.
> There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing
Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.
jeffnash 26 minutes ago | parent
And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.
sodacanner 33 minutes ago | parent
Having limitless webUI ChatGPT usage is much better user experience, though. I'll give them that.
(edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)
basisword 30 minutes ago | parent
rbranson 17 minutes ago | parent
paulmist 26 minutes ago | parent
Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.
On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.
joshstrange 13 minutes ago | parent
Using Agentsview (which might have it's own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).
lifty 25 minutes ago | parent
nwienert 20 minutes ago | parent
jeffnash 9 minutes ago | parent
Is this not the default anymore? I am on the (now closed) 20x plan.
NorwegianDude 23 minutes ago | parent
kornelijus 23 minutes ago | parent
marcd35 22 minutes ago | parent
glub 20 minutes ago | parent
This hasn't been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they're not good for your mental well-being.
> Context window in the harness
Codex now allows 1M for subs with config params. But generally speaking, you shouldn't really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.
> I've subscription hopped a bunch
OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:
you can't buy a $200 sub anymore. So if you cancel, you won't be able to get back in. Hostage situation, essentially.
EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - https://nitter.xitter.cc/_can1357/status/2090075496948060372
rudedogg 9 minutes ago | parent
I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.
joshstrange 17 minutes ago | parent
Compare that to Claude and I can run multiple agents on Opus almost indefinitely. YMMV of course but I was shocked at how quickly I burned through Codex usage.
On the context window, I feel so cramped on Codex, compacting happening every time I turn around is annoying. I didn't realize how much I enjoyed the Claude context window size.
rgbrenner 8 minutes ago | parent
Makes me think they picked Codex, stopped trying Claude, and just hang on to outdated beliefs about the value they're receiving.
rgbrenner 15 minutes ago | parent
That isn't a valid comparison, since Codex 20x is closed. So we should be comparing Claud 20x to Codex 5x + credits.
Also in Codex, even though you can increase the context window to 1m so its on par with Claude, exceeding the default is billed at 2x.
elxr 12 minutes ago | parent
While you're understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it's a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious.
Also, Anthropic has zero models comparable to Luna.
bix6 7 minutes ago | parent
> Also, OpenAI is just a company I'd rather support than Anthropic.
rbranson 11 minutes ago | parent
hintymad 5 minutes ago | parent
I'm quite puzzled about why Anthropic is so hellbent on blocking other coding agent. It's not like Claude Code has any secret sauce, right? And does Anthropic make monkey off API usage, and their magic is on the model side anyway?
impulser_ 3 minutes ago | parent
GodelNumbering 40 minutes ago | parent
Someone1234 40 minutes ago | parent
I actually preferred 5.6-Terra not because it is technically superior (it isn't) but because it had better instincts to NOT do this stuff.
PS - Speaking of better instincts, have they closed the UI-design gap at all? I keep a Claude subscription just because /design produces significantly higher quality UI design/UI feedback/UI refinement than anything I've seen from OpenAI.
cmrdporcupine 35 minutes ago | parent
The GPT / Codex models have always been "overengineer" personalities. I prefer that to "I left a pile of race conditions lying around and big gaps in testing" though, which is what I was getting from Opus at times.
But yes both Astra and Sol veer on the side of paranoid. And honestly that's better for team work. For solo work where you just want to yeet something, it can be tiring.
You learn to tame the GPT "personality" on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.
msp26 40 minutes ago | parent
Incredible.
ggcr 40 minutes ago | parent
> GPT-5.6-Sol is retiring. This conversation will automatically switch to GPT-6-Sol
I don't recall OAI retiring a model so early lol. Similar arch?
scrlk 39 minutes ago | parent
> In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.
https://x.com/ArtificialAnlys/status/2102462962758033624
Given that they had to discontinue sales of the 20x Pro plan after the Astra release due to compute constraints, I wonder if 6 Sol & Luna are smaller vs their 5.6 counterparts?
dmitrygr 38 minutes ago | parent
ghoshbishakh 37 minutes ago | parent
hamburglar1 36 minutes ago | parent
OutOfHere 36 minutes ago | parent
m3kw9 34 minutes ago | parent
seatac76 33 minutes ago | parent
zaik 33 minutes ago | parent
simonw 29 minutes ago | parent
Here's GPT-6 Luna pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
And GPT-6 Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Scroll to the bottom for the GPT-6 Sol max one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For comparison, here are the pelicans I got for GPT-6 Astra: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.
Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...
The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.
jdw64 23 minutes ago | parent
dmazin 22 minutes ago | parent
Is it? It was already too cheap to meter for me. Luna 6 is actually worse on some benchmarks than 5.6. I’d have loved improved performance for 2x the price than ~equal performance for 0.5x the price.
gizmodo59 22 minutes ago | parent
krat0sprakhar 20 minutes ago | parent
jadbox 6 minutes ago | parent
saretup 22 minutes ago | parent
simonw 20 minutes ago | parent
varispeed 16 minutes ago | parent
mkotlikov 12 minutes ago | parent
norman784 9 minutes ago | parent
> GPT‑6 Luna vs. GPT‑5.6 Luna | $0.20 → $0.10 | $1.20 → $0.50 | 50% cheaper
I can read it as follows (below), meaning that GPT-5.6 is 50% cheaper.
- GPT-6 = $0.20
- GPT-5.6 = $0.10
tedsanders 8 minutes ago | parent
simonw 5 minutes ago | parent
+--------------+-------+--------------+--------------+--------+
| Model | Input | Cached input | Cache writes | Output |
+--------------+-------+--------------+--------------+--------+
| gpt-6-luna | $0.10 | $0.01 | $0.125 | $0.50 |
| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 |
+--------------+-------+--------------+--------------+--------+nanook 8 minutes ago | parent
I'm so tired of looking at benchmarks. I always look fwd to the pelicans.
simonw 5 minutes ago | parent
Cu3PO42 6 minutes ago | parent
jumploops 29 minutes ago | parent
For example, I had Fable review Astra’s output yesterday, and it found some issues and fixed them. Passing the fixes back, Astra then uncovered additional issues with Fable’s fixes (and yes, this will go on ad infinitum if you let it, but these were “real” issues).
It seems the big story here is the reduced Luna pricing. It’s a fantastic model that can handle most automation needs (though I still use the big models for day-to-day development).
thm 26 minutes ago | parent
NickHoff 24 minutes ago | parent
therealdrag0 9 minutes ago | parent
But effort is basically how much extra internal scratchpad to use and how much extra questions to ask and answer before producing a result, Exploring more hypotheses, validating consistencies, calling more tools.
If you’re happy with your token spend on Astra then keep doing what you’re doing. but if you feel the need to conserve tokens, then you can do that by switching to smaller models like Luna when the task is straight forward.
miohtama 8 minutes ago | parent
altcognito 4 minutes ago | parent
If you have something that needs to be done right, might be a bit complicated, up the model size.
You can see this in the pelicans. Big model pelicans are pretty accurate by default. Up the reasoning and only more so, but with more detail. For Astra, it is 105 lines for low, 250 lines for max reasoning.
Small model pelicans will lack the fidelity of a large model. Bits will be out of place etc. For luna, it's 90 lines for low, 150 lines for xhigh.
Additionally the amount of time taken is increased for the larger models. Luna takes 11 seconds on low, and 1:33 for xhigh. Astra is 33 seconds on low, 4 minutes on max.
And naturally, there is the cost. There's some overlap in functionality between luna xhigh and Astra low in the sense that luna really can do quite a suitable job for some tasks. But there are just some tasks that just don't make sense for Luna, even at high reasoning.
The other thing to remember is that sometimes high fidelity isn't ideal. It can lead to overdesigning. My recommendation is to commit early, commit often, and review everything you do, which we've all been doing since before LLMs right?
brazukadev 4 minutes ago | parent
ComputerGuru 17 minutes ago | parent
At least it sounds good on paper, the the graphed results do give me pause as it seems the lower cost might come from a slightly nerfed base model combined with more thinking, going by the more erratic scoring curves and the lower no thinking baseline score. I’ll have to try it out but I really hope they haven’t nerfed Luna/Sol to make this price point possible!
rgbrenner 17 minutes ago | parent
Unsure if 5x + credits is a good value compared to Claude 20x.
adamrezich 9 minutes ago | parent
GPT-6 Luna (low medium high xhigh max ultra)
GPT-6 Sol (low medium high xhigh max ultra)
GPT-6 Astra (low medium high xhigh max ultra)
And then there's a fast mode toggle for all of it, too.Not exactly a low-friction user experience!
Like are you supposed to just somehow intuit, “ah yeah, this task is definitely a GPT-6 Sol Medium task,” or something?
Is this just second nature for OpenAI employees? How are end users supposed to know how to optimally choose a model for a given task? Am I missing something completely here?
declan_roberts 8 minutes ago | parent
ElliotAndersonC 6 minutes ago | parent
apitman 6 minutes ago | parent
* Prompt caching dashboard: https://platform.openai.com/usage?usage_section=prompt-cachi...
* Adjust reasoning effort and tool availability without breaking cache
apitman 5 minutes ago | parent