88 points chiefstorm 2 hours ago 32 comments
Imustaskforhelp 2 hours ago | parent
Interesting to see where all this leads us and if other major labs follow suit
Edit: decisions voice looks really interesting as well[1]
[0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev
[1]: https://developers.openai.com/api/docs/guides/decisions-voic...
esafak 1 hour ago | parent
peterson_lock 1 hour ago | parent
mritchie712 1 hour ago | parent
mohsen1 1 hour ago | parent
binlog 1 hour ago | parent
dvt 1 hour ago | parent
Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.
mediaman 1 hour ago | parent
You could ask the same question about why anyone would rent a VPS. I can just run my own hardware, it's just a computer!
Buy vs rent is not just about what's possible, it's about what's economic.
dvt 1 hour ago | parent
TSiege 1 hour ago | parent
Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.
simonw 1 hour ago | parent
Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.
tmhall 1 hour ago | parent
super256 1 hour ago | parent
There are probably a lot more reasons.
jcims 1 hour ago | parent
Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.
csharpminor 1 hour ago | parent
drdexebtjl 19 minutes ago | parent
TSiege 1 hour ago | parent
Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.
If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
gobdovan 1 hour ago | parent
Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.
tripleee 46 minutes ago | parent
st3fan 30 minutes ago | parent
woah 12 minutes ago | parent
simonw 1 hour ago | parent
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $(llm keys get openai)" \
-H "Content-Type: application/json" \
--data '
{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "I am angry about the new product feature"}
]
}],
"questions": [{
"type": "predicate",
"name": "complaint",
"instructions": "Is this a complaint?"
}, {
"type": "predicate",
"name": "compliment",
"instructions": "Is this a compliment?"
}]
}'
Returned: {
"model": "gpt-6-luna",
"answers": [
{
"type": "predicate",
"name": "complaint",
"probability": 0.91
},
{
"type": "predicate",
"name": "compliment",
"probability": 0.06
}
],
"usage": {
"input_tokens": 310,
"input_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 0
},
"output_tokens": 0,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 310
}
}
That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)
sidcool 1 hour ago | parent
Topfi 1 hour ago | parent
Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).
Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).
Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.
Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).
Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
oh_no 57 minutes ago | parent
How these models play out is an open question but existing provider contracts and T&C are important for enterprise.
Topfi 48 minutes ago | parent
Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?
3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.
nostrebored 36 minutes ago | parent
Topfi 17 minutes ago | parent
Ran every task twice on each model. Some were for UI component synthesis and charting insanity that is a bit hard to explain, but some were simple tag selection (basic eval at threshold 60% on whether to use the provided tag given the title of a browser tile):
1. Title: "Mortgage calculator: estimate your monthly payment (Bankrate)"
Tag: "house hunting"
Jev: 0.69 (yes), 0.72 (yes)
Luna: 0.21 (no), 0.21 (no)
2. Title: "S&P 500 index: live chart and news (Bloomberg)"
Tag: "investing"
Jev: 0.91 (yes), 0.90 (yes)
Luna: 0.56 (no), 0.56 (no)
Of course, tags can be a bit subjective, but in these cases, I'd argue the values provided by Jev were far more representative of my subjective assessment over Luna's. If SnP stuff on Bloomberg isn't investing, nothing is.Goal for tagging is mainly a near instant, overwritable, sane default provided to users in the background. Resolve the whole "I love using Notion/Obsidian/PKM software of your choice but spend 80% of my time just thinking about the ideal tag before starting to read" issue. Luna's output is not really helpful here.
Shank 21 minutes ago | parent
scosman 21 minutes ago | parent