69 points tomncooper 2 days ago 17 comments
dominotw 2 days ago | parent
segmondy 2 days ago | parent
LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.
traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.
decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3
lostmsu 2 days ago | parent
IanCal 1 hour ago | parent
Jev is 4.2c/m tokens in and free out.
ShinTakuya 1 hour ago | parent
- https://www.ml6.eu/en/blog/jev-vs-gpt-6-luna-vs-bert-text-cl... - https://tessl.io/blog/jev-is-136x-faster-and-27x-cheaper-tha... - https://x.com/fazxes/status/2100300097695232164 (this last one is Luna 5.6 but that isn't too different from 6 besides accuracy and cost)
AnthusAI 2 days ago | parent
In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.
6thbit 2 days ago | parent
What is openai doing for their decisions API, a finetuned luna?
reexpressionist 2 days ago | parent
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.
deepsquirrelnet 2 days ago | parent
aidiveyt 1 day ago | parent
Havoc 58 minutes ago | parent
eg feed it a weather forecast and ask it whether I need an umbrella. It’s smart enough to make the connection between rain and umbrella.
So somewhere between classifier and fat LLM.
Ultimately boils down to right tool for the job
Garlef 47 minutes ago | parent
I think the abstraction is a useful one - a general purpose classifier that does not need to be specifically trained: Unstructured signal in, structured judgement out - with a focus on speed and cost efficiency.
And since there is not yet a large body of benchmarks, I don't think we have sufficiently explored how to measure these things.
But since there's a market and some hype this will soon happen.
(And it's not like TypesafeAI has a real moat or invented something entirely new here ~ they just managed to put things into one coherent perspecive)
petesergeant 46 minutes ago | parent
bjord 17 minutes ago | parent
"Benchmarking AI decision models against traditional guardrails"
deadbabe 15 minutes ago | parent