105 points jasondavies 1 hour ago 32 comments
warkdarrior 47 minutes ago | parent
kerenskiy 43 minutes ago | parent
didibus 41 minutes ago | parent
petercooper 39 minutes ago | parent
There are a few technical details that can reduce the latency significantly (covered in the post) but the real insight has been from watching the reaction to Jev and seeing that there's enough of a market interest to offer it as a distinct thing. The underlying concept/approach was already there.
theapadayo 4 minutes ago | parent
The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer.
segmondy 33 minutes ago | parent
woah 29 minutes ago | parent
ford 25 minutes ago | parent
zitterbewegung 27 minutes ago | parent
ramoz 17 minutes ago | parent
Anyone can copy that and apply to an array of models - stripped down LLMs or already slim/highly performant traditional classification architectures (just wrap inference with an api that inputs/outputs the same structured data).
Jev, I think, would say their advantage is the intelligence of their models and training data including calibration: https://medium.com/code-applied/calibrated-classifiers-makin... (which i still struggle with in the general application... there's no free lunch with these things).
233mhz 17 minutes ago | parent
If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory
porridgeraisin 15 minutes ago | parent
Getting training data that works well for calibrated classification objectives is difficult.
I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.
But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.
So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.
pizzafeelsright 13 minutes ago | parent
Many people seem to have run into the same question and started working out the answer.
giancarlostoro 13 minutes ago | parent
It seems insanely obvious at least to me, that JEV is the new hot thing for the AI field since they give you stronger output that isn't... flat out wrong, that alone is impressive.
johnecheck 34 minutes ago | parent
hbcdbff 31 minutes ago | parent
aryabakh 28 minutes ago | parent
ssiddharth 22 minutes ago | parent
DesaiAshu 21 minutes ago | parent
open592 21 minutes ago | parent
swingboy 18 minutes ago | parent
yipinwong 17 minutes ago | parent
Or are companies/people already building this based on say an arXiv docs? n
---
The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding
XCSme 9 minutes ago | parent
Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.
Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.
It's not really a new technology, it's more like a new use-case.
bityard 16 minutes ago | parent
MisterMunchkin 14 minutes ago | parent
RGS1811 2 minutes ago | parent
6thbit 9 minutes ago | parent
Perhaps that may be too costly atm
manlymuppet 5 minutes ago | parent
And it's only been a few weeks.