117 points ilreb 5 hours ago 14 comments
minimaxir 5 hours ago | parent
I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.
alberto467 1 hour ago | parent
I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing. I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.
simonw 1 hour ago | parent
For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.
Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.
If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.
(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)
Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.
brokensegue 1 hour ago | parent
flockonus 47 minutes ago | parent
nowittyusername 38 minutes ago | parent
djoldman 33 minutes ago | parent
740M total (270M text, 170M vision, 300M audio)
onlyrealcuzzo 24 minutes ago | parent
Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.
dcl 25 minutes ago | parent
nostrebored 14 minutes ago | parent
sohamactive 21 minutes ago | parent
aabhay 19 minutes ago | parent
Nautman 18 minutes ago | parent
https://developers.google.com/edge/mediapipe/solutions/decis...
sourcecodeplz 10 minutes ago | parent
but you can use this new one and enable/disable what you don't need.
can keep only text for ex.