60 points terryds 1 hour ago 34 comments
pelagicAustral 37 minutes ago | parent
TomGarden 36 minutes ago | parent
pelagicAustral 34 minutes ago | parent
TomGarden 32 minutes ago | parent
pelagicAustral 27 minutes ago | parent
Tepix 37 minutes ago | parent
It would be handy to have a site like this one that takes into account the various deals and attempts to calculate the number of tokens per monthly fee for a chosen model. I realize this makes the task a lot more difficult.
2. It would be handy to have a chart like that for the AI hardware that people own. It helps you decide which model to run (resulting in different levels of intelligence and speed). Also difficult to please everyone (preprocessing vs token generation for example) and to keep updated!
I found https://llm-list.com/ yesterday and when I had a detailed look, I quickly found outdated entries, for example looking at GLM 5.3 flash it listed several providers as "free" that weren't free any longer.
TomGarden 36 minutes ago | parent
jwolfe 36 minutes ago | parent
hsnewman 35 minutes ago | parent
mrngld 14 minutes ago | parent
And if you want Astra/Fable/Opus frontier level, then there's no option at all.
But if you don't need that, or you don't need speed... That opens up the discussion. I've been impressed even with how Siri's been doing with the Apple Foundation Models in MacOS/iOS 27 given how small they are.
Edit: I can't even fully spec the M5 Ultra Mac Studio you'd need for GLM5.3 Flash since 512GB isn't available yet, but it's already at $9500 for 256GB RAM.
aslkalska 32 minutes ago | parent
jeremysalwen 31 minutes ago | parent
shakow 28 minutes ago | parent
SJMG 20 minutes ago | parent
noumenon1111 10 minutes ago | parent
npongratz 10 minutes ago | parent
floppyd 27 minutes ago | parent
For example I've been really enjoying Deepseek v4.1 Flash, it's very "straightforward" to the point of being almost dumb sometimes, but it's absolutely relentless and would solve almost any problem no matter how inefficient the solution is.
No idea how to measure all that, just average CoT length per task is probably a good approximation for some things, but not others.
jrflo 24 minutes ago | parent
lowercased 21 minutes ago | parent
I've had more than a few people tell me "oh, it's so much cheaper to use a $20 claude account" or "i've never hit a limit ever using my openai". Inevitably.. I end up reading/hearing "oh, I need to give it another couple hours to start using it again"... I've never hit that with my approach, even if it's costing me a bit more. Being able to work when I want when I have time has some value.
I also have openai and anthropic direct API billing set up for hosted and client projects that need to call out to an LLM service.
the__alchemist 18 minutes ago | parent
Is it worth the ~10x extra cost over the subscriptions? (This is obviously a leading question). Also, I think you can use OpenAI's subcription login with Air, but not Claude's.
jrflo 17 minutes ago | parent
mmmattt 14 minutes ago | parent
LeBit 12 minutes ago | parent
Otherwise DeepSeek Flash 4.1 is dirt cheap (other "Flash" models are not that expensive either). I pay (very few dollars) out of my own pocket.
There are many things where having an API Key is necessary.
Maybe I’ve missed the boat though: is there now a method to use an api key to access a subscription?
jrflo 5 minutes ago | parent
ford 18 minutes ago | parent
Ex. this type of price estimation is quite naive - some models can require 2-3x the number of tokens to achieve the same level of intelligence. Artificial Analysis' own cost per task is a more fair estimation of cost.
Xeoncross 13 minutes ago | parent
Depending on your memory, you'll need to use the weaker Q4 versions but they still perform well.
It ranks higher than GPT-5.3 Codex (xhigh) or Claude Opus 4.6 (max) so is great for pairing with https://github.com/kunchenguid/gnhf for nightly experimentation, cleanup, or recommendation lists for in the morning.
newsy-combi 12 minutes ago | parent
verytrivial 12 minutes ago | parent