30 points felix089 9 hours ago 5 comments
We wanted to put the popular ones to the test and thought Pac-Man is a good benchmark for simple and fast decision making.
So we let jev 1.13, kev, clef, clef flash, GPT-6 Luna and Laya play Pac-Man against bot ghosts.
The low latency of these models allows for real time play. We had each model play 100 games, published a leader board and open-sourced the repo so anyone can run their own model and join the ranking. Link to repo: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...
You can also join the game and play as Pac-Man yourself, and the ghosts are the models, either a mix of models or all jev, kev, clef etc. A game costs about 2 cent, all models are running via my startup opper, and we added free credits for everyone to try.
It's pretty fun to play and surprisingly difficult to beat jev's highscore. Any feedback is more than welcome!
gsandahl 5 hours ago | parent
ozozozd 1 hour ago | parent
I wish the controls were a little easier on mobile.
NichoPaolucci 38 minutes ago | parent
Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.
nico 15 minutes ago | parent
I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy
I’ll try training one for this benchmark, seems like fun
bayarearefugee 11 minutes ago | parent
What even is this reality.