Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm surprised it is that low. Are not all top AI labs "cheating" and workaround LLMs's low sample efficiency by hiring people to generate more data points - similar problems with answers, so they can train models on those and improve scores? A good benchmark for general intelligence probably should be a complete black box, no sample data given/leaked at all.


Oh it's because the bench is lying. You need to pass each level without failing, if you fail a level, it count as you "lost" the minigame. The fact it start to get a score means it managed to get a 100% score on one of the minigame.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: