Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: