Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That doesn't make any sense, because their 8B is listed as benchmarking above the 13B "model A".


That's why it is very likely it has seen more tokens during training and why the plot is worthless.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: