Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
lolinder
on Nov 30, 2024
|
parent
|
context
|
favorite
| on:
DeepThought-8B: A small, capable reasoning model
That doesn't make any sense, because their 8B is listed as benchmarking above the 13B "model A".
sigmoid10
on Dec 2, 2024
[–]
That's why it is very likely it has seen more tokens during training and why the plot is worthless.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: