Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Most of the above infra is predicated on limiting RAM so that you need so much communication between cards. Bump the RAM up and you could do single card inference and all those connections become overhead that could have gone to more ram. For training there is an argument still, but even there the more RAM you have the less all that connectivity gains you. RAM has been used to sell cards and servers for a long time now, it is time to open the floodgates.


Correct for inference - the main use of the interconnect is RDMA requests between GPUs to fit models that wouldn't otherwise fit.

Not really correct for training - training has a lot of all-to-all problems, so hierarchical reduction is useful but doesn't really solve the incast problem - Nvlink _bandwidth_ is less of an issue than perhaps the SHARP functions in the NVLink switch ASICs.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: