Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

More than that, running hundreds of conversation streams at once is essentially the same cost as running a single conversation. And then you add on the secondary benefit of having the GPUs running nearly all the time rather than mostly idle...

Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: