Hacker Newsnew | past | comments | ask | show | jobs | submit | somnial's commentslogin

throughput scales superlinearly with number of GPUs when networked well and deployed with wideEP, so 1x won't compare.

also it would be interesting to figure from the DSpark paper whether their numbers are consistent with the GPUs still being H800s, since they never actually say...


this is a blog post from a company that hosts open weights LLMs (https://www.doubleword.ai/). I think its possible it might have been tongue in cheek


https://fergusfinn.com/blog/economics-of-speculative-decodin...

good point tho - plus for Deepseek the shared expert increases the overlap slightly


true, but no reason the predictor model couldn't use linear attention (i.e. mamba, GDN etc) to predict KV caches


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: