Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

llama.cpp supports splitting work across multiple nodes on a network already

it essentially just copies a chunk of the model to each one, works well for situations where each machine has limited vram



Any pointers to RTFM / llama repo for that ? I could not find anything on a cursory look. Thanks in advance !





Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: