Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
 help



A lot of llama.cpp contributions come from the community and ecosystem, like Unsloth. If something goes awry, I fully expect lots of forks.

There already are a lot of forks for things they decline to implement. TurboQuant, ROCmFPX, and more. I need to set up an agent that will loop on merging them.

Except when they have less than 16 gb of ram?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: