That speed is significantly lower that what can be attained with streaming from SSDs.
Not only in big desktops, but even in most recent mini-PCs, it is possible to read concurrently from one PCIe 5.0 SSDs and one PCIe 4.0 SSD, at a total sustained reading throughput of around 20 GB/s.
With an optimized inference implementation, it should be possible to overlap completely the computations with streaming weights from SSDs.
This should improve the inference speed to around at least 1 token per second, on a cheap computer, under $2000 even at the current super-inflated prices.
There are enough tasks where this would be useful. Obviously one should use for most tasks a fast small LLM and use the big one only when this actually saves time.
I'd rather not be thinking about tomorrow's meeting all night! Just trust "the robot has got this"... and scramble in the morning when the robot has failed. But at least I got a good night's sleep!
For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.
https://github.com/gavamedia/deltafin