I think (maybe I missed something) that identical size and quant versions of Qwen 3.5 and 3.8 should run at the same speed. It’s the exact same architecture.
There is no way you would run a dense 27b model on that spec. I ran 3.6 27b on a 64gb ram, 24 gb vram, and it felt like the lower limit for this model with a decent context window.
If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.
The issue was caused by Ollama. I've tried again with llama.cpp and mounting the igpu device correctly, I get 18-23 tk/s on the Iris Xe card under Debian, MUCH better than before.
Dense ones like this are more bandwidth-hungry, so you want to try MoE ones like Qwen3.6-35B-A3B (35 Billion params but only 3 Billion Active) or Gemma 4. Unfortunately it seems like we might not be getting a 3.8 MoE.
So the founder of Coursera is starting a new venture backed by Coursera ? That sounds like the Musk playbook.
I don't really get how this is new or even "AI". Can't you already ask Claude to teach you something ?
"Patiently stays with you until you've mastered new skills", that's so nice. Maybe in the future it will be the other way around, your computer will complain about how slow the user is :)
With AI agents assuming roles previously held by humans, it becomes necessary to provide them with guidelines for human task delegation that avoid psychological harm and minimize resistance