Hacker Newsnew | past | comments | ask | show | jobs | submit | aphroz's commentslogin

I think not much can run without a dedicated GPU


Till now I was using successfully Qwen 3.5 and Gemma 4 at a reasonable speed


I think (maybe I missed something) that identical size and quant versions of Qwen 3.5 and 3.8 should run at the same speed. It’s the exact same architecture.


Tried the 4 Bit versions. It loads, bit the <thinking> output isn't even coming out.


There is no way you would run a dense 27b model on that spec. I ran 3.6 27b on a 64gb ram, 24 gb vram, and it felt like the lower limit for this model with a decent context window.

If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.


The issue was caused by Ollama. I've tried again with llama.cpp and mounting the igpu device correctly, I get 18-23 tk/s on the Iris Xe card under Debian, MUCH better than before.


From my experience if it doesn't fit on vram it is rarely worth to bother except for a few narrow tasks.

For example make an essay about something where you don't actively engage with the LLM after the initial prompt. So mostly one-shot prompts.


were you running the MoE models? those perform better speed wise


What's the story with Mac laptops? Worth a try?


The author tested in on an M5 laptop too:

> It feels pretty slow on both the M5 Mac and the DGX Spark.


Dense ones like this are more bandwidth-hungry, so you want to try MoE ones like Qwen3.6-35B-A3B (35 Billion params but only 3 Billion Active) or Gemma 4. Unfortunately it seems like we might not be getting a 3.8 MoE.


Yes. mtplx runs it at 25 tok/sec on a M4 Max with 48GB RAM.


And the resource you need to host so much data


Please add to your prompt: "Make it concise" and "Do not use '—'"


lol, I'm not familiar with English.


That's why I want to help :)


So the founder of Coursera is starting a new venture backed by Coursera ? That sounds like the Musk playbook.

I don't really get how this is new or even "AI". Can't you already ask Claude to teach you something ?

"Patiently stays with you until you've mastered new skills", that's so nice. Maybe in the future it will be the other way around, your computer will complain about how slow the user is :)


Except that it is also quite difficult to assess the quality of a doctor or a software developer if you don't work in the field.

I've heard numerous cases where AI solved medical issues that doctor couldn't.


You mean LTT ?


We type two capital LLs a lot these days.


LoL


yes thx


With AI agents assuming roles previously held by humans, it becomes necessary to provide them with guidelines for human task delegation that avoid psychological harm and minimize resistance


I still use it daily, the business version was swallowed by Teams but I hope Skype will survive, not as bloated and just does the job.


The business version was nothing to do with the real Skype though. It was just a lame rebranding of Lync. Just some Microsoft branding BS.

If it still existed now it would have been called copilot video or something :)


If they don't know, now they know


The Notorious B.I.G. and all his homies hate WordPress


WordPress, HTML, CSS, When I was dead broke, man, I couldn't picture this


"Frequent updates" is something that I would like to avoid in a CRM. Unless you have a team dedicated to update, test, migrate and fix.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: