Funnily enough, pasting your comment straight into Jimmy leads to a... Funnily suboptimal answer that does not answer the question.
As someone else already contributed, this is driven by a Canadian startup taalas that basically makes chips that are llms, so everything is very fast but also, baked into the chip. Once this kind of stuff is a commodity in like 10 years, our world will be very, very different.
Taalas HC1 AI uses Llama 3.1 8B, but takes up a massive 53B transistors and 815mm2 on TSMC N6 (nearly at the reticle limit of 858mm2). N2 is a little less than 3x as dense (110MTr/mm2 vs 313MTr/mm2).
This chip would still be 272mm2 on N2 which is an eye-watering $30k/wafer and bigger than a 9950x or Nvidia 5070.
This just isn't feasible. Some of the latest-gen LLMs seem to have 5-10T parameters or about 1000x more. I don't know that taping out just one chip makes economic sense let alone the 300-1000 chips required for a cutting-edge model. Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips.
There are a TON of uses for an 8B parameter models on the edge, but this is WAY too big to put on the edge of anything. Something like a 10mm2 100m parameter voice model might be feasible on the edge, but only for expensive devices, but most of those are TSMC 28nm (up to 29MTr/mm2) or GF FDX22 (up to 40MTR/mm2) which would increase the AI chip to the point where it would absolutely dominate the BOM.
> Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips.
They probably have a few ideas around that. Me, personally, I'd have one main expensive chip (replaced every 10 years, or whatever), with a secondary cheap chip in front of it that gets replaced every year or so.
The secondary chip could act the way RAG does, or perhaps both chips together can act as LoRA.
Either way, 99.999% of the knowledge is static, you just need to fine-tune the weights with that remaining 0.001% knowledge, which can be done using RAG or LoRA on a much smaller (thus cheaper) disposable chip.
The better solution would be making part of the chip cluster use something like FPGA which can be reprogrammed.
Text to speech or diagnostics equipment where the core model is relatively small and never changes seems like the ideal application. You might be able to fit something in the 25-30B range in 2nm to 14A, but it would need a way to update.
Large models are simply out of the question in my opinion. If you need 400+ different chip designs, it’ll be billions of dollars to tape out before you even make the first chip.
> The better solution would be making part of the chip cluster use something like FPGA which can be reprogrammed.
I'm not sure I follow (It's late, I am tired and I haven't had my dinner yet. That's my stupid trifecta!)
The original chip has the weights, so it's literally just a bunch of on-die (read-only) memory cells. The FPGA, while you could use it for the memory cells, would be way too expensive to use as pure memory. Typically one would hook up (read-only) storage to it, so you still need that read-only chip anyway.
The FPGA is just the compute bits, but this chip has on-die weights, not just compute.
I was proposing that the they have the base weights on a primary (permanent) chip, and have a secondary (replaceable) smaller chip with weights for a specific use-case, or for fine-tuning with new knowledge/updates to the model.
The matrices can be multiplied LoRA style, applying the matrix in the secondary chip to the primary chip, resulting in up-to-date weights through which the prompt is pushed.
I'm wondering about something different. FPGA seems ideal for an AI chip because you can simply flash the latest model. The downsides are low density and low clockspeed. It seems that you can only fit 100-300M parameters in even very large FPGA, but that seems like it would be enough for most finetuning.
I'm thinking of a situation where you do the initial model calculations in hardware on the Taalas chip then hand that off to the FPGA to do the LoRA subset of calculations in hardware that can be continuously re-tuned to keep the model up-to-date. This would probably reduce throughput (or at least increase latency), but would save tons of money by allowing you to use the chips longer.
Yeah, they're clearly just starting out and just shipped their very first proof of concept. But to me, their plans seem generally reasonable https://taalas.com/the-path-to-ubiquitous-ai/, and like I wrote, if this kind of thing succeeds and could become some kind of cheaply producible commodity component, I think there's huge value in that. Alas, maybe not as a frontier model replacement, but say 10 years from now you can drop a cheap raspberry pi like device in your Lan and have a fast local engine for things like text sentiment analysis, text summarisation, voice recognition, basic vision and things like that, that would be pretty exciting to me (but maybe as you outlined, impossible in practice)
There is a reasonable kernel of an idea here, but only if you dial expectations WAY back. The 10 years speculation is just wrong though. Even in 10 years, their 8B param model isn't going to be in consumer devices.
6nm is just 7nm++ and the process will be a decade old in a few months. In the decade since, we've only had a slightly less than 3x increase in transistor density and that's including EUV, BSPD, and GAAFET (which means progress is likely going to slow down even more).
Even if we hit another 3x increase, their 815mm2 design will still be a bit over 90mm2. For comparison, the entire M5 Pro/Max CPU die is just 61.7nm.
If our current progress somehow holds (not likely), even 20 years from now the 8B model would be 30mm2. You need 30 years of dead consistent progress to get it down to an includable 10mm2.
As you can see, this doesn't make sense to invest in. As to the stuff like voice recognition or basic vision, these can often fit within 100m parameter models which would be around 10mm2 on their current 6nm design. That's doable today in custom edge computing devices.
The other possible use is cheap fallback models for AI companies. Moving to N2 and shrinking chips to 600mm2 to improve yields a bit would give about 50B parameters with 3 chips plus another FPGA-ish programmable chip for continuing training and interconnects for everything. You'd need hundreds of thousands of chips produced for that exact AI model just to get costs below $100,000 per board.
That seems like a lot of money for the AI model you are essentially giving away, but maybe it still beats the power and price of GPU server racks.
The government isn't going to be making chip fabs go any faster which is the biggest limitation here.
The second big issue is that it takes months to fab chips meaning your hardware AI is months to maybe a year or more behind the times when it lands.
I do think it makes sense for something like a medical scanner where the model simply doesn't need constant updates, but that doesn't need government involvement to ship.
I don't understand the problem here. If you cannot focus, just focus more?
More structure/checklist to force you to focus will have other side-effects like you found out. When you get rid of the structure, you still need to have a rough map in your mind of where you want to go.
To me, this is similar to being honest. You don't want to depend on a system or checklist for being honest. It is something you always need, as a policy. Focus is like that. If you want to focus seriously on something, just make it a policy, and don't use all these tricks as crutches.
You are right, so I am probably completely on the wrong here. :D
A question to advance the discussion. What I am wondering is, if you can remember to go back to your time tracking system, why can you not remember to go back to your main goals?
Well the truth is we forget the time tracking system too. The solution is to keep the system in your face.
Maybe a programming/assembly analogy can help you understand the issue.
In my case, ADHD makes my brain want to work in a parallel way.
While I'm busy with task A it's like HEY CONDSIDER TASK B. Did you see C?
On good days, we see that and say NO OP - BUSY WITH TASK A. And refocus our mind.
Say it's a bad day...
Instead of CONSIDER TASK B, it's more like GOTO TASK B. And here it's equally harmful.
What should happen is the registers (context) of the CPU (brain) should be saved when task A stops. Likewise before task B starts it should be fully loaded into the brain.
None of these things happen for us.
So task A is left in an unfinished state, the context to finish it dissolved into thin air. Task B is started without properly being prepared which negatively impacts efficiency and performance.
And the moment the going gets hard, dopamine release decreased, you can feel it coming...
INTTERUPT - GOTO TASK C.
So it's managing that that's hard. Writing things down helps a lot, but good luck remembering to write :D
Tangentially related, but after the mindless push in my company for more AI use at any cost, every morning I drive to work thinking to myself if today should be my last day at my job.
One reason I am not giving my two-week notice is that I don't like "difficult conversations" with my manager.
The human brain is not made for multi-tasking. Multi-tasking will always be a productivity and focus hit. This was the point of the article.
> To me, this is similar to being honest. You don't want to depend on a system or checklist for being honest. It is something you always need, as a policy. Focus is like that. If you want to focus seriously on something, just make it a policy, and don't use all these tricks as crutches.
I found that I need a lot of guardrails and "crutches" as you said to be at peak productivity. Maybe something is wrong with me, or maybe it's the Dunning-Kruger effect.
- Better formatting for text: (1) bullet points (2) markdown-like links (3) Slightly different background for code.
- More "sub-reddits". We already have Ask/Show HN. We probably can add a couple more to keep everything organized.
- Option to auto-collapse comments threads deeper than X levels by default. When they are all open by default like today, only the top comment and its children get more of the eyeballs.
The formatting of text is a pain that I haven't figured out a good solution to. To do it right, I'd have to convert the existing HTML to markdown, then convert back to HTML.
It would be nice if HN just put the unstyled text in the page and then used JS to render it, but I'm sure there would be complaints about that too.
2024: I used to split time between IntelliJ IDEA (10% - for Java) and VSCode (90% - for everything else).
2025: Stopped writing so much Java, so used VSCode exclusively for Python, TS etc, with Claude Code or Cline.
2026: Time is split between Codex App (40%), Claude App (30%) and VSCode with Claude Code (30%).
Some other thoughts:
* Overall I feel like opening an IDE in the traditional sense is coming to an end.
* Tech-stack wise I am much more open to trying out new things than before since LLMs will help with the setup and debugging.
* For small teams like ours, code reviews are the bottleneck, and we constantly have to decide what code we review vs what we don't.
* Building seems easy these days, but (1) so much competition no in every field, (2) much more product polish is expected than before, and (3) most products compete with Claude if they realize this or not.
Three of us friends are working on "Data Engineer in a box" or "Cursor for Data": https://getnile.ai
Our thesis is that a lot of Data Engineering practiced today is non-differentiated across companies and they'd rather spend that time on differentiated tasks. What if we can abstract away a lot of the tedious parts of DE work and let companies focus on just the data they want to store, and the questions they have on their data?
So today, we have an MVP that manages the compute, storage, lineage, versioning and pipeline building for datasets. Would LOVE to get feedback on this or your early thoughts!
I don't have a midi piano. Wondering if it is easy for you to support my inputs using microphone. I think there will be several noobs like me with the same problem.
2. lol, why is this $230