Which hyperscalers are doing that? I've heard of one startup trying this, XFRA, currently in early pilot phases. But I'd be pretty surprised to hear that AWS or Google are paying people to host GPU clusters in their home.
Yeah, in essence. This is actually a pretty cool part of working in Lean. It's a somewhat normal convention to write something in a human readable way and then write a second optimized implementation with some kindness of correctness theorem connecting them. There was a whole open "competition" for writing a faster Lean kernel/proof checker that didn't sacrifice on soundness called Lean Kernel Arena. Fun reference point: https://kim-em.github.io/blog/2026-7-24-why-lean-is-faster-t...
I'd wager that the vast majority of the ML research community, especially anyone interested in "AGI", is familiar with Hofstader's work. And I don't think anyone working on contemporary language models would argue that they are somehow an assumption-less "pure" model--the particular inductive bias of the Transformer has been studied by a huge number of researchers and continues to be, and the same is true for things like training data bias.
I think the Hofstader's view of modern LLMs is actually a deeply human and touching one. Looking at his work over the years, his curiosity has always veered towards human thought. He could have written GEB with a focus on completely different examples of self-reference, but he chose three striking humans from history. When he's describing modern systems as "empty intelligence", I think there's a little bit of heartbreak in his perspective, because he sees them as fundamentally different from humans in a way that leaves the part he loves--the "I" in the loop--out of the equation. He gave an interview a few years ago where he explains his feeling as being "diminished" not in a "What will I do if I'm not the best at math?" kind of way, but more specifically as he puts it, that humans are "imperfect, flawed structures".
There is no reason to believe that the transformers couldn't be doing something close to what copycat does (especially with thinking tokens), as an emergent phenomenon of the sheer size of the corpus. The architecture is certainly capable of encoding the actions in copycat anyways.
Love seeing projects like this. The performance benchmarks are nice to see. Have you done any benchmarks against approaches like OpenAI's Symphony for things like token usage or task completion?
I support and use open models as much as possible, but I'm not totally convinced that OAI or Anthropic have no moat, even as open models catch up to the frontier. Serving and inference are still hard problems when you're talking about a 2 trillion parameter model. Fine-tuning, if that remains a realistic need for businesses, is also a difficult infra problem at that scale. In the most bearish case, where there is no competitive advantage to using their models, big labs still have an advantage in this area.
Maybe there is some threshold where the price/quality math for your standard business tips in favor of smaller models and self-hosting the entire stack. I'd certainly love that.
OK, but somehow there won't be companies who will sell you appropriate hardware and a turnkey system to serve inference? Or companies that will help you fine tune popular models?
It's not about money, it's about control. Companies have lots of money and want control over their key technology.
There will be companies doing this. I'm saying the labs are well positioned to be those companies, as they effectively are those companies right now.
The same dynamics that define the public cloud ecosystem are at play here. What AWS sells you is access to appropriate hardware and turnkey infra for your needs. Looking at the cloud industry over the last 20 years, I find it hard to believe that it is impossible to build a moat or a huge business around this.
I think this is happening right now. It's beyond ridiculous and many seem to be praising 5.6 Sol and others, even Grok 4.6 seems to be a breath of fresh air
I bet many people who use Claude at their workplaces still call it "chatgpt" because chatgpt is getting the kleenex treatment and they have no idea what Claude is
I haven't read this in depth yet, though I plan to. If this general line of research is interesting to you, I'd recommend checking out some of the lines of research it touches upon--they're really rich and fascinating, and some are pretty approachable mathematically even if ML research papers aren't usually your thing. The related works section here seems pretty well stocked, but mechanistic interpretability is a pretty interesting peephole into this general vein: https://transformer-circuits.pub/
It gets extremely blurry, because people commonly refer to any model that uses a component associated with the Transformer architecture as a Transformer (i.e. using some kind of QKV-esque attention mechanism). I think it's easier to think of it like this:
A large language model is just what it says--a very large statistical model trained for language tasks. This covers the spectrum of GPT-style models, but also those hard to classify ones, like Liquid's "Liquid Foundation Models", which can get up to 24 billion parameters and use grouped query attention, but are closely related to state-space models as well: https://huggingface.co/LiquidAI/LFM2-24B-A2B
Also, as others have pointed out, a Transformer isn't inherently a language model. So really they're sort of two different axes, one classifying the model size and task, the other referring to a specific architecture.
I dunno. You can argue over whether they're overpaying, but it's not like Huggingface is Clinkle. They hit $150 million in ARR this year, they have tons of runway, and according to reports, have just started to even burn the money they raised a few years ago.
I get that it's fun to be glib about the stupidity of tech elites and investors in general, but Huggingface have been pretty open about their financials and are, in my opinion as a practitioner in the field, one of the most responsible orgs in our space. They've been a pillar of open source ML for years now and have made a very positive impact on our ecosystem.
Nvidia is getting a real business generating revenue, and the center of the universe for open models. Both seem like pretty valuable attributes, from Nvidia's perspective.
Solid point. I'm sure Nvidia went into this deal expecting completely flat growth and no other benefits to their core business. Sorta like how Meta never increased Instagram's revenue from $0 and is still waiting for it to pay off that billion dollar acquisition price.
Or like GitHub, which was generating something like 200 million in ARR and had never hit profitability when Microsoft bought it for $7.5 billion back in 2018. I'm sure it has come as nothing but a happy surprise to Microsoft that GitHub generated $1 billion in 2023. They had initially penciled it in for 38 years til ROI.
With all the spending and valuations in the AI space being so reasonable and grounded in reality, sure. Huggingface will surely be exactly like GitHub and Instagram with their extremely low and stable per-user overhead, and large, extremely dedicated user bases.
I just feel like you're saying things based on a general vibe about "AI companies" but didn't pause to look up the particular company being discussed.
Huggingface reached profitability 2 years ago. They've reportedly just started to touch the cash they raised 3 years ago. They have a tiered pricing model with metered pricing on resource heavy services. Seems pretty stable?
Your point about their user base is even weirder. They play essentially the same role for the ML community that GitHub does for software engineers, so I mean, yeah I'd pretty much expect their relationship with users to be "exactly like" GitHub's for the most part? They're where everyone has published their models for the last 6 years at least, going back to pre ChatGPT and the recent AI boom. There aren't realistically any other major model hubs, certainly not with anything near their footprint.
You’re right. I didn’t really dig into huggingface specifically even though they’re obviously different from the examples I used in my mental model. I’m not going to say I’m wrong because I haven’t personally looked into it, and regardless of HF’s position in it, this industry regularly squeezes out enough bullshit to smother an active volcano. That said, I was clearly speaking glibly using likely flawed references.
Basic napkin math: 5% IRR means they'd only need to 4.5x their revenue to make this roughly work. If they can finance this cheaper and/or do not have better options for their cash, it's even less.
It's not too dissimilar from GitHub, but geared towards ML. They have a 9/mo pro plan for individual users for upgraded storage/usage, and an enterprise version of Hub that larger orgs can pay for. I think the enterprise has some contract minimum + 50/mo per seat. https://huggingface.co/pro
They also have inference endpoints with metered prices, and their spaces product (though i'd imagine this is a smaller portion of revenue).
I’d be curious to see how much that’s subsidized. Maybe they’re different, but when I see ML and monthly plan in the same sentence, I see zero sustainability.
Why I don’t think it’s sustainable to run a compute-intensive business on low monthly plans? If they’re like OpenAI and Anthropic, monthly plan users can spend many times the amount of money in a month than they pay, and unlike regular lower-compute SaaS businesses, overhead increases significantly with usage. That means the more users many of these services get, the more money they lose. Some surmise OpenAI was hiding the actual cost in marketing expenses to make their business look less unprofitable. Anthropic was smart enough to focus on customers most likely to be willing to pay for API pricing, so they’re in a better position. If hugging face is primarily focused on selling monthly accounts rather than token based billing which bills more as the company’s expenses increase, its not likely to ever be sustainable without significant changes.
I think maybe just click around Huggingface's site a bit? They don't run a compute-intensive business on low monthly plans. You're describing frontier labs that sell access to their enormous models with subsidized subscriptions, but that's just an entirely different company/model than Huggingface. They've been around in their current form since around 2018. They provide a GitHub like service for hosting and sharing models primarily, and they also provide infra for optimized compute (for training and inference) that you can purchase through them, but you pay as you go for the compute and they have their premium baked into the price.
Most of HF's revenue comes from their enterprise customers; you essentially get direct access to their MLEs (Slack) as well as their software stack, hence the per-seat cost and annual minimum contract price. In that sense they are not particularly "compute-heavy" in terms of what they "really" sell.
Absolutely none of this makes sense and the thing that amazes me is how long it has continued. Future historians will just laugh at how stupid and obvious the crash was.
Yeah, great point. Nvidia could potentially be overpaying. Not sure how that equates to Huggingface being "a file download mirror with a couple of side features dangling off".
While it's probably easier to say this in retrospect, they were eliminating a direct competitor to their core business. Something so advantageous it should have been blocked by regulators.
I can't see this as being as good of a purchase, especially when it's 13x the price of what was seen as an absurdly large amount back then.
Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020.
If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute
Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.
reply