> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation
Sounds great to me; live by the sword, die by the sword.
Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs.
But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.
I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal.
The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.
If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.
You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.
That's one thing.
The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)
It's pretty common to have such laws. OpenAI can put whatever they want in their ToS, but they cannot go back and sue someone for violating those terms if the government has ruled that clause to be unenforceable.
Sam and Dario keep saying intelligence will be like water or electricity. If that's the case then it should be illegal to deny to anyone. In most parts of the world the power company can't shut off your power for an unpaid bill - they have to get a court order to allow it, which gives you a chance to defend yourself or make a payment plan.
Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.
It's been common in electronic design automation tools to have license terms like that (forbidding use to create a competing product). However, competing companies have often found workarounds, either by finding loopholes or just breaking rules and hoping not to get caught.
How the hell is non-compete legal in market economy? Competition is one of its core strengths. Why would anyone let anyone opt out of this, even a little bit?
No country in the world is full free market economy. It is always a spectrum.
We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.
Chinese companies compete ruthlessly between themselves though. That's how they get this good. Full competition with preventing exploitation by foreign countries seems to be working great for them. American and European protectionism of local rent-seekers can't really compete with that.
The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).
It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.
It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.
Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.
> Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning
"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."
I don't know why you're trying to downplay it.
European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.
>European models are so far behind because...
What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.
China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.
[1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?
> Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States.
No they didn't. Europe has played a role in building all of these. Especially from the software side.
Also anyone that doesn't like the US is free to stop using any American website, device, or service. There are alternatives to everything. For instance I do not like Meta, I refuse to use any of their websites or devices. The domains are blocked on my LAN.
> Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models
This is complete nonsense. First of all most Americans don't give a shit about AI companies like it's a sport and those are our teams we have to support. Second the hysteria around AI is mostly from non-Americans that are completely out of the AI race worried about losing access to good models. That is valid, but also it needs to be recognized.
Yes, they did. Like, look around. Clearly the Western world foolishly and very short-sightedly allowed the US the reigns on far too many things, to its disadvantage.
It is unwinding, but it turns out that having decades of intertwining takes a while to undo.
>Also anyone that doesn't like the US is free to stop using any American website, device, or service.
What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.
> What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.
You want to hate the US but don't want the personal inconvenience of moving away from the daily sites and apps you use. How are these alternatives to american tech going to get enough users if someone like yourself who seems to have a hate boner for the US can't even leave a message board?
The orange idiot can say whatever he wants. The first amendment makes any ban unconstitutional. US companies being advised about potential backdoors in the models isn't a bad thing, even if I personally think it's FUD, open weights is not open source.
> European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.
Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.
This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.
Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).
Don’t know much about how distillation works so please enlighten me here.
> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models
If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?
A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.
Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.
ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.
So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).
The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.
Governments can do anything they want by passing a new legislation. In your example, they could easily pass a law that states that any ToS cannot reject service to a customer based on the color of their attire. In the USA, it's obviously already illegal for a business to reject service to a customer based on some protected classes like race.
Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.
What I'm saying is, doesn't the law already cover 1?
Fair use requires more than you accessing the material legally.
In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.
If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.
> If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?
I mean the choice is: 1) we pass laws that explicitly say training models like this is legal (the original quote, 2) say it’s illegal and requires licenses for the data and ability to opt out, 3) we ignore it and continue because the companies are too big to jail.
This is a silly perspective, inaccurate, and out of bounds framing.
Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.
But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.
But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.
If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.
> If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.
I'm confused now; isn't the LLM that trains on that professor's lectures, videos and textbooks undercutting him?
If the student is really good at generalizing we wouldn't even be having this debate because he would've just generalized from the same source materials the professor used.
One can teach whoever they want to or don't want to. If one joins a university, that changes things. They are now part of an organization larger than themself.
if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.
He is prohibited from regurgitating source material, of course! But if he generalized from the books he read and really learned the subject--and even made new connections between ideas--then he is free to write his own book.
He is not prohibited from "regurgitating source material" in many cases. Facts are free. It doesn't matter who first measured Young's Modulus of aluminum, anyone may state that fact as originally presented.
The professor is free to lift all the facts and formula they want. They just need to rephrase explanations. Which is pretty much what an LLM is going to do.
There is value add in AI irrespective of how the data got to what it is.
Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.
'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.
Sounds great to me; live by the sword, die by the sword.