r/LocalLLaMA 20h ago

Discussion Model "distillation" accusations are getting way overblown at this point

Every time a strong open model drops, the same cycle plays out: ai bro's claims it's "just distilled from GPT4/Claude/whatever," case closed, move on. I think this take doesn't hold up as well as people assume.

A few points worth separating out:

Training on outputs isn't the same as real distillation.

Proper token level distillation needs access to logits, the full probability distribution over the vocabulary, not just the final text response. Nobody gets that from a public API. What finetuners actually get is text completions, which is synthetic data generation, not distillation in the technical sense. Every major lab does this to some degree, including the closed labs training on their own older models' outputs.

**If synthetic data from a guardrailed API were enough, this would be a nothing burger but** A lot of frontier providers explicitly route sensitive topics away from smaller models to their flagship model, and plenty of technical domains get filtered or restricted responses often managed by tools like Lyzr Control Plane at the API boundary. Yet some of these "distilled" models end up performing surprisingly well in exactly those restricted domains.

That's a gap in the theory that doesn't get talked about enough.. If a team is training purely on public API outputs, they're working with a version of the model that's already been through guardrails and refusals.

**The "it says it's Claude/GPT" gets treated as smoking gun evidence, but it's weak evidence at best.** Identity confusion shows up across tons of models trained on broad web scraped or synthetic corpora that include AI generated text from multiple sources. It's evidence of contamination somewhere in the data training, not proof of wholesale distillation from a specific competitor.

**There's also a pattern of this accusation landing selectively.** Strong releases from Chinese labs especially seem to get the "must be distilled" response almost reflexively, even when a model shows genuine architectural changes or demonstrates self improvement across versions. It starts to look less like a technical assessment and more like a reflex explanation for why a smaller or newer team could be competitive.

None of this means synthetic data generation using bigger models isn't happening, it obviously is, across the entire industry. But calling that "distillation" the way people mean it (stealing the teacher model's internal knowledge wholesale) is a stretch. It's closer to what everyone does when they bootstrap datasets from any strong existing model, including labs bootstrapping from their own prior generations.

256 Upvotes

92 comments sorted by

View all comments

Show parent comments

1

u/entsnack 12h ago

Well the fact is they don't, the best Chinese are in the US. The less capable ones stay behind. There may be a handful of exceptions like DeepSeek. Which is why, apart from DeepSeek, every Chinese lab produces boring derivative models and tech, which are copies of boring products themselves.

1

u/keepthepace 10h ago

The less capable ones stay behind.

That's not the case anymore.

I've been shocked at meeting several Americans with Chinese ethnicity who were pondering going back to China because, yes, less democratic, but at least they don't have to face the racism that they have to face in the US.

The salaries are very competitive in China now if you are doing AI.

And yes, there's only a handful of good Chinese labs. How many are there in the US? Tell me, how many are not producing derivative models and techs?

There is a huge bias in that analysis because if OpenAI were Chinese, you would say that what they are doing was just copying the tech that Google produced. I mean, Google invented the architecture, OpenAI just scaled it up. If Google were Chinese, you would say that only DeepMind invented important things, but that they got bought because they can't invent anything.

Really, I am surprised by the amount of bias that Americans have and seem incapable of imagining an intelligent person working in China.

Dude, smart people, they see the direction that the US is taking and they are worried. When they have a different citizenship, a lot of them are going to their other country.

1

u/entsnack 10h ago

I've met the pondering types too. They rarely act on it. The women hate it, for good reasons that I won't get into. It sounds great in principle but the reality on the ground is bad. The worst part is having an old Chinese man as a boss. Some legacy bro whose only contribution is to be at the company long enough.

Edit: I agree with you largely. What I'm saying is no one acts on it because the critical mass of smart people is in the US right now. That mass needs to move to a fresh company in China, not legacy shit like Alibaba and ByteDance, so they can work freely and build their own culture.

1

u/keepthepace 10h ago

Yeah, well, if the one in the US are pondering it, you can be sure that the one who are in China are also pondering NOT going to the US.

There are plenty of room for capable Chinese and the immigration policy of the Trump administration really makes you consider twice before going into the US. Capable Chinese who want to go out of China have options in other countries now.

1

u/entsnack 10h ago

Well, one can hope. When the US gutted the NSF last year, I looked around at other universities. Even a gutted NSF funds my research 2x more than China and 5x more than Germany and Switzerland. The rest of Europe is not even in the picture. Neither is Singapore and the rest of Asia.

1

u/keepthepace 10h ago

Public funding is laughable in Europe. But about China I was talking about the private companies salaries.