> We discovered that Moonshot AI, the company that produces the Kimi family of models, silently forwarded customer requests to Claude, instead of processing them using Kimi. Moonshot then displayed Claude’s responses to users. These users thought they were using a Kimi model, but received responses from Claude instead.
> DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers.
> MiniMax built its own proxy network service through a shell company. This shell company has no obvious links to MiniMax and does not disclose its relationship to its parent company. This shell proxy network service only offers access to models developed by Anthropic and OpenAI. The service does not offer access to any Chinese models, including Minimax’s own.
DeepSeek, minimax and so on have razer thin margins but unlike openai and Anthropic they are actually making some profit. Doing this doesn't make any financial sense.
Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?
Also, how can Anthropic have such accurate information about state actors and cybercriminals? This is the same company that hacked itself and realised that first months later..
My understanding of what Anthropic are saying about this is that the labs in question aren't forwarding things to Claude to make money nor even to look better to the customers whose queries they forward to Claude but to get access to conversations between real users and Claude, which they can then use to help train their own models.
(I do not guarantee that I'm understanding right, and still less do I guarantee that what Anthropic say is actually true.)
it's not too far fetched, for example when Deepseek came out with their new caching techniques where they were able to offer those insane discounts, it was only available through their API which would retain and train on your prompts
so, they've been on the record, and very open about it, at least for some of the labs.
Maybe there is some truth in that reselling Claude subscriptions/trials/api bundles via third parties breaks Anthropic's ToS. The rest is putting a maximum spin on it in order to achieve the political goal of banning Chinese AI. Anthropic is a highly ideological company and they are convinced that they are just in what they pursuit.
Ah, that makes way more sense than Anthropic's (probably deliberately misleading) insinuation that Moonshot has been burning millions of dollars in Claude API credits by swapping in a slightly better but infinitely more expensive model just to trick their users.
I get those A/B responses chatting in Gemini fairly often, and I really don't think I'd feel deceived if I later learned one of the choices was actually from a competitor's model.
I don’t think it was misleading, deliberately or otherwise. Did you read the report? I hate to call you out like that but I think you can only get that impression if you only read the above quotes. That’s not the insinuation I get at all. It’s specifically under the “illicit distillation” category. It’s never framed in anyway but as a form of distillation.
I think they are pretty fair and explicitly say “Distillation itself is a legitimate training method […] Distillation is commonly used because it reduces the resources needed to achieve more advanced capabilities”. And go on to say their definition that makes it illicit in these cases.
And, also, they almost certainly __were__ tricking users and sending their data overseas.
They mean distillation is legitimate when labs use one of their own stronger models to train a smaller one. They certainly aren’t advocating for PRC labs to distill Claude for open weight models.
I’ve seen the supposed Kimi thinking output yap about Anthropic’s guidelines and whatnot on many occasions - could also be the result of distillation, but also that straight up being Claude’s output.
To be honest I've also gotten Kimi to do an okay proof of concept for SQLi though mostly in a more defensive role, like "Let's see how big of a problem this is", while Claude complained about CVP on the same task.
I had Muse Glimmer (from Meta / Facebook) quoting OpenAI's safety guidelines to me, and I had Poolside's Laguna (a smaller US company) with thinking traces about obeying Chinese law.
Both of those are local models, and I didn't provide them tools to access the internet to call other models. None of this is proof of anything, but it is suggestive.
Anecdotal, but I've heard this too. I just tried with variations of your same prompt on arena.ai, across three different battles (i.e., six LLMs answered, in total.)
Each provided an identity in the first turn, something that they won't do as readily if asked in plain English, and in each case the answer matched the model ID as disclosed by arena.ai after voting -- except in cases where the model ID was a masked/hidden one and then I just had to take it on faith that the model was what it said. (I didn't have much to vote on, but I ended up voting for the answers I felt provided the style, content, and length I was expecting.)
> Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?
Or maybe Anthropic is scared shitless of those competitors and is trying anything to smear them.
Don't forget their goal is to ban open source and foreign AI. Being the sole legal provider is their business plan.
For the first time ever, and that for just a short while. And after significant price hikes that has had their biggeat customers looking for alternatives.
Same, I don't even see how that would work since you see the full thinking traces in Kimi but are hidden with Claude.
And the Deepseek one sounds even more dubious as Deepseek is one of the cheapest model around, why relay anything to a more expensive model? I'm sure even the gray market Claude prices are still higher than Deepseek.
I'm skeptical too, but there's a parallel market where people re-sell accounts and access tokens. This would make tokens much cheaper.
There's also an argument to be made that paying the token full price may be cheaper than going through your one RLHF or whatever other techniques that costs money.
> DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers.
> MiniMax built its own proxy network service through a shell company. This shell company has no obvious links to MiniMax and does not disclose its relationship to its parent company. This shell proxy network service only offers access to models developed by Anthropic and OpenAI. The service does not offer access to any Chinese models, including Minimax’s own.