Up to 10 trillion parameters. That’s the size of the model Bytedance has in pretraining right now, according to the Financial Times, and it would make the TikTok parent the owner of the largest AI model built in China by a factor of three.
The current domestic record holder is Moonshot’s Kimi K3. Bytedance’s run, if the number holds, triples it.
How it stacks up against the American labs
Ten trillion puts Bytedance in the same range as Anthropic's top system, Mythos 5, which industry estimates place at around eight trillion parameters. Worth keeping in mind: Anthropic hasn’t disclosed its own numbers. Those estimates are guesses from outside the building.
Bytedance isn’t alone at that scale either. Elon Musk said xAI is training Grok variants at six and 10 trillion parameters on its Colossus 2 cluster.
Parameter counts are the least interesting number here
Parameters set how much a model can store. They don’t decide how well it performs. Data quality and training method do a lot of the work, which is why a bigger model can and regularly does lose to a smaller one.
The more telling detail from the FT’s reporting is about method. One of the sources said Bytedance has gone more than a year without distillation, meaning it hasn’t trained on the outputs of other companies’ models. For a Chinese lab, that’s a harder road than the alternative, and it’s the kind of claim that only shows up in the benchmarks months later.
What Zhang Yiming told the team
Founder Zhang Yiming has told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term. That’s an internal directive, not a shipping date.
Three insiders told the FT the model is in pretraining. That phase typically takes three to six months, which means nobody outside Bytedance gets to judge whether 10 trillion parameters bought anything until well into that window.