Rendered at 20:47:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
HarHarVeryFunny 2 days ago [-]
Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs).
I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.
In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).
yorwba 2 days ago [-]
From the article:
> A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel — which is legal, in most cases — or direct purchases, which are a breach of US regulations. The Information reported this week that Moonshot is seeking additional Blackwell processors to train its next model.
> The 20,000 chips Moonshot accesses via Alibaba, meanwhile, are from Nvidia’s earlier generation of Hopper products, the people familiar with the agreement said.
So the 20k GPUs from Alibaba is only a lower bound on how many you need to train a model like Kimi K3.
The advantage of having more GPUs in any case is not so much that you can train bigger models, but that the turnaround time is faster, so you can run more experiments to dial in training choices. It's entirely possible that Musk has more than enough compute, but can't hire the talent to run all those experiments. (That would also explain why he has excess capacity he can rent to Google.)
gpugreg 2 days ago [-]
> they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).
A few quotes from the transcript:
> Our current computing capacity is approximately 20,000 H-equivalent units, most of which have just arrived within the past month or two
> Regarding the Huawei 950, Huawei currently provides us with 16,000 SIM cards
> A Huawei 950 [cluster] with 16,000 cards is equivalent to only a B-series card [cluster] with 4,000 cards.
vrganj 2 days ago [-]
> Struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs).
This feels like a very American way of designing things - just throw more horse power at it, bigger is better! The rest of the world is usually a bit more resource constrained and efficient at using those resources.
See also Mustangs vs German sports cars, giant American fridges, giant American suburban McMansions vs livable cities etc etc.
> This feels like a very American way of designing things - just throw more horse power at it, bigger is better!
I don't think so. Bruteforcing problems is a well established strategy in any field that involves computation of any form. Once you get something working, you can get results right now if you throw resources at it. In the meantime, any improvement in efficiency can easily be back ported to the same computational resources you're using.
dang 2 days ago [-]
From a superficial look I think the submitted title ("Moonshot built on 20k Nvidia chip cluster from Alibaba") might have been misleading. Article says "That hardware forms a substantial chunk of the overall computing capacity Moonshot uses for its Kimi models", implying that other hardware was also used.
(We've changed the title to what the article says now.)
ux266478 2 days ago [-]
It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity.
qeternity 2 days ago [-]
I don't follow this. Clearly all of the frontier labs are doing these things.
When OAI released gpt-oss it was released as an mxfp4 checkpoint.
OAI, Ant, et al are also obviously employing QAT.
nextaccountic 2 days ago [-]
But perf improvements not only mean you can run the same thing on cheaper hardware, you can also run more capable models in the same hardware
In this sense, any advance in intelligence is a performance improvement and vice versa
infecto 2 days ago [-]
Genuine question. Is there a time factor part of that equation?
HarHarVeryFunny 2 days ago [-]
It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.
upbeat_general 2 days ago [-]
This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute.
Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count.
No part of this pipeline is fixed in stone.
HarHarVeryFunny 2 days ago [-]
OK, but that doesn't address the question of what can be inferred from Moonshot apparently having access to far less compute then the western labs. To what extent does this reflect the need for less compute due to the (published) architectural efficiencies of their model, and to what extent that they just trained for longer (esp. pre-training - 80% of total compute, perhaps)?
I think the word "distillation" needs to be used a bit more selectively here. If their pre-training run was complete before Fable was released that implies that ZERO Fable data went into the base model. Perhaps the timeline allows for a few weeks at best of incremental post-training on some limited amount of Fable data, but calling this "distillation" seems a bit dramatic especially given the redacted outputs that would have been available. A more factual speculation would just be that they may have had time to post-train using a limited amount of Fable output in some fashion (LLM as judge? SFT? Who knows ...).
infecto 2 days ago [-]
Relatedly I also wonder how much the orchestrated distillation bot nets that Anthropic uses against China are real. Not saying it does not happen but I have never seen dating showing how pervasive it is and I always imagine other western labs would absolutely be doing the same.
theblazehen 1 days ago [-]
There's significant amounts of traffic going through transfer stations [1], which is then re-sold to the model labs. For a while the DeepSeek API was serving Fable as the DS 4 Pro model even. [2]
> Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs)
Are you sure they are using all of their compute on training? Didn't they rent out a ton to other AI companies?
foolswisdom 2 days ago [-]
The renting out is a pretty recent development.
andy_ppp 2 days ago [-]
Do we believe them? It seems there’s no reason with the resources OpenAI and Anthropic have they wouldn’t be incentivised to build equally efficient pipelines? I think it’s mostly the fine tuning with distillation that gives Chinese labs their big advantage. However I’m pretty sure all the main labs are poisoning the output now when they detect Chinese activity so while K3 looks frontier on tests when you use it it’s miles off Sol and Fable.
UqWBcuFx6NV4r 2 days ago [-]
“The only reason that China is good is because it steals from us” spoken from atop the unambiguously crumbling has-been empire.
I’m not American nor Chinese, and I’m from a much more US-aligned country. But, Christ, you lot really are asking for it.
andy_ppp 2 days ago [-]
I don’t think you can steal someone’s work if it’s based entirely off theft tbh, I see distillation as smart - are you denying it’s a huge part of why Chinese models are competitive or do you believe the narrative yourself that they are doing this with several orders of magnitude less compute and the US labs are profligate and fond of burning money rather than optimising?
The truth is probably Anthropic/OpenAI/Google are pretty efficient but less efficient than the Chinese labs, the Chinese labs probably have more compute than they say to undermine US spending and distillation is quite efficient at bridging the gap in compute.
HarHarVeryFunny 1 days ago [-]
The way you make an LLM smarter is by training it on more data, which means you need to make it bigger, which means it costs more to train.
High quality data is expensive. Synthetic data will get you so far, but after that you need to start paying experts to create data for you, which has been going on for a long time. The latest thing is paying for human written LLM-as-judge AI-output evaluation "rubrics", trying to extend RLVR into areas where "looks like it checks the boxes" is the best you can do.
When anyone, Chinese or not (Elon Musk cheerfully admits to distilling OpenAI models) uses the output of someone else's model to train their own, then what they are primarily getting is cheap training data, but you still need to train your model on this data! You may have reduced the cost/speed of training data acquisition, but if you are training a 3T param model (Kimi 3) then you still need the compute to do that - that did not change.
There was an interesting mention of the cost of training data in the recently leaked DeepSeek investor meeting, where their CEO referred to the cost of human-generated training data in China (i.e. using Chinese labor) as being the same as that in the US, which seems surprising. He also mentioned the time such data takes to be created. No doubt the Chinese will catch up in this area - this is just time and money, not Dutch technology (ASML) that the US is blocking them from buying.
hereme888 2 days ago [-]
Phew, if that's not clearly anti-American hate, insult towards Christianity, hatred in general.... Then idk what is.
And btw, yes it's an objective fact from every technical angle that the Chinese only have competitive models because of US tech. Why do you think they try so hard to smuggle NVIDIA GPUs, and now exposed infrastructure to create cheap imitations of American tech?
HarHarVeryFunny 1 days ago [-]
Just in case you are interested in any facts ...
The Trump administration, having first blocked China from buying NVIDIA H100's, has since done a U-turn and is now allowing them to buy the more powerful NVIDIA H200, on a case-by-case basis.
Now, the CHINESE government is blocking Chinese companies from buying these H200s, at least in part because it turns out that being denied US tech has been a great accelerator for Chinese tech, with Huawei now producing the entirely domestic Ascend 950 chip, which according to NVIDIA's Jensen Huang performs about the same as NVIDIA's own H100 (which while not NVIDIA's most powerful is still plenty capable, and is what Elon Musk's Colossus-1 data center mostly uses, currently being rented out to Anthropic).
hereme888 23 hours ago [-]
H100 export licensing restrictions began on August 26, 2022, under the Biden administration.
The CCP tried to promote their own chips, but later regressed and started allowing ByteDance, Alibaba, and Tencent to purchase more than 400,000 H200s. And they are trying to acquire millions.
Thus, export restrictions accelerated Chinese substitution, but also made that substitution slower, costlier, less scalable, and technically inferior to unrestricted Nvidia access.
> According to Jensen Huang, Ascend 950 performs about the same as H100
Source?
bigyabai 18 hours ago [-]
It's been great. I hope they keep importing them, Qwen and Z.ai have done far more for humanity with that compute than Anthropic or OpenAI ever did.
hereme888 1 days ago [-]
[dead]
SubiculumCode 2 days ago [-]
One does wonder whether there has been an expiration of the actual weights of Opus at one point.
dzonga 2 days ago [-]
everything about the current 'a.i' models - east vs west points the wastefulness of western labs - whether in terms of gpu clusters being ran, cost of inference.
the only way is down for the massive valuations and 'a.i' revenue projections.
vineyardmike 2 days ago [-]
> west points the wastefulness of western labs whether in terms of gpu clusters being ran, cost of inference.
Can’t possibly be intentional media strategy by a geopolitical target, right? If its going to negatively impact valuations and revenue of major rivals, seems like a desirable strategy?
We just saw OpenAI take the steps to significantly lower the cost of one of their models, which confirms that at least one western lab has a large margin on inference, not a large inference cost.
Meanwhile, we don’t know how many experiments these east/west labs are performing relative to each other. We also know that many western labs have a whole portfolio of models too, which is product breadth not necessarily waste.
nekusar 2 days ago [-]
The wastefulness is akin to having multiple private jets on standby, or workers coming in to a fishbowl office to be gawked at, or massive datacenters to be used as social "proof".
Its all a fucking capitalistic farce to display to other rich elite that "Look at how much clout I have! I can make these peons dance around and do my bidding! Im a slave-owner!"
He noted that that Silicon Valley doesnt really want to SOLVE problems. They want to find already-solved problems with problem matching. And of course, we just throw more people and more compute instead of optimization and understanding.
The Chinese are being actively constrained with bullshit politics around a second Red Scare moment. And, well, they're winning. A lot.
otabdeveloper4 2 days ago [-]
No, it's an attempt to form an artificial moat by buying all the world's GPUs and memory.
nekusar 2 days ago [-]
True. But check out Ebay right now. Networking gear is CHEAP.
I just bought 2 switches, 48 port 10GbE with 4x QSFP+ at 40Gb fiber. $110 each.
You can even get 24 port QSFP+ @100Gb networking devices for $350.
Yeah while ram and gfx is $$$$$, networking is rock bottom prices..
subscribed 1 days ago [-]
Try getting something new from Arista, Juniper or Fortigate. Lead times are sad.
(new in terms of the most recent models, not just "something with X number of ports")
cubefox 2 days ago [-]
> Musk (who freely admits to distilling OpenAI's models)
I’m still baffled that OpenAI can complain about this with a straight face while these models are literally trained on everything regardless of copyright or permission.
nomel 2 days ago [-]
Raw stolen data that is in no way related to AI vs a very very expensive transformative compilation of that raw stolen data that results in usable AI. The frustration is completely understandable to me, since distillation skips much of that "very very expensive" part.
And yes, I understand the stolen data was expensive to make, so I understand the owners of it are also frustrated, but that's partly a problem with current law. Would the authors of the world be rich if OpenAI bought a single copy of their book to legally scan? For best sellers, that's somewhere around pennies, so no. Should the authors get a share in OpenAI? Current laws says, unambiguously, "no".
Frustration all around is reasonable.
HarHarVeryFunny 2 days ago [-]
Not really, in this context the word "distillation" is really being abused, or at least used in a different sense than when it was originally introduced in the Hinton et al "Distilling the Knowledge in a Neural Network" paper, where it was essentially referring to knowledge compression.
The way Anthropic are using "distillation" is just in a very broad vague sense to claim that some data generated by their model was used to help train another one. They are not talking about something like internal logits, expensive to derive, that would be useful to train a smaller model, but rather about any output from their model, even outputs with redacted reasoning (i.e. incomplete outputs that do NOT reflect the underlying knowledge of the source model).
Given the way Anthropic are using the word, IMO it's better just to think of this as cheap training data, and indeed very similar to the way Anthropic themselves got cheap training data just by taking it (even in cases where that was illegal - copyright). The alternative for Moonshot would be to pay for human generated reasoning data, just as the alternative for Anthropic would have been to pay human developers for coding data etc, not just take it from wherever they could lay their hands on it.
So, I guess Moonshot may have violated Anthropic's TOS, in using Anthropic output to compete against Anthropic, but unlike Anthropic they at least didn't break copyright law since model output is not copyright, and technically they may not have even violated Anthropic's Terms of Service unless they owned the accounts used to access Anthropic's models (perhaps not - they may have used one of the anonymizing Chinese token resellers).
So yeah - pot calls kettle black.
sirsinsalot 20 hours ago [-]
You're being reductive. It isn't just one copy of the book. It is using the book to enable, in some part, value to be generated from it that goes in the pocket of a rent seeker (AI business) and not the author.
2 days ago [-]
a-priori 2 days ago [-]
It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.
The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.
HarHarVeryFunny 2 days ago [-]
Yes, continual learning is currently a hot topic since this is what will allow an "AI intern or new employee" to learn on the job rather than be stuck on groundhog day, or pre-trained for every eventuality - needing to anticipate all the quirks and proprietary knowledge of every customer!
Continual learning tends to imply individualized models, else there is no data privacy (the secrets learned on the job at your company now being available to your competition), which really turns the current AI business model of a single centralized model served to everyone on it's head. If every customer has a different model that essentially means the end of batch processing with the same weights loaded into the GPU.
The direction this suggests is a move away from centrally served common models to locally served individual ones, which generally requires them to be smaller, even if some larger companies may be willing to invest in beefier hardware.
I think this is at least in part why the AI companies are trying NOT to implement true continual learning and see if they can instead finesse it by implementing continual compacted(?) memorization instead, since then it's "just" additional context that needs to be recalled and fed into every request, not weights that need feeding into the GPU. I don't think memorization is any substitute for learning, especially learning of practiced skills, but since it's far easier to implement, and non-disruptive to the cloud-based API business model, this is what we will see first.
The recent news of NVIDIA' investment in Sutskever's SSI has a tiny hint of this also, talking about SSI advising on NVIDIA's future architectural direction - apparently pushing it in a different direction than current (cloud-based, pre-trained) models. NVIDIA may be quite happy to see a move towards local models.
wongarsu 2 days ago [-]
The middle ground would be continual learning for a group of users. For a company with a couple thousand employees, one model learning on all your employees might be viable. With some filtering for information that is supposed to stay compartmentalized (HR data, stuff under NDA, core technolgy). But this only really works if you bring enough volume to make it economical to effectively server you a whole other model
btown 2 days ago [-]
It begs an interesting question as well: are frontier labs stuck in the incumbent phase of the "innovator's dilemma," where there's such intense pressure to maintain the flavor/character of their widely-loved existing models, that they dare not invest their resources into radical reimagination of their approach towards architectures for smaller models? A company like Moonshot does not have this cultural limitation.
The frontier labs would be well served in carving out 20k sub-clusters and giving research teams carte blanche in building things with radically different architectures - with full permission to distill whatever they want from the flagship models. We'd expect to see more product lines that feel "different" from the flagship models if this were already being done.
vkaku 2 days ago [-]
Eventually, multi-modality might even get offloaded as workflows, which might allow models to do this at a fraction of the data and compute required.
I am adding multi-modality to https://github.com/guilt/TinyToT, and I see that dis-aggregating capabilities, very similar to how our own sensory organs work, seems to be paying off quite well.
__MatrixMan__ 2 days ago [-]
It feels like we're not really bottlenecked by model performance now anyway. Most cases I find where a model is being an idiot can be remedied by rethinking the context and tooling available to that model.
If they got a lot cheaper and only a little dumber, and we got a little smarter about how we use them, they could appear to the bystander as much more useful than they are.
znpy 2 days ago [-]
> It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller
isn't that, in the end, the case with all/most technologies?
i can get a petaflop of compute capacity in a dgx spark for relatively cheap nowadays. that used to be a whole supercomputer like 20 years ago.
WarmWash 2 days ago [-]
This is likely what Google has been doing, as it suites them best to have small fast models rather than hulking slow giants. Lower intelligence but way more ability to serve. Anthropic would probably need a datacenter the size of a small country to serve Fable on a Google search/services scale.
trollbridge 2 days ago [-]
Getting smaller is how the last revolution in computing happened.
A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.
radialstub 2 days ago [-]
Yes but that's because of the scaling laws for transistors. ML models seem to get better the bigger they are. If you want to compress the world's information, you need to look at all the information in the world.
trollbridge 2 days ago [-]
Qwen 3.6 27b is way ahead of Llama 70b.
infecto 2 days ago [-]
My thesis is still that model building has no moat. Folks continue to migrate around between the big labs. There is a lot of value in having good taste around the harness and how the models are used. The medium to long term winners will be the folks that control the compute.
Aperocky 2 days ago [-]
You can only control the compute if it is not commoditized, which at the current moment, it is not, leading to industry players having 75-80% margins.
Someone will find a way to make cheaper compute, and since nothing fundamental changed there (LLM didn't change how silicon were made), that's bound to happen.
It's also built on the back of downloading all the content of the top ten sites on the list, the fact that all the independent content sites with their "pre-war" steel haven't been acquired at top dollar by U.S. companies shows the existential threat / stop China posturing is a total joke.
No one is using that dataset for any serious models. It’s way too small and old.
fellowniusmonk 1 days ago [-]
Tell that to the nation states that are actively trying to find every exploit they can to pull it down.
ycui7 2 days ago [-]
K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math.
Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.
The article’s statement does not make sense.
throwa356262 2 days ago [-]
Huawei A910 is supposed to have native support for MXFP4. The problem is that huawei doesn't have the capacity to build more GPUs right now. Those yield numbers must be really terrible.
Personally, I think 20K nvidias is a stop gap solution because they really don't have the capacity to serve their models to earn any money right now.
basiccalendar74 2 days ago [-]
Weights will be fetched 4-bit from memory, but compute happens in FP8 (W4A8) or BF16 (W4A16). This is still better than fetching 8-bit weights from memory.
Both W4A8 and W4A16 schemes are supported by Hopper GPUs and commonly used to serve mxfp4 Kimi models on Hopper.
Tempest1981 2 days ago [-]
> Either they have Blackwell
From the article:
A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel...
orliesaurus 2 days ago [-]
What do you mean by "it's meaningless" can you expand?
sho 2 days ago [-]
They mean that K3's native floating point is MXFP4, a newer standard that's not supported in Hopper, an older nvidia chipset, so it's impossible the GPUs in question were Hopper (the article specifically mentions H200s), which is correct.
Probably either the author or the people they interviewed got the specific GPU details mixed up or wrong. Maybe it was B200s.
gpugreg 2 days ago [-]
Notably, MXFP4 was introduced at the (much less costly) supervised fine-tuning stage after pretraining, so the number of B200/B300 GPUs could be relatively small in comparison to the number of H200 GPUs used during pretraining (or maybe not, who knows).
Oh, okay, thank you for the explanation. I was under the impression that we, the United States, were banning exports of such GPUs.
Zigurd 2 days ago [-]
I would not minimize the extent to which AI progress depends on things like the effort that goes into training material acquisition and data labeling. In those areas, improvement is a direct function of how much you invest.
I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.
That leaves a lot of room for efficient aggressive followers.
sho 2 days ago [-]
If Chinese companies are able to just reliably rent the damn GPUs then I really don't see how these restrictions are meant to be effective in the slightest. I guess it makes it less convenient and supply less certain, and retains some optionality for cutting off supply later?
If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV partners where relevant) and then just provide them as a service to its customers in China - what is the point? I'm sure many customers would actually prefer such an arrangement.
downrightmike 2 days ago [-]
"what is the point?" There is no point, just zero-sum thinking which isn't compatible with reality.
I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.
In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).
> A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel — which is legal, in most cases — or direct purchases, which are a breach of US regulations. The Information reported this week that Moonshot is seeking additional Blackwell processors to train its next model.
> The 20,000 chips Moonshot accesses via Alibaba, meanwhile, are from Nvidia’s earlier generation of Hopper products, the people familiar with the agreement said.
So the 20k GPUs from Alibaba is only a lower bound on how many you need to train a model like Kimi K3.
The advantage of having more GPUs in any case is not so much that you can train bigger models, but that the turnaround time is faster, so you can run more experiments to dial in training choices. It's entirely possible that Musk has more than enough compute, but can't hire the talent to run all those experiments. (That would also explain why he has excess capacity he can rent to Google.)
A few quotes from the transcript:
> Our current computing capacity is approximately 20,000 H-equivalent units, most of which have just arrived within the past month or two
> Regarding the Huawei 950, Huawei currently provides us with 16,000 SIM cards
> A Huawei 950 [cluster] with 16,000 cards is equivalent to only a B-series card [cluster] with 4,000 cards.
This feels like a very American way of designing things - just throw more horse power at it, bigger is better! The rest of the world is usually a bit more resource constrained and efficient at using those resources.
See also Mustangs vs German sports cars, giant American fridges, giant American suburban McMansions vs livable cities etc etc.
I don't think so. Bruteforcing problems is a well established strategy in any field that involves computation of any form. Once you get something working, you can get results right now if you throw resources at it. In the meantime, any improvement in efficiency can easily be back ported to the same computational resources you're using.
(We've changed the title to what the article says now.)
When OAI released gpt-oss it was released as an mxfp4 checkpoint.
OAI, Ant, et al are also obviously employing QAT.
In this sense, any advance in intelligence is a performance improvement and vice versa
Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count.
No part of this pipeline is fixed in stone.
I think the word "distillation" needs to be used a bit more selectively here. If their pre-training run was complete before Fable was released that implies that ZERO Fable data went into the base model. Perhaps the timeline allows for a few weeks at best of incremental post-training on some limited amount of Fable data, but calling this "distillation" seems a bit dramatic especially given the redacted outputs that would have been available. A more factual speculation would just be that they may have had time to post-train using a limited amount of Fable output in some fashion (LLM as judge? SFT? Who knows ...).
[1] https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... [2] https://xcancel.com/synthwavedd/status/2078514339552628880
Are you sure they are using all of their compute on training? Didn't they rent out a ton to other AI companies?
I’m not American nor Chinese, and I’m from a much more US-aligned country. But, Christ, you lot really are asking for it.
The truth is probably Anthropic/OpenAI/Google are pretty efficient but less efficient than the Chinese labs, the Chinese labs probably have more compute than they say to undermine US spending and distillation is quite efficient at bridging the gap in compute.
High quality data is expensive. Synthetic data will get you so far, but after that you need to start paying experts to create data for you, which has been going on for a long time. The latest thing is paying for human written LLM-as-judge AI-output evaluation "rubrics", trying to extend RLVR into areas where "looks like it checks the boxes" is the best you can do.
When anyone, Chinese or not (Elon Musk cheerfully admits to distilling OpenAI models) uses the output of someone else's model to train their own, then what they are primarily getting is cheap training data, but you still need to train your model on this data! You may have reduced the cost/speed of training data acquisition, but if you are training a 3T param model (Kimi 3) then you still need the compute to do that - that did not change.
There was an interesting mention of the cost of training data in the recently leaked DeepSeek investor meeting, where their CEO referred to the cost of human-generated training data in China (i.e. using Chinese labor) as being the same as that in the US, which seems surprising. He also mentioned the time such data takes to be created. No doubt the Chinese will catch up in this area - this is just time and money, not Dutch technology (ASML) that the US is blocking them from buying.
And btw, yes it's an objective fact from every technical angle that the Chinese only have competitive models because of US tech. Why do you think they try so hard to smuggle NVIDIA GPUs, and now exposed infrastructure to create cheap imitations of American tech?
The Trump administration, having first blocked China from buying NVIDIA H100's, has since done a U-turn and is now allowing them to buy the more powerful NVIDIA H200, on a case-by-case basis.
Now, the CHINESE government is blocking Chinese companies from buying these H200s, at least in part because it turns out that being denied US tech has been a great accelerator for Chinese tech, with Huawei now producing the entirely domestic Ascend 950 chip, which according to NVIDIA's Jensen Huang performs about the same as NVIDIA's own H100 (which while not NVIDIA's most powerful is still plenty capable, and is what Elon Musk's Colossus-1 data center mostly uses, currently being rented out to Anthropic).
The CCP tried to promote their own chips, but later regressed and started allowing ByteDance, Alibaba, and Tencent to purchase more than 400,000 H200s. And they are trying to acquire millions.
Thus, export restrictions accelerated Chinese substitution, but also made that substitution slower, costlier, less scalable, and technically inferior to unrestricted Nvidia access.
> According to Jensen Huang, Ascend 950 performs about the same as H100
Source?
the only way is down for the massive valuations and 'a.i' revenue projections.
Can’t possibly be intentional media strategy by a geopolitical target, right? If its going to negatively impact valuations and revenue of major rivals, seems like a desirable strategy?
We just saw OpenAI take the steps to significantly lower the cost of one of their models, which confirms that at least one western lab has a large margin on inference, not a large inference cost.
Meanwhile, we don’t know how many experiments these east/west labs are performing relative to each other. We also know that many western labs have a whole portfolio of models too, which is product breadth not necessarily waste.
Its all a fucking capitalistic farce to display to other rich elite that "Look at how much clout I have! I can make these peons dance around and do my bidding! Im a slave-owner!"
https://infosec.exchange/@david_chisnall/116991627711001827
He noted that that Silicon Valley doesnt really want to SOLVE problems. They want to find already-solved problems with problem matching. And of course, we just throw more people and more compute instead of optimization and understanding.
The Chinese are being actively constrained with bullshit politics around a second Red Scare moment. And, well, they're winning. A lot.
I just bought 2 switches, 48 port 10GbE with 4x QSFP+ at 40Gb fiber. $110 each.
You can even get 24 port QSFP+ @100Gb networking devices for $350.
Yeah while ram and gfx is $$$$$, networking is rock bottom prices..
(new in terms of the most recent models, not just "something with X number of ports")
Source?
And yes, I understand the stolen data was expensive to make, so I understand the owners of it are also frustrated, but that's partly a problem with current law. Would the authors of the world be rich if OpenAI bought a single copy of their book to legally scan? For best sellers, that's somewhere around pennies, so no. Should the authors get a share in OpenAI? Current laws says, unambiguously, "no".
Frustration all around is reasonable.
The way Anthropic are using "distillation" is just in a very broad vague sense to claim that some data generated by their model was used to help train another one. They are not talking about something like internal logits, expensive to derive, that would be useful to train a smaller model, but rather about any output from their model, even outputs with redacted reasoning (i.e. incomplete outputs that do NOT reflect the underlying knowledge of the source model).
Given the way Anthropic are using the word, IMO it's better just to think of this as cheap training data, and indeed very similar to the way Anthropic themselves got cheap training data just by taking it (even in cases where that was illegal - copyright). The alternative for Moonshot would be to pay for human generated reasoning data, just as the alternative for Anthropic would have been to pay human developers for coding data etc, not just take it from wherever they could lay their hands on it.
So, I guess Moonshot may have violated Anthropic's TOS, in using Anthropic output to compete against Anthropic, but unlike Anthropic they at least didn't break copyright law since model output is not copyright, and technically they may not have even violated Anthropic's Terms of Service unless they owned the accounts used to access Anthropic's models (perhaps not - they may have used one of the anonymizing Chinese token resellers).
So yeah - pot calls kettle black.
The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.
Continual learning tends to imply individualized models, else there is no data privacy (the secrets learned on the job at your company now being available to your competition), which really turns the current AI business model of a single centralized model served to everyone on it's head. If every customer has a different model that essentially means the end of batch processing with the same weights loaded into the GPU.
The direction this suggests is a move away from centrally served common models to locally served individual ones, which generally requires them to be smaller, even if some larger companies may be willing to invest in beefier hardware.
I think this is at least in part why the AI companies are trying NOT to implement true continual learning and see if they can instead finesse it by implementing continual compacted(?) memorization instead, since then it's "just" additional context that needs to be recalled and fed into every request, not weights that need feeding into the GPU. I don't think memorization is any substitute for learning, especially learning of practiced skills, but since it's far easier to implement, and non-disruptive to the cloud-based API business model, this is what we will see first.
The recent news of NVIDIA' investment in Sutskever's SSI has a tiny hint of this also, talking about SSI advising on NVIDIA's future architectural direction - apparently pushing it in a different direction than current (cloud-based, pre-trained) models. NVIDIA may be quite happy to see a move towards local models.
The frontier labs would be well served in carving out 20k sub-clusters and giving research teams carte blanche in building things with radically different architectures - with full permission to distill whatever they want from the flagship models. We'd expect to see more product lines that feel "different" from the flagship models if this were already being done.
I am adding multi-modality to https://github.com/guilt/TinyToT, and I see that dis-aggregating capabilities, very similar to how our own sensory organs work, seems to be paying off quite well.
If they got a lot cheaper and only a little dumber, and we got a little smarter about how we use them, they could appear to the bystander as much more useful than they are.
isn't that, in the end, the case with all/most technologies?
i can get a petaflop of compute capacity in a dgx spark for relatively cheap nowadays. that used to be a whole supercomputer like 20 years ago.
A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.
Someone will find a way to make cheaper compute, and since nothing fundamental changed there (LLM didn't change how silicon were made), that's bound to happen.
https://archive.is/AvWWX
Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.
The article’s statement does not make sense.
Personally, I think 20K nvidias is a stop gap solution because they really don't have the capacity to serve their models to earn any money right now.
Both W4A8 and W4A16 schemes are supported by Hopper GPUs and commonly used to serve mxfp4 Kimi models on Hopper.
From the article:
A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel...
Probably either the author or the people they interviewed got the specific GPU details mixed up or wrong. Maybe it was B200s.
(Kimi K3 tech report section 4.1.1 https://arxiv.org/pdf/2607.24653)
I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.
That leaves a lot of room for efficient aggressive followers.
If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV partners where relevant) and then just provide them as a service to its customers in China - what is the point? I'm sure many customers would actually prefer such an arrangement.
[1] https://www.tomshardware.com/tech-industry/artificial-intell...