This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
There's a perception in the US that Americans are in a new Cold War, now with China. I'm curious about how they perceive their standing in it. Pulling ahead, like in Raegan's time? Sputnik Moment?
But mainly, I want to hear from Shakes. At the end of June 2026, in another discussion about America Winning the Iran War, I told @Shakes that I'll be returning to his post in which he claimed:
Some updates since then.
First, a recap: in December 2025, the Chinese private launch provider LandSpace attempted a launch and recovery of their methalox powered medium class launch vehicle Zhuque-3 ("Vermillion Bird"), which is very much like Falcon 9 if you don't look at the details (I think it's what Falcon would have been if it were designed in 2020s). The launch and orbital insertion went well, the recovery… not quite, which prompted these kinds of headlines and understandable complacency in some circles.
On July 10th 2026, China Academy of Launch Vehicle Technology (for reference, known as the "First Academy", founded by the exiled communist Qian Xuesen, the co-founder of Jet Propulsion Laboratory which then became the core of NASA) has successfully landed the first stage of their medium class launch vehicle Long March-10B on a sea platform, using a novel net capture mechanism, thus making China the second nation with the capability to reuse first stages of orbital class vehicles, and the only one whose national space agency can do that. There has been commentary to the effect that it's cope for the fact that China can't into precise landing, that you won't have the net capture infrastructure on Mars/Moon/whatever, and that this is a nothingburger and further testament to their backwardness. On August 19th ie today, LandSpace has done this in a more traditional SpaceX manner, with ZQ-3 Y-2 landing on the ground pad. LandSpace was founded in 2015, and ZQ-3 had only been announced in late 2023 – the same year they had put the first methalox rocket into orbit. This, hilariously, puts them technically ahead of SpaceX for the second time. LandSpace is working on a full flow staged combustion methalox engine too, and it's reasonably mature; technologically it's about on par with Raptor 2. Given the track record so far, it's reasonable to say they will have a Starship class vehicle within 5 years. There are multiple similar projects being executed in parallel, both private and state-owned, eg Long March 9. It seems implausible in the extreme that China will find it hard to scale up the production of rocket engines (after all, they do make more WS-15s than F135s now, judging by J-20 vs F35 commission rate), steel tubes (come on) or concrete launch pads (…come on, really), or land permits (lol). So I find it likely that in a fairly short order they can match SpaceX (and thus the US and the world, because SpaceX is a near-monopolist now) in annual mass to orbit if they so wish. With Blue Origin's recent disaster and anklebiters like Rocketlab not doing anything interesting, it appears inevitable that the space race is just SpaceX vs China.
On July 16th, the private startup Moonshot AI has unveiled Kimi K3, the 2.8T, 104B active multimodal MoE, with very innovative architecture and ability to execute on long-horizon self-improvement-related tasks such as chip design or ML research/engineering, delivering performance close to the best American public models, solidly exceeding the previous domestic champion GLM 5.2. Right now it scores 60 on Artificial Analysis, 1 point behind Grok 4.6, and 2-3 behind Opus 5, Fable 5 and GPT 5.6 Sol. (GLM 5.2 reached 53).
Also on Jul 16th, Chinese memory company CXMT completed its IPO subscription. Currently it's worth around $500B, underpinned by its central role in supplying DRAM chips to Chinese (and soon global) industry, including AI. They target 30% global market share by 2030. This is doable, given what I know about the velocity of upstream tool supply chain in China.
On July 18th, during the World Artificial Intelligence Conference in Shanghai, Huawei has demonstrated their Atlas 950 SuperPoD, boasting of the largest scale-up domain in systems for training advanced AI models (yes larger than anything Nvidia ships right now, and this is more important than raw FLOPS as we continue to increase the parameter count). You might be interested in this writeup on Huawei's design philosophy, it's pretty special and promises to compensate for their lack of advanced lithography. In short, their thesis is that the performance comes not so much from Moore's law as from minimization of latency across the entire architecture from transistor to cluster level, and their idea of a solution is 3D-native chip design for multi-layer logic, with very precise (1.5 µm currently, <1µm scheduled, well ahead of the competition) wafer-on-wafer stacking and the first chips demonstrating its viability (mobile Kirin SoCs) coming out in September. Multiple other companies, such as Alibaba, have also shown supernode-based designs, including one absolutely bonkers system from Oriental Computing that uses 14nm chips; as Jensen says, lithographic process is overrated compared to design, so I'm bullish on this line. At the opening ceremony, Xi Jinping delivered a pretty impressive speech on Chinese strategy with regard to AI, committing to support open source and international collaboration.
Since then we've learned of multiple 100K GPU class cluster projects in China (Sugon, completed, Alibaba token factory apparently as well, unclear what's up with Zhipu's 1GW cluster; DeepSeek will also have gigawatt-class systems in Ulanqab, Inner Mongolia, as well as their own chips). These should be sufficient to design and train 10T models, ie comparable to the alleged size of Mythos Preview (and larger than the deployed Mythos/Fable; I am not privy to these details, though). Ryan Fedasuyk of Georgetown estimates that «No matter how we slice the data, we find China is well on its way to producing large numbers of AI accelerators».
On Jul 31st, DeepSeek has deployed and open sourced V4-Flash-0731, getting performance around GLM 5.2 (and much higher on some hard evals like ARC-AGI-2) at a ludicrously low price and parameter count, doing even better than GPT 5.6 Luna after the much-hyped 80% price cut. The situation with DeepSeek is a bit ambiguous, it's not clear if they're flailing (the subsequent Pro was barely any better, their harness project is insanely ambitious but clearly not even half-done); but it speaks to the fact that the Chinese tech ecosystem is now very large and dense and nobody can be champion for long. On Aug 14th, Zhipu has responded to Kimi with GLM 5.3 which is on par with K3 at a fraction of the cost and scale; one of their priorities has been cyberdefense (and thus cyberoffense) capability, plainly driven by concerns around Mythos/Fable. They'll release the weights in <2 weeks, as did Kimi, as did Alibaba Qwen with their 2.4T 95B MoE that's roughly in the same ballpark. The CEO of Zhipu, Jie Tang, is a professor at Tsinghua, and thus essentially a state official, a CCP member who regularly contributes to People's Daily, so we can consider him speaking for the Chinese policy (Xi's speech has much the same tenor). His philosophy is roughly as follows:
As an aside, on Jun 29 it became known that Meituan (a food delivery company) had trained a 1.6T 50B MoE on previous generation Chinese chips, almost certainly Ascend 910Bs. It went under the radar because the model isn't that good, but it sets a lower bound on what can be done going forward, by more competent actors, with more advanced hardware. Today, the globally dominant Hangzhou humanoid/quadruped robot maker Unitree also went public, and is now worth about $50B, ahead of the fraudulent American company FigureAI with $39B and no publicly sold robots to show for it. They clearly have the best hardware at the moment and unmatched development velocity.
I could go on. It's been a rather eventful period. But these are, I think, the most interesting and strategically significant domains: AI, hardware for AI, hardware for manufacturing AI hardware, robotics, and space.
Shakes, do you think China can't catch up?
Didn't you give me a lot of shit about a year back for believing that developing novel and new ML architectures was still practical and useful? I think your paraphrased words were "Data is all that matters, anyone believing the architecture advancements matter is an imbecile" You going to walk that back now?
I'm not sure what you refer to. It's innovative relative to the field, within the dominant paradigm of decoder-only Transformers, and particularly at this scale (from what I know Meta is still cowardly doing basic GQA+SWA on a DeepSeekMoE carcass, for their last generation that they're so proud of, and xAI is similar). It's more efficient and cheaper to serve. It's possibly slightly cooler than internal designs at OpenAI/Anthropic/Google. But it's not like Kimi has reinvented ML. It's just a sign of certain research/engineering maturity and boldness. To the extent that K3 is a strong, good model and not just "cheaper than expected for 104B active 2800B total" model, those qualities are overwhelmingly downstream of their work on data.
I misremembered the exact quote even if I captured the core of overall discussion on Feb 26th 2025. It still being a transformer-decoder arch is not what is really under contention. Your overall stance was that data + compute is all that really mattered and that worrying about architecture, or seeking to make architectural improvements demonstrated a fundamental poor understanding of ML. The problem with taking such a provocatively maximalist position is that now you must defend it in the future when architectural improvements to the transformer arch are made and which you call innovative. Does MoonshotAI have a mediocre understanding of ML? Or instead were you wrong?
You act like the arc of history is settled when you make comments like that. And like every pronouncer of "History has ended, all discoveries that will be made have been made", the march of progress leaves you blacked in the soot of arrogance. It's frankly an anti-science position. Algorithm/Architecture design is as much a core part of ML as is data and compute. The transformer, as it was invented, is unlikely to be the endpoint of ML arch research, just has the CNN, or LSTM were not the endpoint of ML research a decade prior. The transformer of today is different from the transformer of 2018, and I would not bet against the transformer of 2036 being different than that of today. I would not bet against the core elements of the transformer are metastasized into another architecture in 2046. I don't have a crystal ball, I don't know, but I do know planting a flag and saying "The Transformer has solved all architecture problems, no improvement of consequence will ever be made", and then calling anyone who disagrees with you an idiot, is liable, as it has now, to required you to defend increasingly convoluted arguments.
First, you did say that real big boy algorithmic research means leaving this entire basin, and dismissed the kind of innovation I praise in K3 as tinkering with the assembly level (that's not all DeepSeek did, of course, but that was your understanding; and spiritually you were right, it's all about pumping compute more efficiently through a Transformer). I didn't remember that, but it does reinforce my point about Kimi. You said:
Then I clarified my claim:
Now let's look at what the Chinese are actually doing. Their strongest model right now is arguably GLM 5.3. What is GLM 5.3? A basic DSMoE reusing DeepSeek's discarded architecture experiment ["DeepSeek Sparse Attention-prototype"] from October 2025, with one small twist. How did it become so strong? They're very blunt about this:
They explain some aspects of how they did it, might illuminate why scoffing at "mere data engineering" is misguided.
What's the second/equally strongest Chinese model? Kimi K3. It's much more innovative in architecture, but their attention still depends on MLA (invented by DeepSeek in early 2024). No MLA model is this strong. How did it come so far? See image, which illustrates nicely a part of what I was going on about with my breakdown. (Edit: seems like we don't have images. pages 14-16 in the tech report).
Meanwhile, DeepSeek itself went on to redesign attention the third time (fourth if we count the apparently unsuccessful NSA project), to wring even more capacity out of their limited compute, and now for all their sophisticated V4 architecture they are, as @dailydogma tells me with a sneer, "a second rate company in China". What's their most impressive recent result? Flash-0731: «We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview… DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version.» (Recently updated Vision-Exp is the same model with a vision encoder). Just more post-training, just better training signal. If they reach the domestic frontier again, it'll be because they do this more and better. Architecturally, they are ahead, and it's just not doing enough for them.
Grok itself is very competitive now, largely because Elon has bought Cursor, which had a lot of valuable data and expertise on post-training. We don't know its architecture, but from rumors and what I can infer (high cache hit costs, for starters), it's very banal, probably behind all these Chinese models. Inkling from Thinking Machines is clearly banal («The MoE design largely follows DeepSeek-V3»), though it makes some small departures which are basically judgement calls. And these are researchers from frontier American labs.
I could go on (eg this small model from a third tier lab does surprisingly well on ARC-AGI 2, and it's just DSA + SWA again, and uses a bit different RL algo and data). The bottom line is, architecture really does not decide peak model intelligence, and innovations here are overwhelmingly about economics of inference, and the high-leverage research is all happening on the training signal side.
But China is China. I believe my point was much more true for large American companies we were discussing, who are not so compute-constrained. They'll build very strong models with conservative algorithms, and then use those to disassemble all published tricks, make new ones and overtake the crafty Chinese on efficiency too. That's the plan, at least (I don't know how close they are to doing this; GPT 5.6-Luna suggests they are not very far). For them, investing more effort into algo research over data is plainly an opportunity cost.
Maybe you still think that True Geniuses like Hinton or Goodfellow would be disappointed by this. If that is so, I say they were geniuses in a small and uncompetitive pond, and the current crop of talent knows better. It certainly knows better than LeCunn, who by the way got ousted out of Meta by the Pinoy slavedriver Wang we've discussed back then. (You said: «Maybe you can compete with ScaleAI, they do data engineering. Definitely the top AI research company.») Anyway, I'm not walking back shit, my point stands.
If we're still doing Transformer of any kind in 2036, that'll be pretty wild. It'll suggest that even superhuman AI with like a yottaflops for parallel experiments can't find a better primitive than a bunch of Googlers found in 2017 by going through literature and thinking at it. I wouldn't be so optimistic as to predict that. I am very secure in claiming that the priority on data remains rational and empirically backed from the perspective of reaching AGI faster, and it's more rational the more resources a company has; and that people who try to wriggle out of this reality with clever architectures will flounder (case in point: SakanaAI).
You suggested sarcastically that I found an AGI company. Honestly, I don't think that my vision, as outlined here and before, has enough alpha to get anywhere with a yet another company; I believe that AGI is now a resource-intensive heavy industry field, kind of like fracking. But in fact, many people who thought they know better did just that! Have you heard anything from Keen lately? What's your favorite non-transformer lab? What do you think of LeCun's Advanced Machine Intelligence, would you bet they ship anything competitive by 2028? How about you start one?
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link