site banner

Culture War Roundup for the week of July 27, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

AI safety & open models

Yesterday, Kimi K3's weights were published on HuggingFace. We now have an open model comparable to GPT 5.5 and Opus 4.8, which were SOTA only a couple months ago.

An example of what it can do on its own: this website that imitates macOS Desktop (background). Click around, every app has lots of features, and again, this was implemented in one shot. Do you not think that's impressive?

You can, at least in theory, run this on your own equipment. Millions of dollars of equipment, sure, but far more attainable than running GPT or Opus. A medium-sized corporation or small government can.

Which presents problems: its impressive capabilities can be used for evil, without guardrails or surveillance unlike GPT or Opus. Like Fable, Kimi has already found several vulnerabilites. For examples of real evil, see how other LLMs are being used for terrorism by Boko Haram (and almost certainly other groups).

Regulation concerns

Allegedly some US officials are considering restricting US companies from using Chinese open models. Despite this claim being repeated across many outlets, I didn't actually find any evidence. However, I did find plenty of tweets criticizing the models' development and themselves, like this tweet by Treasury Secretary Scott Bessent:

We support open-source AI and the innovation it unlocks. But open source is not open season on American IP. When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.

IP theft from Anthropic, who themselves are disregarding IP? Really?

What seems more indicative of a potential future restriction, the US government is already rushing unclear regulations for US models Fable/Mythos and GPT 5.6. If an open-weight model reaches somewhere around their capability, intuitively it would also be restricted, and Kimi is close.

Tech companies...support open weights models?

You have people like Dean Bell arguing for regulation, but many companies including Andressen Horowitz, Dell, IBM, Meta, Microsoft, and front and center Nvidia came out in support of open models, in this letter.

Key paragraphs (emphasis mine):

To be sure, open weights carry real and distinct risks. Once released, the weights are beyond the original developer's control, and modified versions are difficult to trace or reverse. But the right response to this risk is not to prohibit open weights. In a world where cybersecurity attackers use advanced AI, defenders need access to models with comparable capabilities so they can detect, simulate, and respond to emerging threats. Open models broaden defensive capability, increase transparency, and allow vulnerabilities to be discovered and remediated across many teams

In fact, openness may be one of the most important paths to AI safety and security. Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time. Just as open-source software demonstrated that transparency can be more secure than obscurity, AI safety may depend on giving more people the ability to test and strengthen the models on which society relies. It allows for rigorous benchmarking and evaluation, red teaming, and protections tied to real and demonstrated harms rather than assuming that closed systems are safer by default.

Also

In shaping this ecosystem, policymakers should be careful not to conflate legitimate model-development techniques with misappropriation. Distillation, or the practice of using one model's outputs to help train or improve another, is a widely used technique for model improvement, evaluation, and validation. It reflects a long tradition of learning from, building upon, and improving existing technologies, a tradition that has helped drive innovation since the rise of the open-source software movement. By contrast, unlawful efforts to extract value from closed models raise legitimate concerns. Those concerns should be addressed through targeted legal and commercial frameworks rather than sweeping restrictions on techniques that play an important role in AI innovation.

Unlawful extraction from companies that have themselves unlawfully extracted? Again, really?

Regardless, I think overall it's a good sign.

Anthropic's position

Key quotes (emphasis theirs)

Anthropic has never advocated for a ban on open-weights models.

I do support the following three measures, which I and Anthropic have consistently advocated for:

  1. We should not sell powerful chips or chipmaking equipment to China...
  2. We should crack down on industrial-scale distillation operations...
  3. All sufficiently capable models, open and closed, should go through mandatory safety testing...

This brings me to the open letter. I agree with much of it...But I don't agree with the letter's assertions that open-weights models necessarily make it easier to develop safeguards or that broad access to capabilities necessarily helps defenders more than attackers...For example, I worry that biology will have a strong attacker-defender asymmetry...Questions like this should be empirically answered by rigorous pre-release testing, not assumed in advance.

The important part is that Anthropic claims they don't want to ban open models, but want "mandatory safety testing" applied to all models. They use the threat of an AI-assisted superbug, which admittedly could cripple civilization, but so far is merely plausible. But that could effectively ban open models if it's implemented such that only closed models pass, like how "nobody can sleep under a bridge" applies equally to rich and poor but only affects the latter.

But it's also part of Plan A, proposed by the AI rationalists, who are supposed to be experts on this topic (what else have they been doing the past 10+ years?). Plan A actually argues against open models entirely, although it specifies that access to the models should be open, and all development and regulation discussions should be public. But how can we publicly develop models without making them open?

I'm curious what the Plan A authors think about Kimi K3, the open letter, and Anthropic's response; I haven't seen anything on lesswrong.com yet, although admittedly I only skimmed the front-page and recent.

My thoughts

For now, I support open weights models.

If someone comes up with a way to regulate AI development that doesn't eventually consolidate power into corrupt hands, sure. But who can be trusted? Even if future atrocities are caused by open models, they may be lighter than the atrocities committed in an alternative timeline, by a tyrant who gained power with the help of regulation, or lack of open models that prevented them.

As for "unlawful training": I still maintain the position that IP should gradually be completely abolished. AI training has already been ignoring IP, so I believe that should continue.

That includes China training on American models. I doubt China will surpass American companies if they're training on American models, especially since America has more hardware. And incentive? Come on, the American companies have enough incentive even if they had to distribute their models freely, from the dream of ASI.

There seems to be a rhetorical trick being employed by safety proponents where "being in support of open models" means "sharing, in some small part, responsibility for when open models end up causing damage" when in reality this is a total non-sequitur.

As a first matter of practicality, unless you are a high-ranking member of the CCP, there is approximately nothing you can do to prevent proliferation of open models if it's deemed in the interest of the Chinese state that it happens; Pakistan and North Korea got the bomb, and the world got Kimi K3, despite American seething to the contrary. Banning Chinese models from being hosted on American soil or sanctioning the Chinese model developers is simply pretending that this isn't going to happen, and entrenching the market position of the existing closed-weight providers, rather than achieving anything meaningful.

As a second matter of practicality, from what we can tell the majority of actual black-hat activity eventuated via LLM has happened via jail-breaking of closed-weight models rather than the use of the feared abliterated open-weight models e.g this Claude hack of Mexico. If you're a black hat with ~infinite access via black market token proxies and don't fear legal consequences, you can already just put effort into jailbreaking the frontier models and using them for evil, while legitimate white hats are the ones that have to use Chinese models because they fear reputational and legal consequences. This is a fully intractable problem unless you go actually closed-weights i.e no access other than by via white-list.

If we already live in a fragile world, then so be it; attempting to restrict open weights is not going to make it any less fragile. It's better to do what we can to harden the world via proliferation rather than sticking our heads in the sand, hoping that will reverse the changes happening in the world.

One angle I rarely see discussed is that open weight models are inherently decelerationist for the frontier. They greatly reduce the profit motive to push the frontier, especially if distillation is allowed because the make it much harder to recoup the investment thus reducing the capital available to invest. Ultimately though they're definitionally unsafe and can never be made safe, any work put in to align them can be sanded off with post training and even with just mundane levels of uplift the idea that the offense/defense equilibrium will favor the defender in most areas, let alone all areas is just impossible to believe.

They greatly reduce the profit motive to push the frontier, especially if distillation is allowed because the make it much harder to recoup the investment thus reducing the capital available to invest

That's the argument of a few OpenAI goons, namely roon, Will Depue and most articulately Dean Ball:

Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models… One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end.

It's also innumerate bullshit from nervous, greedy bagholders. Since I've blocked you on X (deservedly, seeing the unearned confidence in the latter half of your post), you probably haven't seen my multiple posts on this, but to sum it up: margins on inference at frontier labs are obscene, frankly if people understood this, AI lab CEOs would be at nontrivial risk of vigilante justice from mobs with pitchforks. The talk about bubbles, subsidized subscriptions, capital to reinvest into R&D – this is all one big pile of bullshit obscuring unhinged rent extraction, which is viable because demand still exceeds supply, because they physically cannot mint enough tokens with gigawatts of deployed capacity. Frontier AI is wildly profitable, it has high software margins, and Chinese distillation is powerless to dent that.

Let's do the math. DeepSeek comes to the rescue, as usual, since they are 1) all time champions of low API pricing and 2) the only ones who periodically disclose some of their economics (starting in the V3/R1 era). Soon after V4's deployment, they have cut prices by 4x and started using DSpark, an advanced lossless speculative decoding technology that «accelerates per-user generation speeds by 60%–85% at matched throughput levels». They've open sourced it, so even if the frontier wasn't yet on that level (dubious; they are technically overrated and from what I hear arguably behind China on infrastructure and generally on engineering culture, but not really incompetent), it is now. It's general-purpose (eg see here how it's adapted to accelerate Kimi K3, a wildly different architecture, by up to 4x). Obviously frontier labs can customize it further for their scenarios, and not tell us about it.

DeepSeek uses it to get average ≈2000 tokens/second/GPU decode for V4-Pro at "decent speeds" (≈70 tokens/second per user stream), in live traffic. Presumably their inference GPUs are something like H200 or Ascend 950DT, with whatever better stuff they have smuggled reserved for training and experiments (at the time of publishing, they had 20K H100-equivalents total). Prefill and cache hits are much cheaper but we can assume they're priced roughly proportionally to decode in wall clock cost, this follows from kv cache difference in volume and pricing between V4-Pro and V4-Flash and is also obvious from timing (eg it takes their first party API seconds to process a 500K input). So what does this get us? At this utilization, 500 seconds of GPU-time for 1 million tokens, priced at $0.87.. $6.26 per hour. $54862 of revenue per year. If we use the usual rent cost of $2 per hour of a Hopper generation GPU, we get $37.3K of pure annual profit. If we assume it's owned and the per-GPU consumption is 2 kW (very high), that's just $1.7K of electricity costs at Chinese industrial pricing. It's mostly owned now, but they'll be renting to expand. So, something like $45K/year of average profit per GPU. At sub $1/million output tokens. Is this plausible? Well, in a recently leaked investment call, Liang Wenfeng insists that his target is 10 months to breakeven on hardware: «91. Our API pricing is designed to generate a reasonable profit. Roughly speaking, if we buy a batch of equipment on the market, recovering the investment in about ten months represents a reasonable level of profitability … » 10 months is $37K of profits as per the above. H100 tier cards were going for $35K a piece recently, though there's more scarcity recently.

In other words, DeepSeek for all this Confucian restraint has something like 70-84% profit margin. Well would you look at that, The Information reports: «"The gross profit margin from selling cloud based access to V4 is somewhere between 70 to 80%."» I don't have the subscription, but seems like being able to do 1st grade arithmetic is enough. By normal business logic, «10 months to breakeven» allows you to raise multiples for the next round of capex, which is exactly what DeepSeek successfully did, and will do again soon, and what all American frontier labs have done repeatedly.

How big are frontier models? I think they're not big, judging by knowledge coverage which is the most honest indicator of scale, due to physics of LLMs – eg Opus 4.8 is in the same band as Kimi K3, V4 and GPT 5.6 Terra. How costly are they to produce? For V4, we have an idea of training costs: 49B active, 33T tokens, 6ND = 9.7e24 FLOPs. V3 was 37e9*14.8e12*6 = 3.3e24, and trained over 55 days on 2048 H800s, giving us 2,664M GPU-hours ($5.3M) and a reasonable (for large sparse MoE) MFU of maybe 25% (depending on how you count mixed fp8, 20-40%; this is genuinely annoying). Anyway, we can say that V4 is 3 times larger = $15M pretrain at rental costs. Sparsity is informed by scaling laws and inference optimality, and is likely 3-4ish % across model lines (frontier models most likely use the same DSV3-type designs). So, Kimi K3 has twice the active parameters of V4, at 104B; Opus is likely similar; Claude Fable may be twice that again, something like 200B. GPT 5.6 Sol – 150B? That's roughly what I get from insiders. So these models still cost at most low hundreds of millions to pretrain, likely under <100M. They also are more overtrained (GPT-OSS hinted at 60-120T pretraining tokens vs 30-50 in Chinese models), with more rigorous and intense post-training (10% of pretraining cost vs 50-100% or more in the US; Cursor threw 7x more FLOPs at Kimi K2.5 RL to make Composer 2.5, though this is crude Muskian maximalism in a desperate attempt to catch up), so they don't need to be large to perform better in downstream tasks. They also operate on a vastly larger scale, so can do even more efficient fleet level optimizations. All in all, a frontier model production goes for <$500M today. Anthropic's ARR reached $47 billion in May and continues to explode. So let's talk of reinvestment. Fine! Let's say, for the sake of argument, that Americans are lower IQ, can't do math, are unable to capitalize on their years of first mover advantage, on staff that had founded the whole field, on reams of user data, unwilling to waste time to optimize for costs, whatever – so they're a whopping 2 times less efficient than the frugal (and distilling) China. So what do we get? $50/1M tokens with Fable, $30 with Sol, $15 with Terra. They also have access to radically more efficient Blackwell hardware, soon Rubin (also TPUs, Trainium, Cerebras etc.), which more than makes up for the increase in compute demands per token/second.
Do you see where I'm going with all this? Their API margins are over 90%. They get payback on their compute in under Wenfeng's 10 months. Their "generous" subscription limits do very little to offset this. There's no "subscription war", it's kayfabe and a price fixing cartel. The only reason they're in the red is that they are, indeed, reinvesting tens of billions into scaling compute for R&D and larger-scale training. How much capital do they need to keep accelerating? Hundreds of billions? Trillions? Mind you, algorithms and hardware are constantly improving (stuff like DSpark is pretty mundane, and will get trivial in the age of AI doing AI – which even K3 is already doing). And they're still nowhere close to exhausting the demand, even as they keep jacking up prices. Frontier intelligence that can automate competitive white collar labor in the US is priced against not even the hour of labor but against its revenue, and would make economic sense even at costs an order of magnitude higher, and that's where we're going. How much revenue do you generate in an hour of work for your employee?

Enough. This is all noise. All plausible negative impact of Chinese labs, distillation or whatever, can be negated by roughly 2 months of regular progress. The only serious question is how the multi-trillion valuation pie gets sliced – how much goes to "labs", how much to Jensen, and how much to clouds serving a mix of closed and open weights. The championing of open weights comes from the clouds and Jensen, because he benefits from maximum demand for CUDA systems and doesn't want to help OpenAI et al. to switch to in-house compute and cut him off (which they are trying to do, and have the capital for). The bullshitting about "deceleration" comes from the first camp, because they're SF dorks high on their own supply and want their equity to get high enough to buy entire landmasses with genetically engineered catgirl slaves (if not star systems in the coveted "light cone"). My opinion is that the rest of us can play the world's smallest violin if they miss those IPO targets; American AGI will arrive on schedule regardless. Finally, as Wenfeng correctly observes from Communist China, it's a political impossibility that any self-appointed «AGI lab» would be allowed to capture 10% of the world's or even the nation's GDP; so he pursues «restraint» and «reasonable profit» to preclude the painful enforced correction that the Chinese tech/finance scene is familiar with. The logic in the US, however, will be similar.

How much capital do they need to keep accelerating? Hundreds of billions? Trillions?

Scale is all you need. We are at the beginnings of RSI. Anyone (except, apparently, Google) can build a frontier model, capable of research, for the chump change of $1B, or likely less. An extra $10B, though, and you can build it and its successor faster. Promises of the light cone let you raise that money easily, but threats to have to share it make investors start asking (somewhat irrationally) how much of the light cone they get. That slows you down. Not that OAI and Anthropic are exactly hurting for capital, but at the margins it makes a decelerationist difference.

All plausible negative impact of Chinese labs, distillation or whatever, can be negated by roughly 2 months of regular progress.

Curious how you got that number, though I agree any delay they introduce is on the order of months. Even two months, though, is a long time nowadays.

(The strongest argument for a Chinese accelerationist effect is the closed labs can easily incorporate their research and experiments into their own models. I don't think that matters: whatever ideas are coming from humans in any lab are being automated.)

I don't think that matters: whatever ideas are coming from humans in any lab are being automated.

Not remotely true now. But when it becomes true, well, the side with more GPUs will move faster.

Still, Wenfeng is likely to be simply correct that "labs" get 0% of the light cone.

They use the threat of an AI-assisted superbug, which admittedly could cripple civilization, but so far is merely plausible. But that could effectively ban open models if it's implemented such that only closed models pass,

It's plausible, but we can't wait until someone makes the superbug to do something to prevent it.

I am extremely aware of the risks of an Anthropic/OpenAI duopoly, but I don't see an easier alternative to the superbug problem (which also has a parallel problem in cybersecurity). For the superbug, do we instead put enforcement mechanisms on the facilities and businesses that would be the obvious path in making a superbug? Maybe, but a sufficiently capable model (and open weights are probably only ~3 months behind the leading closed labs), it's likely possible to get around it. There are lots of places that attackers have significant advantages over defenders.

On the other hand, maybe it's for the best we have a disaster now killing a million people, to make us willing to take action that will prevent a a Chicxulub event. It's also unclear whether banning open weights would actually do much to prevent bad actors from doing the naughty.

At some point, surely this year, the first really big “happening” is going to happen to do with this. Then things get interesting. Still, I don’t think there’s anything anyone can do. This isn’t like nuclear proliferation, where really anybody except state level or wannabe-state-level (ISIS, etc) has much reason to go for it - oh you’re going to threaten to just nuke all your enemies? Come on - essentially everybody doing anything benefits from AI at some point, to some degree.

Open weight frontier models are indeed incredibly dangerous. Closed-weight frontier models are marginally less dangerous but significantly more classically dystopian. We are essentially weighing an increased probability of omnicide against an increased probability of objectivist utopia.

Millions of dollars of equipment, sure, but far more attainable than running GPT or Opus.

Technically speaking, you can run a 3T model on 25x DGX Sparks. Not fast, but it's around 150k if you include networking and power costs. If you really don't care about speed, you could probably get a couple second-per-token on older server hardware running fully CPU under 70k, though I wouldn't recommend it.

Allegedly some US officials are considering restricting US companies from using Chinese open models. Despite this claim being repeated across many outlets, I didn't actually find any evidence.

I'm also just not seeing a route to do it. Putting Moonshot on the BIS Entity List is the nuclear option, and it'd fuck up Moonshot's business and operations a lot, but I don't think it'd actually stop anyone from using their weights legally, and might not even do much to prevent them from doing future training. And I just don't see any better avenue.

Unlawful extraction from companies that have themselves unlawfully extracted? Again, really?

To be fair, a distillation attack typically involves massive numbers of smurfed requests, and doesn't have the fair use backing that normal transformative works would. I agree it doesn't make much philosophical sense, but it's not quite the same thing categorically.

But that could effectively ban open models if it's implemented such that only closed models pass, like how "nobody can sleep under a bridge" applies equally to rich and poor but only affects the latter.

There's also an issue where it might be only possible to pass without letting anyone access the model directly, regardless of how public. Pretty much every model is vulnerable to 'heretic' modification which changes the weights in minor ways to make it see any request as legitimate. These are widely available and, once discovered, can be performed by anyone with the weights and some hard drive space. It's mostly been helpful for writing smut in more 'circumspect' models, but there's already a lot of 100B-sized models that can independently discover serious dangers that are not well-known among the general populace (and, to be fair, some legitimate and safe uses of the model that also trigger those safeguards and thus are good reason to want a heretic model).

I doubt China will surpass American companies if they're training on American models, especially since America has more hardware.

Not so sure. Qwen3.6 was a major step forward compared to other models of the same time, and even now it remains a good example of good ideas being able to best raw hardware or parameter count to some degree. And there is the power-and-regulation issue, where especially if China decides to really hammer on the matter, the US advantage might not stay around.

I understand why large LLM developers are against distillation. It is essentially patent infringement, basically a way to copy their weights. It destroys their business model when someone can just wait until they finish the expensive training process and copy their results. But seeing as these corporations did not care about the intellectual property concerns regarding the data they used for training, I have a hard time seeing that they really have a patent to infringe upon. They neither produced nor owned the data they used for training. It follows then that they also don't really have a claim to their weights.

I have a hard time seeing that they really have a patent to infringe upon

They can't on account of them not having patented it in the first place.

I understand why large LLM developers are against distillation. It is essentially patent infringement

Honestly, given how much FOSS code they ate to train these models, they should be compelled by law to release everything about them.