site banner

Culture War Roundup for the week of July 27, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

One angle I rarely see discussed is that open weight models are inherently decelerationist for the frontier. They greatly reduce the profit motive to push the frontier, especially if distillation is allowed because the make it much harder to recoup the investment thus reducing the capital available to invest. Ultimately though they're definitionally unsafe and can never be made safe, any work put in to align them can be sanded off with post training and even with just mundane levels of uplift the idea that the offense/defense equilibrium will favor the defender in most areas, let alone all areas is just impossible to believe.

even with just mundane levels of uplift the idea that the offense/defense equilibrium will favor the defender in most areas, let alone all areas is just impossible to believe

If you're just referring to different areas of software security, the long-term equilibrium is decidedly in favor of defense. Bugs, especially exploitable bugs, are a consequence of the fact that on any objective scale our programming languages still suck (especially the fact that the ones which suck less on security tend to suck more on performance) and our programmers suck (yeah, including me, sorry), not because it's actually impossible to write a program that does what you want but doesn't also let you get p0wned as soon as someone figures out just the right corrupt input data to send. In the meantime, before we have languages that aren't larded with Undefined Behavior pitfalls and/or the ability to cheaply write programs that never stumble into exploitable pitfalls, initiatives like Project Glasswing are a pretty good way to keep the short-term equilibrium in favor of the defender too.

If you're also thinking about e.g. biological security ... well, yeah, there is a chance we'll all be dead soon. I'd like to hope that, since our immune systems are pretty versatile, maybe natural pathogen evolution is already stress-testing us near the limit of what a bioweapon could do, and in the worst case maybe AI could quickly figure out a vaccine to prepare us in advance of infection by any even-more-dangerous inventions ... but "we can make software good enough" is practically a theorem, whereas "we can make immune systems good enough" is more of a prayer. That prayer would have to be answered, not just for humans directly, but for all the life that humans depend on. Florida's orange production is down over 90% since the first spread of "greening disease" there, and we've had decades of inability to fix it, and if something similarly unstoppable ever infects corn, rice, and/or wheat too then there are going to be a lot of starving people.

Bugs, especially exploitable bugs, are a consequence of the fact that on any objective scale our programming languages still suck (especially the fact that the ones which suck less on security tend to suck more on performance) and our programmers suck (yeah, including me, sorry), not because it's actually impossible to write a program that does what you want but doesn't also let you get p0wned as soon as someone figures out just the right corrupt input data to send.

This is heavily disputed to put it lightly. The ask here is something like all software everywhere be written/rewritten incredibly defensively and continuously upgraded as the models get better at finding and exploiting vulnerabilities. Even hardening what we have now, even with mythos helping via glasswing, this is long term project. The trouble is that the defender has to win every fight in every domain for every surface, while the attacker needs to only win once. The effort disequilibrium is massive.

initiatives like Project Glasswing are a pretty good way to keep the short-term equilibrium in favor of the defender too.

Yes, and I've personally read reports on my software that have gone through this program. Importantly the defends have better tools here.

If you're also thinking about e.g. biological security ... well, yeah, there is a chance we'll all be dead soon.

Yes. Seems pretty bad. I personally do not want to die.

They greatly reduce the profit motive to push the frontier, especially if distillation is allowed because the make it much harder to recoup the investment thus reducing the capital available to invest

That's the argument of a few OpenAI goons, namely roon, Will Depue and most articulately Dean Ball:

Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models… One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end.

It's also innumerate bullshit from nervous, greedy bagholders. Since I've blocked you on X (deservedly, seeing the unearned confidence in the latter half of your post), you probably haven't seen my multiple posts on this, but to sum it up: margins on inference at frontier labs are obscene, frankly if people understood this, AI lab CEOs would be at nontrivial risk of vigilante justice from mobs with pitchforks. The talk about bubbles, subsidized subscriptions, capital to reinvest into R&D – this is all one big pile of bullshit obscuring unhinged rent extraction, which is viable because demand still exceeds supply, because they physically cannot mint enough tokens with gigawatts of deployed capacity. Frontier AI is wildly profitable, it has high software margins, and Chinese distillation is powerless to dent that.

Let's do the math. DeepSeek comes to the rescue, as usual, since they are 1) all time champions of low API pricing and 2) the only ones who periodically disclose some of their economics (starting in the V3/R1 era). Soon after V4's deployment, they have cut prices by 4x and started using DSpark, an advanced lossless speculative decoding technology that «accelerates per-user generation speeds by 60%–85% at matched throughput levels». They've open sourced it, so even if the frontier wasn't yet on that level (dubious; they are technically overrated and from what I hear arguably behind China on infrastructure and generally on engineering culture, but not really incompetent), it is now. It's general-purpose (eg see here how it's adapted to accelerate Kimi K3, a wildly different architecture, by up to 4x). Obviously frontier labs can customize it further for their scenarios, and not tell us about it.

DeepSeek uses it to get average ≈2000 tokens/second/GPU decode for V4-Pro at "decent speeds" (≈70 tokens/second per user stream), in live traffic. Presumably their inference GPUs are something like H200 or Ascend 950DT, with whatever better stuff they have smuggled reserved for training and experiments (at the time of publishing, they had 20K H100-equivalents total). Prefill and cache hits are much cheaper but we can assume they're priced roughly proportionally to decode in wall clock cost, this follows from kv cache difference in volume and pricing between V4-Pro and V4-Flash and is also obvious from timing (eg it takes their first party API seconds to process a 500K input). So what does this get us? At this utilization, 500 seconds of GPU-time for 1 million tokens, priced at $0.87.. $6.26 per hour. $54862 of revenue per year. If we use the usual rent cost of $2 per hour of a Hopper generation GPU, we get $37.3K of pure annual profit. If we assume it's owned and the per-GPU consumption is 2 kW (very high), that's just $1.7K of electricity costs at Chinese industrial pricing. It's mostly owned now, but they'll be renting to expand. So, something like $45K/year of average profit per GPU. At sub $1/million output tokens. Is this plausible? Well, in a recently leaked investment call, Liang Wenfeng insists that his target is 10 months to breakeven on hardware: «91. Our API pricing is designed to generate a reasonable profit. Roughly speaking, if we buy a batch of equipment on the market, recovering the investment in about ten months represents a reasonable level of profitability … » 10 months is $37K of profits as per the above. H100 tier cards were going for $35K a piece recently, though there's more scarcity recently.

In other words, DeepSeek for all this Confucian restraint has something like 70-84% profit margin. Well would you look at that, The Information reports: «"The gross profit margin from selling cloud based access to V4 is somewhere between 70 to 80%."» I don't have the subscription, but seems like being able to do 1st grade arithmetic is enough. By normal business logic, «10 months to breakeven» allows you to raise multiples for the next round of capex, which is exactly what DeepSeek successfully did, and will do again soon, and what all American frontier labs have done repeatedly.

How big are frontier models? I think they're not big, judging by knowledge coverage which is the most honest indicator of scale, due to physics of LLMs – eg Opus 4.8 is in the same band as Kimi K3, V4 and GPT 5.6 Terra. How costly are they to produce? For V4, we have an idea of training costs: 49B active, 33T tokens, 6ND = 9.7e24 FLOPs. V3 was 37e9*14.8e12*6 = 3.3e24, and trained over 55 days on 2048 H800s, giving us 2,664M GPU-hours ($5.3M) and a reasonable (for large sparse MoE) MFU of maybe 25% (depending on how you count mixed fp8, 20-40%; this is genuinely annoying). Anyway, we can say that V4 is 3 times larger = $15M pretrain at rental costs. Sparsity is informed by scaling laws and inference optimality, and is likely 3-4ish % across model lines (frontier models most likely use the same DSV3-type designs). So, Kimi K3 has twice the active parameters of V4, at 104B; Opus is likely similar; Claude Fable may be twice that again, something like 200B. GPT 5.6 Sol – 150B? That's roughly what I get from insiders. So these models still cost at most low hundreds of millions to pretrain, likely under <100M. They also are more overtrained (GPT-OSS hinted at 60-120T pretraining tokens vs 30-50 in Chinese models), with more rigorous and intense post-training (10% of pretraining cost vs 50-100% or more in the US; Cursor threw 7x more FLOPs at Kimi K2.5 RL to make Composer 2.5, though this is crude Muskian maximalism in a desperate attempt to catch up), so they don't need to be large to perform better in downstream tasks. They also operate on a vastly larger scale, so can do even more efficient fleet level optimizations. All in all, a frontier model production goes for <$500M today. Anthropic's ARR reached $47 billion in May and continues to explode. So let's talk of reinvestment. Fine! Let's say, for the sake of argument, that Americans are lower IQ, can't do math, are unable to capitalize on their years of first mover advantage, on staff that had founded the whole field, on reams of user data, unwilling to waste time to optimize for costs, whatever – so they're a whopping 2 times less efficient than the frugal (and distilling) China. So what do we get? $50/1M tokens with Fable, $30 with Sol, $15 with Terra. They also have access to radically more efficient Blackwell hardware, soon Rubin (also TPUs, Trainium, Cerebras etc.), which more than makes up for the increase in compute demands per token/second.
Do you see where I'm going with all this? Their API margins are over 90%. They get payback on their compute in under Wenfeng's 10 months. Their "generous" subscription limits do very little to offset this. There's no "subscription war", it's kayfabe and a price fixing cartel. The only reason they're in the red is that they are, indeed, reinvesting tens of billions into scaling compute for R&D and larger-scale training. How much capital do they need to keep accelerating? Hundreds of billions? Trillions? Mind you, algorithms and hardware are constantly improving (stuff like DSpark is pretty mundane, and will get trivial in the age of AI doing AI – which even K3 is already doing). And they're still nowhere close to exhausting the demand, even as they keep jacking up prices. Frontier intelligence that can automate competitive white collar labor in the US is priced against not even the hour of labor but against its revenue, and would make economic sense even at costs an order of magnitude higher, and that's where we're going. How much revenue do you generate in an hour of work for your employee?

Enough. This is all noise. All plausible negative impact of Chinese labs, distillation or whatever, can be negated by roughly 2 months of regular progress. The only serious question is how the multi-trillion valuation pie gets sliced – how much goes to "labs", how much to Jensen, and how much to clouds serving a mix of closed and open weights. The championing of open weights comes from the clouds and Jensen, because he benefits from maximum demand for CUDA systems and doesn't want to help OpenAI et al. to switch to in-house compute and cut him off (which they are trying to do, and have the capital for). The bullshitting about "deceleration" comes from the first camp, because they're SF dorks high on their own supply and want their equity to get high enough to buy entire landmasses with genetically engineered catgirl slaves (if not star systems in the coveted "light cone"). My opinion is that the rest of us can play the world's smallest violin if they miss those IPO targets; American AGI will arrive on schedule regardless. Finally, as Wenfeng correctly observes from Communist China, it's a political impossibility that any self-appointed «AGI lab» would be allowed to capture 10% of the world's or even the nation's GDP; so he pursues «restraint» and «reasonable profit» to preclude the painful enforced correction that the Chinese tech/finance scene is familiar with. The logic in the US, however, will be similar.

Since I've blocked you on X (deservedly, seeing the unearned confidence in the latter half of your post)

Honestly man, I had a lot of respect for you and always approach in good faith, part of me still hopes this is some kind of crazy misunderstanding of mistaken identity. When you blocked me you implied I was a Trump supporter or something, which is just confusingly not true. And if smug confidence is a sin, well I guess you've already cast the first stone.

margins on inference at frontier labs are obscene

Yes, I've argued this point before on this forum several times against people convinced the frontier labs were losing money even on inference.

Frontier AI is wildly profitable, it has high software margins, and Chinese distillation is powerless to dent that.

No, actually, sating demand does in fact dent profitability of selling a scarce resource even if it only goes from extremely obscenely profitable to merely obscenely profitable, and the projected profitability is what dictates investment levels. A competitor that copies your R&D at a lag shifts your optimal resource allocations from research to buildout optimization. Research is what accelerates the frontier thus Chinese distillations slow acceleration.

> The only reason they're in the red is that they are, indeed, reinvesting tens of billions into scaling compute for R&D and larger-scale training. How much capital do they need to keep accelerating?

This is literally the whole of the argument. Nothing else in this post matters at all.

Enough. This is all noise

agreed.

All plausible negative impact of Chinese labs, distillation or whatever, can be negated by roughly 2 months of regular progress.

Wait, how? do you mean 2 months of the ability to reap high margins with paused competition?

The only serious question is how the multi-trillion valuation pie gets sliced – how much goes to "labs", how much to Jensen, and how much to clouds serving a mix of closed and open weights.

These are all considerations that people consider carefully when they decide how many trillions of dollars to put into building out datacenters. Dario famously underbuilt in the last cycle on these mundane economic concerns.

On the more important topic though, and you're certainly allowed to only reply to the part of my post you disagree with, but I'm much more concerned with the, clear to me fact, that open weight models are unsafe and nothing can make them safe. The deceleration thing might even be good given this. Even given just mundane uplift I can't seem to get away from the offense/defense equilibrium disruption. Those are heavy dice to roll.

How much capital do they need to keep accelerating? Hundreds of billions? Trillions?

Scale is all you need. We are at the beginnings of RSI. Anyone (except, apparently, Google) can build a frontier model, capable of research, for the chump change of $1B, or likely less. An extra $10B, though, and you can build it and its successor faster. Promises of the light cone let you raise that money easily, but threats to have to share it make investors start asking (somewhat irrationally) how much of the light cone they get. That slows you down. Not that OAI and Anthropic are exactly hurting for capital, but at the margins it makes a decelerationist difference.

All plausible negative impact of Chinese labs, distillation or whatever, can be negated by roughly 2 months of regular progress.

Curious how you got that number, though I agree any delay they introduce is on the order of months. Even two months, though, is a long time nowadays.

(The strongest argument for a Chinese accelerationist effect is the closed labs can easily incorporate their research and experiments into their own models. I don't think that matters: whatever ideas are coming from humans in any lab are being automated.)

I don't think that matters: whatever ideas are coming from humans in any lab are being automated.

Not remotely true now. But when it becomes true, well, the side with more GPUs will move faster.

Still, Wenfeng is likely to be simply correct that "labs" get 0% of the light cone.