site banner

Culture War Roundup for the week of August 31, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

3
Jump in the discussion.

No email address required.

I know what regulatory capture is. I'm asking how the type of regulatory capture being discussed could possibly apply to a product class that takes billions of dollars to bring to market in any scenario. Stop pattern matching and look at the actual fact pattern it's absurd. If a small business was able to bring a frontier model to market alone then we have a whole lot of other things to worry about.

I don't believe these regulations will only apply to the very top multi-billion dollar companies. As far as I've seen nobody has proposed any specific rules yet so there's plenty of space to build an expensive system that everyone in the AI space needs to participate in.

At the very least I expect a regulatory system that applies to frontier AI would need to check models to determine whether they count as frontier AI before they release. Otherwise we'd just be taking the word of companies that their model is totally fine and too dumb to count.

They call for "pacing the frontier" and all the previous proposed bills have had frontier scale compute reporting requirements. Essentially if you aren't using frontier levels of compute and don't believe you're pushing the frontier then you're exempted. We can go over any specific calls for state intervention if you want but as far as I know the labs are not proposing any regulatory framework that applies to smaller models.

I went for a look at specific numbers from Anthropic and bills they support and found numbers like 500 million annual revenue or 10^25 FLOPs. That doesn't quite sound like a multi-billion dollar requirement but it would certainly require a substantial business.

While I'm not entirely convinced this isn't an anti-competitive tactic, it does at least look like the two areas I raised a concern about (self-hosting or fresh startups) would be exempt from this regulation unless compute availability dramatically improves in the near future.

Does it really take billions of dollars? Wasn't Mistral 7B developed for $1 million or so, and Mistral Large for less than $100 million? Didn't DeepSeek or whoever launch a successful model with a low-end stated cost of a few million dollars? Google says that estimates are that they did that with a hardware investment of less than $2 billion, so I guess there's a number of different ways to creatively make the costs seem lower than the true cost of doing business...but if compute costs really are dropping it stands to reason that developing new models will be cheaper in the future, not more expensive...right?

I realize the obvious counter-response is "well but those aren't cutting edge" but as far as I can tell, most people using Anthropic and ChatGPT aren't using their cutting edge models. Thus is seems like edging out smaller providers protects the majority of the primes' customer base (although to be fair I am not sure how much of the majority of the primes' customer base is actually paying for their compute.)

but if compute costs really are dropping it stands to reason that developing new models will be cheaper in the future, not more expensive...right?

Compute costs are arguably not even falling, as the utility of compute is rising. See memory price hikes per GB, on the same process.

Very interesting.

From what I understand, a lot of AI boosters were saying that falling inference costs were going to help the primes out, so that's not necessarily good for them, I take it.

On the other hand, it probably does make it harder to get into the market.

You can make a tiny model pretty cheap - I've toyed with a 200m one for specialized purposes - but the bigger the model, and the longer the training, the more expensive it gets. There's been a lot of tricks developed so that training costs aren't exponential or super-exponential, and there's more to a model's intelligence than how many parameters it has, but the costs scale quickly.

DeepSeek's costs are also... probably not comparable. They say about 6m for DeepSeek v4, and it's probably not a lie, but it depends on a training architecture built for specialized silicon that you or I can't get, and highly optimized decisions about power costs. Meta 4 Scout at 12m for 109B parameters is probably more representative for anyone not directly supported by a major world power (which, tbf, is not a matter specific to Moonshot or China). Plus the whole MOE thing can make comparison to the real dense models

That said, while smaller models do fine for constrained projects, most have struggled pretty badly with things like tool-calling or multi-step problem solving. Mistral7b's pretty out-of-date as small models go, but Gemma4-12B is still stuck at 'well-read but not-bright intern'. Qwen3.8-Flash-Next can do some amazing stuff, but it's about Opus 4.6-grade if you're being generous. For a lot of purposes, that's fine -- the only one of these items it can't do is the 3d modeling one -- but users tend to notice where the model is brain-damaged as much or more than the broad areas it succeeds.

((I'm not convinced that the small models are safe, but they're probably not capable of hacking the NSA and definitely aren't close to serious self-improvement.))

DeepSeek's costs are also... probably not comparable. They say about 6m for DeepSeek v4, and it's probably not a lie, but it depends on a training architecture built for specialized silicon that you or I can't get, and highly optimized decisions about power costs

They don't, you mean V3. V4 is $17-30M. Also, what they specifically said then (and everyone was eager to first misquote them and then accuse them of lying) was:

Lastly, we emphasize again the economical training costs of DeepSeek-V3, summarized in Table 1, achieved through our optimized co-design of algorithms, frameworks, and hardware. During the pre-training stage, training DeepSeek-V3 on each trillion tokens requires only 180K H800 GPU hours, i.e., 3.7 days on our cluster with 2048 H800 GPUs. Consequently, our pretraining stage is completed in less than two months and costs 2664K GPU hours. Combined with 119K GPU hours for the context length extension and 5K GPU hours for post-training, DeepSeek-V3 costs only 2.788M GPU hours for its full training. Assuming the rental price of the H800 GPU is $2 per GPU hour, our total training costs amount to only $5.576M. Note that the aforementioned costs include only the official training of DeepSeek-V3, excluding the costs associated with prior research and ablation experiments on architectures, algorithms, or data.

R1 was additional $280K or so, according to a later paper. Current RL runs are much more expensive.

Anyway you're correct on the core point that this isn't a game for small actors. Tens of millions in GPU-hours is just the entry fee.

Ah, thanks for the correction. I took the first value off Google, and should have known better. 30m still feels low for the core costs of V4, even with the architecture benefits, but that's plausible low rather than 'not cost of electricity'.

We don't have a lot of data points for costs. But for what it's worth, V4-Pro-Preview scored 44 on Artificial Analysis. Semianalysis says of Korean Motif:

Motif is a sub 30 person startup that released their first pre-trained model, which was only 2.6B total parameters, last June. They’ve raised just $17M and have access to a mere 768 B200s (< 2MW).

They have reached 47 with Motif 3. Its pretraining budget is ≈10x below V4 (given active parameters x tokens). It's definitely doable to train something like V4 for under $30M. I mainly hedge because of apparent troubles DeepSeek had with Huawei.

Huh. I know Motif's gotten mixed reviews, but even assuming it benchmaxxed a little, that's still well outside my intuitions. And A lot of the South Korea stuff is explicitly competitive to get government support, so that excludes the free electricity option. Guess I'm behind the power curve, here.

That said, Motif lost the competition to chaebols, for usual anti-meritocratic Worst Korea reasons.

It looks like people have been able to train models large enough to trigger proposed oversight without spending billions, though. I'd totally buy that they aren't good, but if they aren't good and they are still good enough to trigger oversight then the entire argument here that the only point of the regulations being proposed is to monitor good models is just wrong.

Do you think that inference costs are dropping quickly enough to make a difference in training-large-model costs or nah?

Wasn't Mistral 7B developed for $1 million or so, and Mistral Large for less than $100 million?

Mistral is not a frontier lab and no model they've ever trained would trigger any of the proposed regulatory scrutiny.

Didn't DeepSeek or whoever launch a successful model with a low-end stated cost of a few million dollars?

No they didn't. The number you're thinking of was the money they spent getting an already existing model to work with chain of thought.

but if compute costs really are dropping it stands to reason that developing new models will be cheaper in the future, not more expensive...right?

Ok sure, but in the coutnerfactual world where there is no regulation the current frontier labs will have substantially better models.

I realize the obvious counter-response is "well but those aren't cutting edge" but as far as I can tell, most people using Anthropic and ChatGPT aren't using their cutting edge models.

If they're not cutting edge then the regulations do not apply. They're trying to pace the frontier, not dinky merely hundreds of millions of dollars models.

Mistral is not a frontier lab and no model they've ever trained would trigger any of the proposed regulatory scrutiny.

I think you are wrong about this - Mistral Large and Large 2 were believed to be past the 10^25 FLOPs trigger proposed by Anthropic at the beginning of 2025. So was Duobao-Pro by ByteDance and Pangu Ultra by Huawei.

And by the way, (as per a quick Google) Mistral's operating on a total of $3 billion in funding, and made Mistral Large after raising less than half a billion. So you don't need billions to hit the EU (and apparently Anthropic's) "systemic risk" category.

Either Mistral is a frontier lab or Anthropic's proposed regulation would apply to non-frontier labs. Either way, the thrust of your arguments here seems to be based on a misunderstanding.

Yes. There's a lot of money around. Getting a billion dollars, or 10 billion dollars, is much easier than getting past a regulator controlled by the firms you want to compete with.

So the dastardly plan here is that the frontier labs destroy their own leads on the theory that they'll be able to control the regulatory state so thoroughly as to choke out all other competition and avoid becoming a commodity. Except when the frontier is paused they'll instantly become a commodity with several other players in the same tier already. This just doesn't make any sense. If progress stalls then the margins on current models get competed down to the marginal inference cost with or without other market entrants. Half or more of the nation hates these companies, the idea that a regulatory state is going to be particularly sweet to them is hard to imagine and you're making a general argument against all regulation. I'm sorry but one way or the other we're going to be regulating the production of models with nationstate level hacking ability. The public and state will not stand for any random with a bone to pick having the ability to shut down the power grid.

the idea that a regulatory state is going to be particularly sweet to them is hard to imagine

They don't need to be explicitely sweet to them, but the regulation when it gets written up, will be written with the realities of the current market in mind, which is to say it will be written in a way that won't strangle the entirety of the profits out of the the current incumbants to comply with, very little thought will be given to future competitors because the shape of these competitors, and any way their technology or organisational structure might be different and not fit well with the regulation, is currently unknown.

you're making a general argument against all regulation

Yes, he is, and? This is a valid general argument against all regulation (with perhaps the exception of anti-trust regulations). That doesn't mean that the tradeoff is never worth it, but regulation in general has that flaw.

That doesn't mean that the tradeoff is never worth it,

Agreed. To be clear, I am not arguing that there should or shouldn't be AI regulation. My position in this discussion is simply that it's reasonable to hypothesize that leaders in the AI field support regulation, in whole or in part, for selfish reasons, namely that it will hinder potential competition.

So the dastardly plan here is that the frontier labs destroy their own leads on the theory that they'll be able to control the regulatory state so thoroughly as to choke out all other competition and avoid becoming a commodity.

That's kind of how regulatory capture works, yes.

But it clearly cannot work in this case for the reasons I've listed. So the idea that they're going to try to do that instead of continuing to the trajectory that has them rising faster in value than any other companies in history is hard to justify on grounds other than the belief that going forward without regulation is legitimately dangerous. These are at the very least dual use technologies, the idea that we're not going to regulate them is crazy.

The "reasons you have listed" are just you assuming your own conclusion.

The entirety of your case is assuming a boat load of contradictory incentives. You've not bothered to actually make a case for how this regulatory capture could work or why it would be preferable to just continuing the status quo.

Because winning pretty big for all time is much better and more legible for IPO than a small chance of winning very big before being beaten out by a competitor who doesn't have to service your massive and entitled customer base or deal with your office politics and legal woes.