This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Perhaps because they anticipate or hope that the applicable regulations would impose a lot of bureaucratic hurdles and red tape on new players in the industry.
There are already multi-billion dollar hurdles for new players in the industry to put together enough compute to train a frontier model, the idea that having a legal department is going to be the difference is absurd. This isn't a haircutting industry.
Big companies start off as small companies. Anti-competitive regulation rarely targets the hard-to-kill large companies and instead strangles the small companies to ensure there's nobody around to grow into a big company.
Frontier models also aren't their only competition. Anthropic and OpenAI have both shown a desire to raise prices which is going to make budget and self hosted models more appealing alternatives as they do. Adding large costs to everyone in the industry doesn't just put a big barrier to entry up but also brings budget options closer on price to the big players.
I know what regulatory capture is. I'm asking how the type of regulatory capture being discussed could possibly apply to a product class that takes billions of dollars to bring to market in any scenario. Stop pattern matching and look at the actual fact pattern it's absurd. If a small business was able to bring a frontier model to market alone then we have a whole lot of other things to worry about.
Does it really take billions of dollars? Wasn't Mistral 7B developed for $1 million or so, and Mistral Large for less than $100 million? Didn't DeepSeek or whoever launch a successful model with a low-end stated cost of a few million dollars? Google says that estimates are that they did that with a hardware investment of less than $2 billion, so I guess there's a number of different ways to creatively make the costs seem lower than the true cost of doing business...but if compute costs really are dropping it stands to reason that developing new models will be cheaper in the future, not more expensive...right?
I realize the obvious counter-response is "well but those aren't cutting edge" but as far as I can tell, most people using Anthropic and ChatGPT aren't using their cutting edge models. Thus is seems like edging out smaller providers protects the majority of the primes' customer base (although to be fair I am not sure how much of the majority of the primes' customer base is actually paying for their compute.)
You can make a tiny model pretty cheap - I've toyed with a 200m one for specialized purposes - but the bigger the model, and the longer the training, the more expensive it gets. There's been a lot of tricks developed so that training costs aren't exponential or super-exponential, and there's more to a model's intelligence than how many parameters it has, but the costs scale quickly.
DeepSeek's costs are also... probably not comparable. They say about 6m for DeepSeek v4, and it's probably not a lie, but it depends on a training architecture built for specialized silicon that you or I can't get, and highly optimized decisions about power costs. Meta 4 Scout at 12m for 109B parameters is probably more representative for anyone not directly supported by a major world power (which, tbf, is not a matter specific to Moonshot or China). Plus the whole MOE thing can make comparison to the real dense models
That said, while smaller models do fine for constrained projects, most have struggled pretty badly with things like tool-calling or multi-step problem solving. Mistral7b's pretty out-of-date as small models go, but Gemma4-12B is still stuck at 'well-read but not-bright intern'. Qwen3.8-Flash-Next can do some amazing stuff, but it's about Opus 4.6-grade if you're being generous. For a lot of purposes, that's fine -- the only one of these items it can't do is the 3d modeling one -- but users tend to notice where the model is brain-damaged as much or more than the broad areas it succeeds.
((I'm not convinced that the small models are safe, but they're probably not capable of hacking the NSA and definitely aren't close to serious self-improvement.))
They don't, you mean V3. V4 is $17-30M. Also, what they specifically said then (and everyone was eager to first misquote them and then accuse them of lying) was:
R1 was additional $280K or so, according to a later paper. Current RL runs are much more expensive.
Anyway you're correct on the core point that this isn't a game for small actors. Tens of millions in GPU-hours is just the entry fee.
Ah, thanks for the correction. I took the first value off Google, and should have known better. 30m still feels low for the core costs of V4, even with the architecture benefits, but that's plausible low rather than 'not cost of electricity'.
We don't have a lot of data points for costs. But for what it's worth, V4-Pro-Preview scored 44 on Artificial Analysis. Semianalysis says of Korean Motif:
They have reached 47 with Motif 3. Its pretraining budget is ≈10x below V4 (given active parameters x tokens). It's definitely doable to train something like V4 for under $30M. I mainly hedge because of apparent troubles DeepSeek had with Huawei.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link