This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
The only reason I'm not saying it's over is because I've said it before. It's been over for a while. And we've barely begun.
The good news, for mathematicians, is that their employability (in academia) did not hinge very strongly on tangible economic output. Nobody funds an algebraic geometer expecting a return on investment, which means nobody can defund one on the grounds that a datacenter in Texas now offers a better one.
The bad news? Everything else.
Just look at this fucking sweep. Just look at it. Hilbert's tenth over the rationals. Rational Hodge for CM abelian varieties. Integer multiplication below n log n, which I had mentally filed under "the floor, go home." A quasi-Riemann hypothesis thrown in like a free tote bag. Any one of these would have been the defining result of a human career, and they were released as a batch, in a GitHub repo, on a Tuesday, alongside hundreds of others.
And it wasn't brute force at absurd cost. OpenAI says the average result used roughly three hours of ChatGPT Pro-equivalent compute. I agree with your Pareto intuition, but the deeper point is that three hours is a price, and prices in this industry fall monotonically. Gwern spent 2020 arguing that neural nets would keep absorbing compute and sprouting new abilities, while noting the idea was so unpopular that it would only be accepted as a fait accompli. Well, here is the fait, and it is quite accompli. Next year the same theorem costs twenty minutes. The year after, it's a rounding error on someone's API bill.
Your own experiment deserves more attention than you gave it. As others have already done to good effect, your research methodology consisted of being a motivational poster ("you can do it!", "I believe in you") and you got a materially stronger result in half an hour. The binding constraint has shifted to elicitation and persistence.
Some might claim that human mathematicians are necessary to humanize and understand the proofs, or to come up with new frameworks and fields of research.
Maybe. I don't know. The strongest version of the argument is that mathematics is a conversation between humans about what humans find interesting, and a proof nobody understands is a tree falling in an empty forest. I'm sympathetic to that! I also notice that this position has been retreating to progressively smaller hills since the first computer-assisted proofs, and each hill gets defended with the same conviction as the last.
I do know that the current state of affairs is unlikely to last for more than six months anyway. Just get the models to come up with their own novel conjectures, or to attempt an overarching synthesis. I don't think turning abstruse Lean proofs into something human-readable will be particularly difficult, though human readability has long ceased to be a major concern for anyone. Note that OpenAI is shipping Lean formalizations for many of these proofs, so "trust me bro" is off the table; a type checker cares nothing for anyone's feelings. We're at the point where people are coping that the tsunami of new discoveries hinge on hither-to undiscovered bugs in Lean. You wish.
Every day, I feel the pace of progress accelerate, and then the practical ramifications show up in my life a few months later. I used to think of this as a lag. These days it feels more like the gap between seeing lightning and hearing thunder, and the gap keeps shrinking.
We've just had Scott Aaronson announce a new UT Austin course, CS395T: AI Alignment Theory, with Yudkowsky's AGI Ruin as the first assigned reading. To quote:
I'll give credit where it's due, but also: when the designated Reasonable Skeptic of the rationalist diaspora concedes the point in writing, the Overton window has relocated.
We've also just had an American company, Nolla Health, receive regulatory approval in Utah for AI to issue initial prescriptions. Yes, it's acne. Yes, the model can only choose from a short list of physician-approved topicals, and clinicians reportedly agreed with its recommendations in over 96% of cases. Everything starts as acne cream. Utah let Doctronic's AI handle prescription renewals back in January; nine months later, it's writing first-line scripts. You can extrapolate the line yourself. That's the harbinger of medicine's fall, as far as I'm concerned. I'm impressed we held the moat this long.
At least I've got a moat. The NHS is famously a slow ship to steer, and sclerotic at that. The future arrives unevenly distributed, and we can count on the service the last postcode on the delivery route. For once, I find that reassuring <3
Good career choice, past self_made_human. Turns out there are concrete benefits to taking theoretical but plausible concerns seriously, and preparing accordingly.
If I understand your comments below, this is actually false, you didn't have any expectations about the bounds for this problem. Presumably this claim was just AI generated.
The improvement in the paper goes to n(log n)^(1-(2^-182)), which while kind of neat is essentially the same as n log n (unless we're multiplying numbers on a computer larger than the universe). It's a toy result. Your amazements suggests you probably didn't even look at the paper.
In other words, I consider your thoughts on this totally untrustworthy. Please don't use AI to write your posts.
Ahem. Please look at this comment I had left a scant few minutes back.
https://www.themotte.org/post/3960/culture-war-roundup-for-the-week/485779?context=8#context
I am well aware that this is practically indistinguishable from O(n log n). I know how exponents work. I haven't checked yet, but a sufficiently ridiculous constant factor would make it even more impractical. That's been known to happen.
I had read mathematicians, on Twitter, expressing surprise that we went below the previous SOTA by any margin, no matter how minuscule. The model rephrased my notice of secondhand surprise into a first person version.
That is such an innocuous, irrelevant change that I wouldn't have bothered to correct it for it's own sake. I hadn't even told the model to only wrap my references in links. It was entirely at liberty to add minor context. The only reason I point it out is because I respect @2rafa, and wanted to declare it as a concrete example of a phrase I hadn't hand-typed.
Totally untrustworthy? What a massive overreaction. A totally untrustworthy person would deny the whole thing. Instead, I'm punished for admitting any use. You're lucky I think the price of honesty is acceptable.
If you write me off, that's your problem rather than mine. Particularly since I didn't use AI to "write" my post, I had already written a post, and I threw into it for minor improvements and explicit citations I didn't have the time to source. By my standards, the changes I've quoted from the original are positive or at least benign.
More options
Context Copy link
If for the longest time the best exponent was some nice number like an integer or rational with small denominator, and the algorithm reaching it relatively simple, then even the tiniest improvement is notable. That it never performs better in cases small enough to be encourted in the real world, is not what mathematicians care about. Were it so, Big O notation would not be as common, and ultrafinistism would be the majority position.
The best exponent was some nice number because those are easier for humans to develop proofs for.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I see no reason why you couldn't replace the hype man with a cheap AI either. One LLM tries to find the proof. The other is instructed to hype it up. No human involved, other than in deciding what proof to look for next.
They do, I've heard of it being done for at least 2 years.
More options
Context Copy link
I consider that so obvious I didn't find it worth saying. Unless, for some stupid reason, the models learn that they're being prompted by print.ln instead of a human.
Given that reasoning effort is something that can be adjusted (usually by simply a privileged part of the prompt, rather than some fancy dial), it's moot.
More options
Context Copy link
More options
Context Copy link
Who? Aaronson? The "took his family to Disneyland and aborted the trip after having a meltdown over seeing people not wearing masks" guy? He's many things, but nobody is designating him any combination of "reasonable" and "skeptic".
(By the way, concurring with @2rafa's suspicion below, after you have acquired your reputation as an unapologetic meat proxy, it does feel extra suspicious when you make such catchy irresistible logit-attracting utterances whose substance seems bizarre when you actually think about it.)
Some people have forgotten that I'm capable of sarcasm.
I'm more than aware of Aaronson's quirks. In some regards, he's a pathological quokka, and in other cases, he's had a rather overblown threat response and reacted out of proportion. See his posts about his fear of anti-semitism for the latter, or the fact that he inspired the other Scott to write "Radicalizing the Romanceless."
In the Rationalist sphere, he's the one who's always painted him as the reasonable person, not prone to theoretical flights of fancy.
Anyway, unapologetic? Not at all. I just think it's a tradeoff I'm willing to make, even if it annoys you.
pushes up glasses
Well actually, he inspired Scott (PBUH) to write “Untitled”, which came out a year or so after “Radicalizing the Romanceless”. Both absolute bangers, though.
I am ashamed of my error, and will commit sudoku out of repentance.
(That's what I get for not running a throw away comment through an LLM for fact checks. I'm more likely to 'hallucinate' than they are.)
Post the sudoku when done so we can confirm you are a man of honor and good repute.
More options
Context Copy link
I lold.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Have you heard about Poe's Law? Sarcasm doesn't work when it's equally plausible that you would say the same thing in earnest, and attaching a clichéd epithet like that without any deeper reasoning behind it is the sort of thing an LLM would do in earnest.
All's well. You're willing to annoy me with AI content, and I'm willing to annoy you by derailing your threads with speculations about AI content :)
I consider that a reasonable compromise for all involved. It's the people who think they have the right to dictate how I spend my free time that I take umbrage with. I otherwise aim to be honest if anyone explicitly asks.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Sure, AI may be embarrassing humanity's proudest intellectual achievements. But there's a bright side. We can finally stop arguing about how AI is an overhyped bubble that's about to burst! Haven't even seen the term "stochastic parrot" in a while.
Another plus: AI can now compose its own Disney villain song. I'm not exactly sure where "writing a catchy diss track against people who think you're not intelligent" lies on the spectrum of intelligence, but it's probably a little ways past the Turing Test.
We aren't arguing about it because there's no point. At this point, both sides have tired of trying to convince each other, each still thinks they are right, and each still thinks the other is willfully ignoring evidence that disproves their POV. What's there to argue about any more? Life's too short.
More options
Context Copy link
I'm sure we can do that, but maybe we shouldn't.
There's been some pretty not-great signs for the financial state of the industry lately: Microsoft looking to wean off of Anthropic, Harvey and Thompson Reuters moving to their own models, the force majeure notice by Oracle in New Mexico, the failure of SB Energy, Holtec, and Aggreko to IPO (not to mention OpenAI and Anthropic!), Firmus missing its rental payment, and generally increased skepticism on the part of investors.
I think it's been a huge problem for thinking clearly about the situation that the financial concern regarding the AI industry has been latched onto by the worst AI skeptics, who tend to flatly deny the capabilities of the models.
Meanwhile on the flip side, I suspect there may have been a parallel problem that the people who are the biggest boosters of the technology are the ones most likely to be suffering from mild AI psychosis from talking with them all the time.
Thus, the AI debate has been, somehow, between "AIs suck and there is a massive bubble" and "AIs are the best thing since sliced bread and nuh-uh," which excludes two entire quadrants of possibility from the conversation, "AIs suck and will be profitable" and "AIs are good and there is a bubble."
This state of discourse creates epistemic closure on the topic in a way that I think is preventing a lot of people from assessing where, exactly, AI is sitting.
The word "bubble" is hurting more than it's helping here. It's beyond question that the money spent on training frontier models will produce an absolutely massive amount of value for the world. But it's still to be seen how much of that value can be captured by the labs themselves. This isn't really the same shaped as tulip mania or people paying absurd amounts of money for pets.com. If the labs fail economically(Which I think could only happen if scaling laws collapse roughly tomorrow) then the weights will still be around and still worth quite a bit to serve on the hardware which it will still be around and be worth it to keep running inference on. The labs could fail but they wouldn't be failing because AI is overhyped or a scam, just that running a research company in a competitive environment without state granted monopolies on the produce of your research is a brutal business to be in.
It seems pretty plausible that the frontier labs could fail simply because demand doesn't meet their projections. Anthropic, for instance, has something like $200 billion in commitments to Google and Amazon out to 2036 where, according to the terms of the deal, they pay regardless of demand.
There's already signs that demand may soften for Anthropic specifically (the OpenRouter trendline is switching away from Anthropic towards open source models iirc, Microsoft is looking to cut their spend with them, ditto (we can infer) Harvey and Thompson Reuters, Astra is apparently universally beloved by coders). If people start switching from Anthropic (and it doesn't take many: keep in mind that 80% of Anthropic's revenue (like OpenAI's) comes from 1% of their users, and 2 of Anthropic's customers generated about 25% of their 2025 revenue), they can either
Except they can't cut a lot of their spend (as per above), if they raise prices, OpenAI eats them anyway, and perhaps Astra or OpenAI engineers are good enough that they are simply locked out of producing a superior product. That leaves #4, die.
Maybe this seems good for OpenAI, except that if the timing is bad it sends the market into an AI panic and could spoil their IPO, and OpenAI is also on the hook for a bunch of infrastructure bills (although they may have structured them more flexibly, I'm not sure offhand). And of course there are a lot of other things that could also ruin an IPO: another pandemic, major war breaking out in Europe or the Pacific or Middle East, political unrest: pretty much any little thing that goes wrong and tightens the belt could crack up the revenue stream, OpenAI is not a profitable company, and the open-source models are nipping at its heels.
Yes, I agree (and have said before) that "AI is not going anywhere." I agree this isn't the tulip mania. But we're in a subsidized era of AI right now, and it may look very different once that subsidy ends.
Perhaps the frontier labs fail for some other reason (or don't fail at all) but I think it's pretty fair to say that AI has been overhyped. OpenAI said it was going to spend $1.4 trillion on infrastructure by 2030. That's hyping. Then they slashed their public infrastructure commitments by more than half, to $600 billion (because their CFO was worried that they were overhyping), although they've brought the number back up since to $750 billion.
If they actually revise their numbers back up to $1.4 trillion in 2030 and meet that infrastructure goal, feel free to ping me and I will agree that this was a bad example. Then I will point you to the 2027 Project where it postulates that the robots would kill us all by now, as my fallback example.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
There are still holdouts, but at this point I regard them with more psychiatric curiosity and pity than I do as entities to reason with.
As an analogy: people who ignore the insurance on their beach side property shooting up vs those ignoring an active hurricane evacuation warning. Both could be doing better. One side is past my ability to help.
Even Gary Marcus has shifted to claiming that the modern models are not "pure LLMs".
Modern models use tool calls, so this seems straightforwardly correct, at least in a certain technical sense.
The best kind of correct, perhaps, but that interpretation really takes the wind out of the sails of his conclusions. He wanted to "start easy: 5 digit multiplication problems". GPT4 had a 6.67% success rate, ha ha! But what do you think Gary Marcus's success rate would be? Modern LLMs might still need chain-of-thought to answer such a question reliably without an external tool, but I'd bet that Gary needs the same and I'd even bet he needs external pencil-and-paper to keep track of the intermediate steps. He says "The LLM never induces such an algorithm", but when I ask for a tool-less multiplication I get a reasonable digit-by-digit algorithm executed in the output text, which is pretty weird if "the LLM based system is generalizing by similarity, doing better on cases that are in or near the training set, never, ever getting to a complete, abstract, reliable representation of what multiplication is."
This guy still gets quoted as an "expert".
I don't think using tool calls reflects poorly on LLMs are products at all - if anything it enhances their value. This isn't an "LLMs suck" post.
However a lot of people view intelligence as a "unified" property. I've been pushing back on that idea on here for a while because I doubt that will be correct; my guess (and so far I've been proven correct; see LLMs getting worse at writing as they specialize for coding) is actually that there are benefits to intelligence specialization. This doesn't mean you cannot wrap those specialized compartments together into a unified process, of course. The human brain, for instance, is "unified" but it has specialized regions that seem to be optimized for specific tasks; same with your computer. Arguably an LLM using tool calls and the like is doing something similar.
I think it matters because 1. I'm petty and like being right, and 2. I think it's worth thinking clearly about these things.
For quite some time, frontier models have generally been Mixture of Experts models (or, possibly, something even more proprietary and advanced - I'm not an insider). I think this fits with your intuition.
Thanks for flagging that link! Yes, I do think that fits with my intuition.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I implore you not to sane-wash Gary Marcus. Modern models like Astra are more than capable of feats that he swore up and down couldn't be done by LLMs, even if they could be run as bare as possible, without tool calls.
The most symbolic component in OpenAI's release is the Lean checker, and its job is to grade the LLM's work.
Gary Marcus could be literally insane and it would still not be right to mock him for saying something that is true.
Do you know for sure that no tools were used? My understanding is that the reasoning traces released for the most recent models were summarized; however, we know that the agents that solved Napier-Stokes had tool access. I'm not sure I'd assume something different was done here, but maybe they said so somewhere.
I'd be fascinated to be corrected if the facts disagree, but just based on priors it would honestly be irresponsible to do something different here. LLM output is indispensable for problems that aren't well-posed or don't have a straightforward algorithm to find a solution, but inefficient and risky for problems that are and do.
When I was a first-year grad student, a TA gently pointed out at the bottom of half a page of my handwritten symbolic calculus simplifications that my doing that much by hand was (although correct in this case!) both risky and a waste of time, and that we all had Mathematica/Maple/etc licenses we were allowed and encouraged to use. I'd guess an LLM might prefer SymPy in the same circumstances, but the general principle is surprisingly just as true.
I'd be a little less surprised if the available tools other than Lean weren't at all useful for most of these problems, but surely they were useful for many of them. There are some really good open source tools out there these days (just looking around: SymPy has had a differential geometry module for a decade!) and OpenAI's models surely know about effectively all of the best ones.
Yeah I agree, I assume you'd want to give them the best tools, if they might be relevant at all.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Stock market can still be a bubble, even if the underlying tech is amazing. Research has cost an absurd amount of money, and it is unclear whether any of these companies will recoup their investments. Open-weight models are increasingly powerful and there is a lot of competition in the field. No one has a monopoly at the moment, and it doesn't look like one is forming either. Instead, everyone is trying to release the best possible models at the lowest possible cost.
Dang it, I spoke too soon! :) But fair enough; I was being a bit snarky. I do think a lot of the "AI is a bubble!" folks were arguing that AI would never be valuable, not that it's incredibly valuable but hard to capitalize on. Hopefully the former argument, at least, has been put to bed.
My theory is that a lot of the anti-AI crowd has simply been motivated by fear and ideology. They latch onto any argument that can be used to stop investment into these things, either because they distrust the tech bros, or because they view continued improvement as an existential threat to their livelihoods.
Since neither of those factors have changed and the existential threat has only worsened, I think the anti-AI crowd will only become louder in the coming months. If they can no longer credibly point to the abilities of the robots, they will just find a different angle. Water consumption, energy supply, or profitability are all arguments that are still being spread around.
You don't stop these people by being right. You stop them by convincing them they will benefit from the technology, or at least that it won't harm them or a cause they care about.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
If only it were so simple!
Alas almost all technological hype cycles are for technology that works and is transformative. Trains, canals, optic fiber, you name it.
There is most definitely a world where the massive investments in AI don't meet enough demand quickly enough and Anthropic goes the way of Cisco.
More options
Context Copy link
More options
Context Copy link
Be honest. How much of the above did you actually write? I’m reserving judgment. I’m just curious.
95% of it? I typed it out during my lunch break (I had lost my appetite), and then the only extent of LLM involvement was to dig up and then wrap in Markdown the relevant links.
Which are all real, anyway. I just didn't have the time to bother.
Hmm. Let me check closely:
Several throwaway statements had a sentence or two of context added. For example, "The binding constraint has shifted to elicitation and persistence".
If you want to be technical, it inserted the full title of Aaronson's post, or surfaced it in the first place. Initially, I just happened to come across it secondhand on Twitter and had pasted the quote in bare. Uh, now I see that it claims that I had strong opinions on the algorithmic floor of integer multiplication. I did not, beyond being vaguely aware that there was a better option than a naive n^2 based off an article I think I read on Quanta. I couldn't have told you off the top of my head that the previous SOTA was O(n log n).
So the LLM inserted something dumb that you have no personal knowledge of? I'd be embarrassed, myself, but you do you I guess.
Oh really? What's so dumb about it? When I checked X earlier, there were plenty of mathematicians expressing surprise at getting any lower than O (n log n). I would have been surprised too, if I had remembered the exact threshold.
https://www.themotte.org/post/3960/culture-war-roundup-for-the-week/485819?context=8#context
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
In my very scholarly and medical opinion, I think you're suffering from Human to AI writing transmorigification. I thought your post was also at the very least highly AI modified but it may well be that AI has addled your brain so much that now you just naturally write that way.
A curious specimen indeed.
Count, I always give your opinions as much weight as they're due. Just don't ask me how much right now.
More options
Context Copy link
More options
Context Copy link
Aw, dang, I was actually pretty impressed you had such specific CS knowledge!
Hey man, I read HN religiously ever day, and I once solved a LC medium in 2021.
Jokes aside, I do like maths and CS. The former all the more so when I'm not forced to for the sake of exams. And I try to learn what I can as the mood takes me. I even tried to learn Lambda calculus once, even though it provided literally no practical utility.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link