site banner

Culture War Roundup for the week of October 5, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

Even Gary Marcus has shifted to claiming that the modern models are not "pure LLMs".

Modern models use tool calls, so this seems straightforwardly correct, at least in a certain technical sense.

The best kind of correct, perhaps, but that interpretation really takes the wind out of the sails of his conclusions. He wanted to "start easy: 5 digit multiplication problems". GPT4 had a 6.67% success rate, ha ha! But what do you think Gary Marcus's success rate would be? Modern LLMs might still need chain-of-thought to answer such a question reliably without an external tool, but I'd bet that Gary needs the same and I'd even bet he needs external pencil-and-paper to keep track of the intermediate steps. He says "The LLM never induces such an algorithm", but when I ask for a tool-less multiplication I get a reasonable digit-by-digit algorithm executed in the output text, which is pretty weird if "the LLM based system is generalizing by similarity, doing better on cases that are in or near the training set, never, ever getting to a complete, abstract, reliable representation of what multiplication is."

This guy still gets quoted as an "expert".

I don't think using tool calls reflects poorly on LLMs are products at all - if anything it enhances their value. This isn't an "LLMs suck" post.

However a lot of people view intelligence as a "unified" property. I've been pushing back on that idea on here for a while because I doubt that will be correct; my guess (and so far I've been proven correct; see LLMs getting worse at writing as they specialize for coding) is actually that there are benefits to intelligence specialization. This doesn't mean you cannot wrap those specialized compartments together into a unified process, of course. The human brain, for instance, is "unified" but it has specialized regions that seem to be optimized for specific tasks; same with your computer. Arguably an LLM using tool calls and the like is doing something similar.

I think it matters because 1. I'm petty and like being right, and 2. I think it's worth thinking clearly about these things.

For quite some time, frontier models have generally been Mixture of Experts models (or, possibly, something even more proprietary and advanced - I'm not an insider). I think this fits with your intuition.

Thanks for flagging that link! Yes, I do think that fits with my intuition.

I implore you not to sane-wash Gary Marcus. Modern models like Astra are more than capable of feats that he swore up and down couldn't be done by LLMs, even if they could be run as bare as possible, without tool calls.

The most symbolic component in OpenAI's release is the Lean checker, and its job is to grade the LLM's work.

I implore you not to sane-wash Gary Marcus.

Gary Marcus could be literally insane and it would still not be right to mock him for saying something that is true.

The most symbolic component in OpenAI's release is the Lean checker, and its job is to grade the LLM's work.

Do you know for sure that no tools were used? My understanding is that the reasoning traces released for the most recent models were summarized; however, we know that the agents that solved Napier-Stokes had tool access. I'm not sure I'd assume something different was done here, but maybe they said so somewhere.

I'm not sure I'd assume something different was done here

I'd be fascinated to be corrected if the facts disagree, but just based on priors it would honestly be irresponsible to do something different here. LLM output is indispensable for problems that aren't well-posed or don't have a straightforward algorithm to find a solution, but inefficient and risky for problems that are and do.

When I was a first-year grad student, a TA gently pointed out at the bottom of half a page of my handwritten symbolic calculus simplifications that my doing that much by hand was (although correct in this case!) both risky and a waste of time, and that we all had Mathematica/Maple/etc licenses we were allowed and encouraged to use. I'd guess an LLM might prefer SymPy in the same circumstances, but the general principle is surprisingly just as true.

I'd be a little less surprised if the available tools other than Lean weren't at all useful for most of these problems, but surely they were useful for many of them. There are some really good open source tools out there these days (just looking around: SymPy has had a differential geometry module for a decade!) and OpenAI's models surely know about effectively all of the best ones.

Yeah I agree, I assume you'd want to give them the best tools, if they might be relevant at all.