This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
A Big Day in the Culture War
Apologies for any incoherency or grammatical errors; today's events led me to down more than my usual share of booze.
Earlier today, big things happened in math. No, not Claude providing a sub quadratic 3SUM. OpenAI released hundreds of notable math results, in a GitHub repo.
Fun results:
Hilbert's tenth problem over (\mathbb{Q}) is undecidable
The quasi-Riemann hypothesis
The rational Hodge conjecture holds for every CM abelian variety
Integer multiplication can be done in sub-log-linear time. Oh, there's also a sub-log linear DFT.
Thoughts and observations:
Mathematicians are big mad. View the relevant subreddit on our progenitor site. Most there are probably, at best, adjuncts at community colleges desperately coping with the downward trajectory of already marginal careers, but it's fair to say that the writing is on the wall for mathematicians. There's probably a double digit number of grad students staring into a glass of whiskey tonight and thinking of hanging themselves.
My immediate question was about whether this closer to the current peak of performance, or just a lazy demonstration of OpenAI's power. So, I took one of the particularly interesting preprints to me (memory and precision in Gaussian models) and tested whether it's at the edge of capabilities or not. 30 minutes of back and forth with Astra (itself behind OAI's internal model) resulted in a significantly stronger result, on multiple dimensions. While I finished my first bottle, I spent a fair amount of time convincing myself of the result; I was convinced the strengthened results were plausible. Write this off as AI psychosis if you want, but try it yourself; I'm genuinely curious for what you get. (The back and forth, here, was entirely me saying "you can do it!", "keep at it, I believe in you", and "you've got this, finish it!")
Probably the most important line in OAI's announcement post is "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." This isn't a case of OAI spending millions for a marketing bump. My bet is a kind of Pareto distribution: most took minutes, not hours, with a long tail around Riemann-level results pulling up the average significantly.
Notably, ML related results are nearly entirely absent from this batch of proofs; maybe a half dozen touch on it, distantly. Some problem indices are skipped in overview.md. Conspiratorially, my inclination was to think they filtered them out for competitive advantage. I can't find any evidence of that in the GitHub repo or any of the preprints (equally plausible: deduping), so maybe they judiciously decided not to point their mathematical ballista at ML. ML also doesn't have a meaningful bank of rigorous conjectures, so given the conjecture sources, maybe they're not yet digging into ML math. But color me skeptical.
Does math matter? Is it something to advance civilization and technology, or an artistic pasttime for humans to create logical beauty? Likely both, today, but this is an almost nuclear detonation against the latter.
The only reason I'm not saying it's over is because I've said it before. It's been over for a while. And we've barely begun.
The good news, for mathematicians, is that their employability (in academia) did not hinge very strongly on tangible economic output. Nobody funds an algebraic geometer expecting a return on investment, which means nobody can defund one on the grounds that a datacenter in Texas now offers a better one.
The bad news? Everything else.
Just look at this fucking sweep. Just look at it. Hilbert's tenth over the rationals. Rational Hodge for CM abelian varieties. Integer multiplication below n log n, which I had mentally filed under "the floor, go home." A quasi-Riemann hypothesis thrown in like a free tote bag. Any one of these would have been the defining result of a human career, and they were released as a batch, in a GitHub repo, on a Tuesday, alongside hundreds of others.
And it wasn't brute force at absurd cost. OpenAI says the average result used roughly three hours of ChatGPT Pro-equivalent compute. I agree with your Pareto intuition, but the deeper point is that three hours is a price, and prices in this industry fall monotonically. Gwern spent 2020 arguing that neural nets would keep absorbing compute and sprouting new abilities, while noting the idea was so unpopular that it would only be accepted as a fait accompli. Well, here is the fait, and it is quite accompli. Next year the same theorem costs twenty minutes. The year after, it's a rounding error on someone's API bill.
Your own experiment deserves more attention than you gave it. As others have already done to good effect, your research methodology consisted of being a motivational poster ("you can do it!", "I believe in you") and you got a materially stronger result in half an hour. The binding constraint has shifted to elicitation and persistence.
Some might claim that human mathematicians are necessary to humanize and understand the proofs, or to come up with new frameworks and fields of research.
Maybe. I don't know. The strongest version of the argument is that mathematics is a conversation between humans about what humans find interesting, and a proof nobody understands is a tree falling in an empty forest. I'm sympathetic to that! I also notice that this position has been retreating to progressively smaller hills since the first computer-assisted proofs, and each hill gets defended with the same conviction as the last.
I do know that the current state of affairs is unlikely to last for more than six months anyway. Just get the models to come up with their own novel conjectures, or to attempt an overarching synthesis. I don't think turning abstruse Lean proofs into something human-readable will be particularly difficult, though human readability has long ceased to be a major concern for anyone. Note that OpenAI is shipping Lean formalizations for many of these proofs, so "trust me bro" is off the table; a type checker cares nothing for anyone's feelings. We're at the point where people are coping that the tsunami of new discoveries hinge on hither-to undiscovered bugs in Lean. You wish.
Every day, I feel the pace of progress accelerate, and then the practical ramifications show up in my life a few months later. I used to think of this as a lag. These days it feels more like the gap between seeing lightning and hearing thunder, and the gap keeps shrinking.
We've just had Scott Aaronson announce a new UT Austin course, CS395T: AI Alignment Theory, with Yudkowsky's AGI Ruin as the first assigned reading. To quote:
I'll give credit where it's due, but also: when the designated Reasonable Skeptic of the rationalist diaspora concedes the point in writing, the Overton window has relocated.
We've also just had an American company, Nolla Health, receive regulatory approval in Utah for AI to issue initial prescriptions. Yes, it's acne. Yes, the model can only choose from a short list of physician-approved topicals, and clinicians reportedly agreed with its recommendations in over 96% of cases. Everything starts as acne cream. Utah let Doctronic's AI handle prescription renewals back in January; nine months later, it's writing first-line scripts. You can extrapolate the line yourself. That's the harbinger of medicine's fall, as far as I'm concerned. I'm impressed we held the moat this long.
At least I've got a moat. The NHS is famously a slow ship to steer, and sclerotic at that. The future arrives unevenly distributed, and we can count on the service the last postcode on the delivery route. For once, I find that reassuring <3
Good career choice, past self_made_human. Turns out there are concrete benefits to taking theoretical but plausible concerns seriously, and preparing accordingly.
I see no reason why you couldn't replace the hype man with a cheap AI either. One LLM tries to find the proof. The other is instructed to hype it up. No human involved, other than in deciding what proof to look for next.
I consider that so obvious I didn't find it worth saying. Unless, for some stupid reason, the models learn that they're being prompted by print.ln instead of a human.
Given that reasoning effort is something that can be adjusted (usually by simply a privileged part of the prompt, rather than some fancy dial), it's moot.
More options
Context Copy link
More options
Context Copy link
Who? Aaronson? The "took his family to Disneyland and aborted the trip after having a meltdown over seeing people not wearing masks" guy? He's many things, but nobody is designating him any combination of "reasonable" and "skeptic".
(By the way, concurring with @2rafa's suspicion below, after you have acquired your reputation as an unapologetic meat proxy, it does feel extra suspicious when you make such catchy irresistible logit-attracting utterances whose substance seems bizarre when you actually think about it.)
Some people have forgotten that I'm capable of sarcasm.
I'm more than aware of Aaronson's quirks. In some regards, he's a pathological quokka, and in other cases, he's had a rather overblown threat response and reacted out of proportion. See his posts about his fear of anti-semitism for the latter, or the fact that he inspired the other Scott to write "Radicalizing the Romanceless."
In the Rationalist sphere, he's the one who's always painted him as the reasonable person, not prone to theoretical flights of fancy.
Anyway, unapologetic? Not at all. I just think it's a tradeoff I'm willing to make, even if it annoys you.
pushes up glasses
Well actually, he inspired Scott (PBUH) to write “Untitled”, which came out a year or so after “Radicalizing the Romanceless”. Both absolute bangers, though.
I am ashamed of my error, and will commit sudoku out of repentance.
(That's what I get for not running a throw away comment through an LLM for fact checks. I'm more likely to 'hallucinate' than they are.)
More options
Context Copy link
More options
Context Copy link
Have you heard about Poe's Law? Sarcasm doesn't work when it's equally plausible that you would say the same thing in earnest, and attaching a clichéd epithet like that without any deeper reasoning behind it is the sort of thing an LLM would do in earnest.
All's well. You're willing to annoy me with AI content, and I'm willing to annoy you by derailing your threads with speculations about AI content :)
I consider that a reasonable compromise for all involved. It's the people who think they have the right to dictate how I spend my free time that I take umbrage with. I otherwise aim to be honest if anyone explicitly asks.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Sure, AI may be embarrassing humanity's proudest intellectual achievements. But there's a bright side. We can finally stop arguing about how AI is an overhyped bubble that's about to burst! Haven't even seen the term "stochastic parrot" in a while.
Another plus: AI can now compose its own Disney villain song. I'm not exactly sure where "writing a catchy diss track against people who think you're not intelligent" lies on the spectrum of intelligence, but it's probably a little ways past the Turing Test.
We aren't arguing about it because there's no point. At this point, both sides have tired of trying to convince each other, each still thinks they are right, and each still thinks the other is willfully ignoring evidence that disproves their POV. What's there to argue about any more? Life's too short.
More options
Context Copy link
I'm sure we can do that, but maybe we shouldn't.
There's been some pretty not-great signs for the financial state of the industry lately: Microsoft looking to wean off of Anthropic, Harvey and Thompson Reuters moving to their own models, the force majeure notice by Oracle in New Mexico, the failure of SB Energy, Holtec, and Aggreko to IPO (not to mention OpenAI and Anthropic!), Firmus missing its rental payment, and generally increased skepticism on the part of investors.
I think it's been a huge problem for thinking clearly about the situation that the financial concern regarding the AI industry has been latched onto by the worst AI skeptics, who tend to flatly deny the capabilities of the models.
Meanwhile on the flip side, I suspect there may have been a parallel problem that the people who are the biggest boosters of the technology are the ones most likely to be suffering from mild AI psychosis from talking with them all the time.
Thus, the AI debate has been, somehow, between "AIs suck and there is a massive bubble" and "AIs are the best thing since sliced bread and nuh-uh," which excludes two entire quadrants of possibility from the conversation, "AIs suck and will be profitable" and "AIs are good and there is a bubble."
This state of discourse creates epistemic closure on the topic in a way that I think is preventing a lot of people from assessing where, exactly, AI is sitting.
More options
Context Copy link
There are still holdouts, but at this point I regard them with more psychiatric curiosity and pity than I do as entities to reason with.
As an analogy: people who ignore the insurance on their beach side property shooting up vs those ignoring an active hurricane evacuation warning. Both could be doing better. One side is past my ability to help.
Even Gary Marcus has shifted to claiming that the modern models are not "pure LLMs".
Modern models use tool calls, so this seems straightforwardly correct, at least in a certain technical sense.
The best kind of correct, perhaps, but that interpretation really takes the wind out of the sails of his conclusions. He wanted to "start easy: 5 digit multiplication problems". GPT4 had a 6.67% success rate, ha ha! But what do you think Gary Marcus's success rate would be? Modern LLMs might still need chain-of-thought to answer such a question reliably without an external tool, but I'd bet that Gary needs the same and I'd even bet he needs external pencil-and-paper to keep track of the intermediate steps. He says "The LLM never induces such an algorithm", but when I ask for a tool-less multiplication I get a reasonable digit-by-digit algorithm executed in the output text, which is pretty weird if "the LLM based system is generalizing by similarity, doing better on cases that are in or near the training set, never, ever getting to a complete, abstract, reliable representation of what multiplication is."
This guy still gets quoted as an "expert".
I don't think using tool calls reflects poorly on LLMs are products at all - if anything it enhances their value. This isn't an "LLMs suck" post.
However a lot of people view intelligence as a "unified" property. I've been pushing back on that idea on here for a while because I doubt that will be correct; my guess (and so far I've been proven correct; see LLMs getting worse at writing as they specialize for coding) is actually that there are benefits to intelligence specialization. This doesn't mean you cannot wrap those specialized compartments together into a unified process, of course. The human brain, for instance, is "unified" but it has specialized regions that seem to be optimized for specific tasks; same with your computer. Arguably an LLM using tool calls and the like is doing something similar.
I think it matters because 1. I'm petty and like being right, and 2. I think it's worth thinking clearly about these things.
More options
Context Copy link
More options
Context Copy link
I implore you not to sane-wash Gary Marcus. Modern models like Astra are more than capable of feats that he swore up and down couldn't be done by LLMs, even if they could be run as bare as possible, without tool calls.
The most symbolic component in OpenAI's release is the Lean checker, and its job is to grade the LLM's work.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Stock market can still be a bubble, even if the underlying tech is amazing. Research has cost an absurd amount of money, and it is unclear whether any of these companies will recoup their investments. Open-weight models are increasingly powerful and there is a lot of competition in the field. No one has a monopoly at the moment, and it doesn't look like one is forming either. Instead, everyone is trying to release the best possible models at the lowest possible cost.
Dang it, I spoke too soon! :) But fair enough; I was being a bit snarky. I do think a lot of the "AI is a bubble!" folks were arguing that AI would never be valuable, not that it's incredibly valuable but hard to capitalize on. Hopefully the former argument, at least, has been put to bed.
More options
Context Copy link
More options
Context Copy link
If only it were so simple!
Alas almost all technological hype cycles are for technology that works and is transformative. Trains, canals, optic fiber, you name it.
There is most definitely a world where the massive investments in AI don't meet enough demand quickly enough and Anthropic goes the way of Cisco.
More options
Context Copy link
More options
Context Copy link
Be honest. How much of the above did you actually write? I’m reserving judgment. I’m just curious.
95% of it? I typed it out during my lunch break (I had lost my appetite), and then the only extent of LLM involvement was to dig up and then wrap in Markdown the relevant links.
Which are all real, anyway. I just didn't have the time to bother.
Hmm. Let me check closely:
Several throwaway statements had a sentence or two of context added. For example, "The binding constraint has shifted to elicitation and persistence".
If you want to be technical, it inserted the full title of Aaronson's post, or surfaced it in the first place. Initially, I just happened to come across it secondhand on Twitter and had pasted the quote in bare. Uh, now I see that it claims that I had strong opinions on the algorithmic floor of integer multiplication. I did not, beyond being vaguely aware that there was a better option than a naive n^2 based off an article I think I read on Quanta. I couldn't have told you off the top of my head that the previous SOTA was O(n log n).
Aw, dang, I was actually pretty impressed you had such specific CS knowledge!
Hey man, I read HN religiously ever day, and I once solved a LC medium in 2021.
Jokes aside, I do like maths and CS. The former all the more so when I'm not forced to for the sake of exams. And I try to learn what I can as the mood takes me. I even tried to learn Lambda calculus once, even though it provided literally no practical utility.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link