This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
A Big Day in the Culture War
Apologies for any incoherency or grammatical errors; today's events led me to down more than my usual share of booze.
Earlier today, big things happened in math. No, not Claude providing a sub quadratic 3SUM. OpenAI released hundreds of notable math results, in a GitHub repo.
Fun results:
Hilbert's tenth problem over (\mathbb{Q}) is undecidable
The quasi-Riemann hypothesis
The rational Hodge conjecture holds for every CM abelian variety
Integer multiplication can be done in sub-log-linear time. Oh, there's also a sub-log linear DFT.
Thoughts and observations:
Mathematicians are big mad. View the relevant subreddit on our progenitor site. Most there are probably, at best, adjuncts at community colleges desperately coping with the downward trajectory of already marginal careers, but it's fair to say that the writing is on the wall for mathematicians. There's probably a double digit number of grad students staring into a glass of whiskey tonight and thinking of hanging themselves.
My immediate question was about whether this closer to the current peak of performance, or just a lazy demonstration of OpenAI's power. So, I took one of the particularly interesting preprints to me (memory and precision in Gaussian models) and tested whether it's at the edge of capabilities or not. 30 minutes of back and forth with Astra (itself behind OAI's internal model) resulted in a significantly stronger result, on multiple dimensions. While I finished my first bottle, I spent a fair amount of time convincing myself of the result; I was convinced the strengthened results were plausible. Write this off as AI psychosis if you want, but try it yourself; I'm genuinely curious for what you get. (The back and forth, here, was entirely me saying "you can do it!", "keep at it, I believe in you", and "you've got this, finish it!")
Probably the most important line in OAI's announcement post is "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." This isn't a case of OAI spending millions for a marketing bump. My bet is a kind of Pareto distribution: most took minutes, not hours, with a long tail around Riemann-level results pulling up the average significantly.
Notably, ML related results are nearly entirely absent from this batch of proofs; maybe a half dozen touch on it, distantly. Some problem indices are skipped in overview.md. Conspiratorially, my inclination was to think they filtered them out for competitive advantage. I can't find any evidence of that in the GitHub repo or any of the preprints (equally plausible: deduping), so maybe they judiciously decided not to point their mathematical ballista at ML. ML also doesn't have a meaningful bank of rigorous conjectures, so given the conjecture sources, maybe they're not yet digging into ML math. But color me skeptical.
Does math matter? Is it something to advance civilization and technology, or an artistic pasttime for humans to create logical beauty? Likely both, today, but this is an almost nuclear detonation against the latter.
They have absolutely pointed the ballista at machine learning. What you're seeing is them briefly pointing the ballista anywhere else to produce something to share, because there is no way they are sharing their trade secrets and competitive advantage publicly.
I'm quite sure they have. These math results are likely almost an afterthought; I wouldn't be surprised if they've already put 10x of the compute they used for these math results into researching ML.
10x? More like 100-10,000x I'd wager. See the piddling $4 million estimate above. It makes little difference even if it cost them $40 million instead. >>$100m is the ballpark of where it might make sense for them to not try, because the reputational aura and PR is worth it.
What did people think RSI meant? Vibes? Papers?
"sub-incremental counterexamples to previous conjectures with no practical applications"?
No doubt. I'm sure you could describe the majority of human-derived mathematical findings in the exact same words. I'm sure figuring out optimal square-packing has made billions of dollars for someone.
Nobody is making any money on a 1-10^xx exponent on integer multiplication's big O -- the fact that you think this shows that you are relying on an unreliable interlocuter to form your opinions.
I don't dispute that some of the other results may have some useful application, but the ones of this form are only interesting in the sense that humans tend to expect round numbers, and have been mostly correct in this assumption to date. (I do share this expectation and personally think it's more likely that there's some kind of mistake in these proofs -- but if I'm wrong, that's interesting!)
Your bot can probably explain this to you if you ask, but briefly: Big-O notation typically disregards the portion of algorithmic complexity that's constant doesn't and depend on the size of the dataset; ie. whether a loop takes a millisecond or a minute to execute once is not considered. Something that takes a minute per cycle but is O(1) may be better than something that does it in a millisecond but is O(n^2) -- it depends on how many "n" you are interested in though.
TBH I haven't looked into the actual algorithms involved here, but given the length of the proof I'd expect the constant term (call it overhead) to be quite high compared to existing methods -- considering the tiny difference in the exponents, you would need n to be literally approaching infinity for the proposed algo to make any difference at all. In the case of the integer math, this would mean that multiplying integers of ~infinite length could be somewhat faster -- but we can't even test it, because computers don't have infinite memory.
Do I think this? Nobody told me. It's not true. I've never said it, I've never thought it, and you can go look.
I'd prefer you stick to what I've actually said. Or if that's what you thought I implied, I ask that you at least acknowledge when I clarify otherwise.
I also know how Big O works, for the record. I took CS courses in my free time once. I am a nerd.
You just said it though! (well, maybe your LLM did)
It doesn't really seem like you do; an optimal square packing algorithm that is better only in the case of infinite squares is not making anybody any money either.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I'm sure they have pointed it at ML, but ML isn't exactly the same as formal math. I'm probably describing this poorly but ML is very experimental/empirical, not theoretical like formal math. The math heavy side of ML would be finding better optimizers, regularizers, improving backpropagation or finding a better method than SGD. The question becomes how much of an edge do these provide? Because better algorithmic components still have to deal less-better data or compute components and how well all those now perform better is not theoretically provable.
What most of these breakthroughs tell me is that supervised learning is highly effective at formal math, and the benefit of something like Lean + Solvers to allow for computational checking of math proofs has been a fundamental driver. Whether or not strategies learned on this frontier are applicable to broader fields that are less formalizable is unknown to me.
Much of the empirical part of ML, though, is also something the LLM can do autonomously. The agent can write up some pytorch or jax to train a model on an existing dataset, observe whatever quantities you want, repeat. It does have a longer feedback loop and more compute requirements than pure math, but quite automatable.
The pure math part of ML is unfortunately quite weak right now and hasn't played a huge role in the current boom; pure math research could provide a much stronger basis for understanding learning and creating new approaches, and I think it will, but that's speculative.
We'll find out soon.
The longer feedback loop is unfortunately the expensive part, as we talked about last time. The real game changer would be to reduce the compute cost and training time for LLMs and/or figure out a more efficient, non-quadratic self-attention mechanism. However, at some level the cost is the moat frontier labs have, so there is almost perverse-conflicting incentives not to improve it but also requiring it for RSI to happen.
Agreed.
I'm sure the mathematicians left out in the cold by the big data revolution of ML will rejoice that there was indeed a mathematical approach to LLMs over the pesky engineering approaches that were developed by the peons from MIT.
Damn this reminds me of the ML class I took back in my maths degree. We spent the whole term proving theorems about perceptrons rather than being taught any sort of ML that was being used in the world.
Rite of Passage honestly. I had one similar about the fundamental theory of AI, lots of proofs. While fascinating I actually wanted a class of how to train neural nets back when Torch was still in Lua and how to use Cuda.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link