This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
I'm quite sure they have. These math results are likely almost an afterthought; I wouldn't be surprised if they've already put 10x of the compute they used for these math results into researching ML.
10x? More like 100-10,000x I'd wager. See the piddling $4 million estimate above. It makes little difference even if it cost them $40 million instead. >>$100m is the ballpark of where it might make sense for them to not try, because the reputational aura and PR is worth it.
What did people think RSI meant? Vibes? Papers?
"sub-incremental counterexamples to previous conjectures with no practical applications"?
No doubt. I'm sure you could describe the majority of human-derived mathematical findings in the exact same words. I'm sure figuring out optimal square-packing has made billions of dollars for someone.
Nobody is making any money on a 1-10^xx exponent on integer multiplication's big O -- the fact that you think this shows that you are relying on an unreliable interlocuter to form your opinions.
I don't dispute that some of the other results may have some useful application, but the ones of this form are only interesting in the sense that humans tend to expect round numbers, and have been mostly correct in this assumption to date. (I do share this expectation and personally think it's more likely that there's some kind of mistake in these proofs -- but if I'm wrong, that's interesting!)
Your bot can probably explain this to you if you ask, but briefly: Big-O notation typically disregards the portion of algorithmic complexity that's constant doesn't and depend on the size of the dataset; ie. whether a loop takes a millisecond or a minute to execute once is not considered. Something that takes a minute per cycle but is O(1) may be better than something that does it in a millisecond but is O(n^2) -- it depends on how many "n" you are interested in though.
TBH I haven't looked into the actual algorithms involved here, but given the length of the proof I'd expect the constant term (call it overhead) to be quite high compared to existing methods -- considering the tiny difference in the exponents, you would need n to be literally approaching infinity for the proposed algo to make any difference at all. In the case of the integer math, this would mean that multiplying integers of ~infinite length could be somewhat faster -- but we can't even test it, because computers don't have infinite memory.
Per wikipedia, the existing O(n log n) algorithm for integer multiplication that the AI improved on is already considered "galactic" - i.e. only optimal for impossibly large data sets. The practical algorithm for the largest numbers we can realistically multiply is a FFT-based approach which is O(n log n log log n)
To my non-expert eye, using the new slightly-faster DFT in the AI papers in the bog standard FFT-based multiplication algorithm should give you a faster multiplier than the new multiplier the AI found explicitly.
More options
Context Copy link
Do I think this? Nobody told me. It's not true. I've never said it, I've never thought it, and you can go look.
I'd prefer you stick to what I've actually said. Or if that's what you thought I implied, I ask that you at least acknowledge when I clarify otherwise.
I also know how Big O works, for the record. I took CS courses in my free time once. I am a nerd.
You just said it though! (well, maybe your LLM did)
It doesn't really seem like you do; an optimal square packing algorithm that is better only in the case of infinite squares is not making anybody any money either.
You have missed clear sarcasm. Now you know better.
I don't think I've agreed with a word out of @BurdensomeCount 's mouth in my life -- but you truly are a curious case.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I'm sure they have pointed it at ML, but ML isn't exactly the same as formal math. I'm probably describing this poorly but ML is very experimental/empirical, not theoretical like formal math. The math heavy side of ML would be finding better optimizers, regularizers, improving backpropagation or finding a better method than SGD. The question becomes how much of an edge do these provide? Because better algorithmic components still have to deal less-better data or compute components and how well all those now perform better is not theoretically provable.
What most of these breakthroughs tell me is that supervised learning is highly effective at formal math, and the benefit of something like Lean + Solvers to allow for computational checking of math proofs has been a fundamental driver. Whether or not strategies learned on this frontier are applicable to broader fields that are less formalizable is unknown to me.
Much of the empirical part of ML, though, is also something the LLM can do autonomously. The agent can write up some pytorch or jax to train a model on an existing dataset, observe whatever quantities you want, repeat. It does have a longer feedback loop and more compute requirements than pure math, but quite automatable.
The pure math part of ML is unfortunately quite weak right now and hasn't played a huge role in the current boom; pure math research could provide a much stronger basis for understanding learning and creating new approaches, and I think it will, but that's speculative.
We'll find out soon.
The longer feedback loop is unfortunately the expensive part, as we talked about last time. The real game changer would be to reduce the compute cost and training time for LLMs and/or figure out a more efficient, non-quadratic self-attention mechanism. However, at some level the cost is the moat frontier labs have, so there is almost perverse-conflicting incentives not to improve it but also requiring it for RSI to happen.
Agreed.
I'm sure the mathematicians left out in the cold by the big data revolution of ML will rejoice that there was indeed a mathematical approach to LLMs over the pesky engineering approaches that were developed by the peons from MIT.
Damn this reminds me of the ML class I took back in my maths degree. We spent the whole term proving theorems about perceptrons rather than being taught any sort of ML that was being used in the world.
Rite of Passage honestly. I had one similar about the fundamental theory of AI, lots of proofs. While fascinating I actually wanted a class of how to train neural nets back when Torch was still in Lua and how to use Cuda.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link