site banner

Culture War Roundup for the week of October 5, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

A Big Day in the Culture War

Apologies for any incoherency or grammatical errors; today's events led me to down more than my usual share of booze.

Earlier today, big things happened in math. No, not Claude providing a sub quadratic 3SUM. OpenAI released hundreds of notable math results, in a GitHub repo.

Fun results:

  1. Hilbert's tenth problem over (\mathbb{Q}) is undecidable

  2. The quasi-Riemann hypothesis

  3. The rational Hodge conjecture holds for every CM abelian variety

  4. Integer multiplication can be done in sub-log-linear time. Oh, there's also a sub-log linear DFT.

Thoughts and observations:

  1. Mathematicians are big mad. View the relevant subreddit on our progenitor site. Most there are probably, at best, adjuncts at community colleges desperately coping with the downward trajectory of already marginal careers, but it's fair to say that the writing is on the wall for mathematicians. There's probably a double digit number of grad students staring into a glass of whiskey tonight and thinking of hanging themselves.

  2. My immediate question was about whether this closer to the current peak of performance, or just a lazy demonstration of OpenAI's power. So, I took one of the particularly interesting preprints to me (memory and precision in Gaussian models) and tested whether it's at the edge of capabilities or not. 30 minutes of back and forth with Astra (itself behind OAI's internal model) resulted in a significantly stronger result, on multiple dimensions. While I finished my first bottle, I spent a fair amount of time convincing myself of the result; I was convinced the strengthened results were plausible. Write this off as AI psychosis if you want, but try it yourself; I'm genuinely curious for what you get. (The back and forth, here, was entirely me saying "you can do it!", "keep at it, I believe in you", and "you've got this, finish it!")

  3. Probably the most important line in OAI's announcement post is "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." This isn't a case of OAI spending millions for a marketing bump. My bet is a kind of Pareto distribution: most took minutes, not hours, with a long tail around Riemann-level results pulling up the average significantly.

  4. Notably, ML related results are nearly entirely absent from this batch of proofs; maybe a half dozen touch on it, distantly. Some problem indices are skipped in overview.md. Conspiratorially, my inclination was to think they filtered them out for competitive advantage. I can't find any evidence of that in the GitHub repo or any of the preprints (equally plausible: deduping), so maybe they judiciously decided not to point their mathematical ballista at ML. ML also doesn't have a meaningful bank of rigorous conjectures, so given the conjecture sources, maybe they're not yet digging into ML math. But color me skeptical.

  5. Does math matter? Is it something to advance civilization and technology, or an artistic pasttime for humans to create logical beauty? Likely both, today, but this is an almost nuclear detonation against the latter.

My understanding impression of the navier-stokes "result" is that it is a mound of Lean and a paper no person can parse. Lean kernels say proof is valid, but Lean kernels may be buggy. The problem itself is so convoluted there is no agreement if even the formulation of the problem itself is correct, so Lean not being buggy is not enough either. Sounds unreliable, I don't care and presume the result is wrong.

In near-ish future, human understanding not needed, I'll take whatever pure math result AI publishes on faith alone. For now, show me something genuinely brilliant that any competent mathematican can verify in under a week (any of that in this pile? probably, but let's see in a month or 3), or a practical, usable result.

Break some encryption scheme, show a big advancement in ML, right, you totally can (true) but won't release. Any big applied math result, is that permitted, can american AI corpos into public science beyond IPO pump curios?

The problem itself is so convoluted there is no agreement if even the formulation of the problem itself is correct

Who doubts this? The formulation has been part of mathlib for some time already AFAIK.

Dumb of me to phrase it that way, entire sentence is confused. I meant to say, correct in the sense of formulating the core problem the big deal is about, not just a problem; enough, to count towards the core problem - no bearing on correct. Who doubts, I'd pretend to know if I listed anything concrete, it only is an impression from some scattered reading.

My heuristic to such results is that a counterexample is easier to verify then a positive proof. With Navier-Stokes conjecture in particular, the fact that another team was close to publishing a (different) counterexample makes me almost certain that Navier-Stokes conjecture doesn't hold. So if I were working on something that relied on the conjecture to be true in all cases, I would certainly redirect my efforts. (Or, as is quite common in research of pure mathematics, restrict my results to the conditions where the N-S conjecture holds true.)

Can you explain to a guy that's not deep into the math trenches what it would mean for that conjecture to be false in concrete terms irl? Or is that part of those abstractions on abstractions the mathematicians sometimes get themselves into?

In concrete terms the direct meaning of a blowup in the incompressible Navier-Stokes equations is nothing - incompressible fluids do not exist. For that matter, fluids do not exist - it's all atoms, and we pretend that atoms blend into a perfect continuum because it's a lot easier to approximate a problem using a million finite elements rather than a septillion atoms. Perhaps most fluid flow simulations are performed using incompressible Navier-Stokes anyway, though. If some of the fluid gets significantly warmer than the rest we'll partially take compressibility into account by adding a Boussinesq buoyancy term, but otherwise we generally pretend a fluid is incompressible until we're dealing with speeds near the speed of sound (in that fluid - e.g. "Mach 1" might mean 767 mph in standard air, but Mach 1 is 4 times faster in water), and we pretend flowing matter is a fluid until we're dealing with such tiny length scales or rarified gases that you have to worry about atoms slipping right past each other.

IMHO the rest of what's interesting is "abstractions on abstractions", especially for now, but ones that might lead to practical consequences in the future.

Half of what's interesting in the most practical sense here is: we want to control the error in our fluid flow simulations, and this might be another step toward figuring out how to do so more rigorously.

That "million finite elements" (in practice more likely finite volumes; unimportant distinction here) doesn't give you an exact solution to your flow problem, just an approximate solution, and that's fine because it's engineers who need the approximate solution and engineers use safety factors and the safety factors mean that if we give them a solution that's 10% off they're still fine. I once found a bug I'd written that made my code 0.0001% off, which snowballed through other equations to make a ton of engineering solutions 0.1% off. Coming from academia I was mortified, but the engineers hadn't noticed anything was wrong until I fixed the bug and thereby tripped some regression tests against past "gold" solutions. They hadn't yet bothered setting up problems where they could compare simulation solutions to exact solutions to verify the error magnitudes. Their idea of responsible testing was just validating against experiments, and the experimental measurement error was well over 0.1% so the validations said "this matches experiments well enough", and thankfully we never were simulating anything where simulation and reality diverged more than they did in experiments.

The trouble is that, even with an approximate solution, we always like to know how approximate it is, and if we always had an experiment to validate against then we wouldn't bother running a simulation. Grossly oversimplifying: with the simplest partial differential equations, we can prove results that look like "if the boundary conditions and 'forcing function' are of magnitude ‖f‖, then the solution u must be of magnitude ‖u‖≤c‖f‖, for a constant c that depends on the boundary". We can't actually find the solution u exactly, but we can approximate it, and we can turn those results about the exact solution into results about the error e=u-uₕ on an approximate solution uₕ, results that look something like ‖e‖≤mc‖f‖, for a constant m that depends on what type of approximation we use and how expensive we make it. If we want to guarantee a small error, for a given c and ‖f‖ we can pick an approximation with m small enough to get it.

We couldn't do this with Navier-Stokes in the most general cases, and now it's clear why: this does not work if c is infinity. Techniques that use the fact that the solution can't grow too badly to prove that our error can't grow too badly do not work if the solution can blow up. Often we can still prove that we're dealing with a "nice" case where the fluid viscosity can "overpower" a weak forcing enough to prevent a blowup, but now we know that "often" isn't "always"; we weren't just missing some clever trick to get there. This counterexample might stop people from wandering down a ton of blind alleys looking for one, or by showing us what we need to avoid it might point the way to expanding our criteria for "nice".

And the other half of what might become interesting down the road here: We don't understand turbulence nearly well enough, and yet most of the liquid flow and essentially all of the gas flow problems we want to solve are turbulent.

If we have a turbulent flow, that grid of "a million finite elements" becomes laughably small for a direct numerical simulation of it, because smaller and smaller turbulent eddies quickly become so small that we can't represent them accurately, while still being powerful enough to greatly affect the flow on the larger "macroscale" we care about. Instead we need to use, say, "35 trillion grid points", if we want a solution to just the equations of fluid flow that resembles the actual fluid flow. We can't afford to make every engineering problem millions of times more expensive to solve. So we've come up with "turbulence models", additional equations that try to approximate the statistically-averaged-out effects of turbulence on the macroscale flow, rather than precisely tracking every microscale vortex from its creation until viscosity damps it out. Now our equations are only 10% more expensive to solve (with the cheapest models) or 50% or 200% or whatever as we switch to better and better models, and we hope the results are good enough. It really is a mix of "hope" and "this matches experiments well enough" (for a much looser definition of "well enough" this time), because we don't have any really good turbulence models, not in the same "so close we never notice a difference" or "this is an inevitable consequence of the laws of physics with a few simplifications" senses in which incompressible-Navier-Stokes is a good model for viscous macroscopic fluid flow. Heisenberg supposedly said "When I meet God, I am going to ask him two questions: why relativity? And why turbulence? I really believe he will have an answer for the first." - apocryphal quote from nearly a century ago, but the turbulence experts I've talked to during the 21st century didn't disagree. But now we've got some fairly unexpected Navier-Stokes results, specifically dealing with the ways in which viscosity can or can't damp down a microscale vortex? This might go nowhere, or it might end up leading someone to a much better turbulence model down the road. I haven't looked at this stuff in a decade, but at a brief glance right now it looks like the cutting-edge in turbulence modeling is split between "tweak some 1990-era models to make them a bit better" and "if we enforce a few laws of physics on top of a machine-learning black-box it sort of works". Breaking out of those ruts wouldn't have a flashy Millennium Prize in the headlines but it could be a huge breakthrough.

Thanks, I enjoyed reading that.

That was super-informative! I just gained more useful insight from 2 minutes of reading your comment then I did from an hour of reading other explainers of the Navier-Stokes result. Many thanks!

A result which had two others downstream from it has been withdrawn by OpenAI. I think this is the first time a math result a frontier lab published was shown to be incorrect.

Not sure why @gradation is blocking me, but for the benefit of others, anyway, I think there's another astonishing bit here. In the version (hash 3014888) I'm reading: "This brings the total percentage of top-line results formalized to 300 / 719 = ~42%."

I'm generally the first to rave about how over the last 3 years AI has gone from "I'm lucky if I can talk it out of its hallucinated math mistakes" to "a mid-tier model catches me in as many math mistakes as vice-versa" ... but I still wouldn't trust anything long and AI-generated to be correct without some kind of verification! I would not have expected OpenAI's latest internal model to be so bad at autoformalization that they wouldn't just turn everything into Lean as a matter of course, nor for it to be so good at informal mathematics that they could even semi-safely publish an AI-generated proof without formally verifying it first. What exactly is happening there?

It'll be very interesting to see how much of the remaining 419 informal proofs get retracted like these 3, or at least revised for correctness like they did with another half dozen, versus simply being belatedly formalized like they did with an additional half dozen.

Ya, I think it was strategic to start this whole math run with that simple, easily-verifiable counterexample, because what we've received since then is not of that character. What we're getting now is mountains of slop that nobody has any hope of parsing in a thousand years, and that people can't even check in Lean because the "proofs" are so big they can't even be checked on consumer hardware.

But hey, if you pay into the totally-not-Ponzi-racket and buy industrial hardware at astronomically-inflated prices, you too can have the privilege to actually load the proofs in Lean!

and that people can't even check in Lean because the "proofs" are so big they can't even be checked on consumer hardware.

Did this retracted paper have a Lean formalization (IIUC not every paper in the repo does)? How much would it cost to run Lean on AWS to verify one of these proofs?

the "proofs" are so big they can't even be checked on consumer hardware

Fermat's Last Theorem verification squeezes into 256GB, which definitely isn't "consumer hardware", but it is on the workstation my job bought me 5 years ago; this is "prosumer" tier not "industrial". It's possible that OpenAI has come up with even trickier-to-verify formalizations, but I doubt it. Anthropic's FLT proof is 13.5M lines of Lean. All the recent OpenAI formalizations put together add up to 26M lines, but that's for 400 proofs. There's a lot of "library code" shared between proofs, so it's not as straightforward as 26M/400, but you're probably still going to be able to check at least some of them. This blogger/streamer tossed 3 of them at a tiny 7GB RAM cap and found that two hit the cap but one of the 3 topped out under 3GB.

How much would it cost to run Lean on AWS to verify one of these proofs?

For FLT it looks like you'd need a 16xlarge instance for 30-ish hours, so a bit over $100?

This particular one, no.

The concern isn't that Lean won't confirm the proofs. It's unlikely major AI vendors would brazenly lie like that. Rather, it's that you need to be able to use Lean like an IDE to poke around at the proof to investigate whether it even makes sense, whether they proved the thing they claimed to (a bug in the problem statement), or whether they found another bug in Lean itself.

As for running on some cloud host, my napkin math + Googling shows on the order of $1000/mo. But the thing is it's not just the cost: nobody that uses Lean knows anything about cloud hosting or how to set Lean up that way, because that's not the way anyone ever uses Lean. It would be like asking asking a Pokemon streamer to play with their GameBoy emulator on AWS instead of on their local PC. It's possible, yes, but the skillset to do so is completely orthogonal to what they usually do and they'd be lost attempting it. It's hard enough to even get mathematicans to use Lean in the first place, much less learn the basics of Linux sysadmin work.

As for running on some cloud host, my napkin math + Googling shows on the order of $1000/mo.

So $1.40 an hour? Doesn't seem that bad.

But the thing is it's not just the cost: nobody that uses Lean knows anything about cloud hosting or how to set Lean up that way, because that's not the way anyone ever uses Lean.

It's 2026, grandpa. There's no more excuses for not knowing how to open a PDF. You can just ask an LLM.

"Just pay into the racket bro"

Yes, we get it. Everyone has to throw all their money at the golem to be allowed to participate in society. Some of us aren’t buying it, in both senses.

You can use a free model. Or pay a few dollars (is that really all your money?) and use a better model.

I don't understand your perspective. Either LLMs are useless, in which case, what are we even talking about in this thread, or they aren't, in which case putting your head in the sand is not an effective long term strategy. "LLMs are eating my field and they're a racket" does not seem like a coherent position.

Doomer responses to this is just so sad. I'm an artist and designer and with every improvement of ai art/creativity I welcome it as augmenting possibilities of creativity rather than thinking my own creativity is worthless. Math people should be so happy they have new knowledge of math opening up to them and excited about the new answers rather than sad that a computer did it better than them. Surely there are now tons of new applications of knowledge and information from the new findings they can work from, I don't understand why they'd be upset about it. I'm sure it's very much like in art where most of the good artists are curious about the new AI tools while the bad ones worry.

Centaurs (humans using AI tools) don't last forever. They are a transitional stage until the AI gets good enough to replace the humans entirely.

And maybe that would be fine if we had figured out a way to compensate the losers with money and sexual access, but here in the real world, going from being an artist to being unemployed because now suddenly everyone has an artist in their laptop is a huge blow. It's not just envy and greed that other people now have access to your talents; it has real material consequences.

Most artists are artists because they love art, not because of the money (it's a notoriously badly-paid field). Telling them, "nope, sorry, you can't even keep earning a marginal living by drawing now, you will have to take some other job flipping burgers or sweeping floors and maybe draw in your spare time if you want but keep in mind you will never be as good as the machine and nobody will ever look at your drawings because everyone can have infinite high-quality drawings of whatever their niche interest is created on demand" is soul-crushing.

until the AI gets good enough to replace the humans entirely

Even with rapid AI advances, I predict this is a long time away. Sometime in the 1900s, amidst rapid technological development and automation (e.g. TVs, microwaves, washing machines), we thought the world would be like the Jetsons by 2000. AI’s creativity has actually regressed since around GPT 3; the newest models can produce images that look better and follow more instructions, but unprompted details are no more interesting.

Even the fine details like strokes are worse than an (albeit good) artist, so those artists still stand out from the centaurs, just less so. And those who want to up their game will still go to art school.

going from being an artist to being unemployed because now suddenly everyone has an artist in their laptop is a huge blow…” nope, sorry, you can't even keep earning a marginal living by drawing now, you will have to take some other job flipping burgers or sweeping floors and maybe draw in your spare time if you want…”

How many artists today make a living wage? If anything, AI improving productivity and automating boring tasks may give them more time for art.

I would argue that a large part of art is connecting with another human being through a shared experience. I think that's why so many artists are up in arms about making art so "easy," because it drops the human connection in many ways. The only human in the loop is prompting.

My view is that things will adapt. The human connection is still needed and there will always be art done the "old fashioned way." People still drag around large-format cameras to take pictures or simply shoot on film. My friend does painting and is trying to make a business of it.

At the same time, things have gotten easier. You don't have to be a "good" photographer to take a good photo with your phone's camera -- because there's already so much AI in that loop making things better. Same with so much other human-made art (and beyond). It's so much easier to get to the joy of creation. With photography, you can be creative with your phone and don't need the traditional thousands of dollars of gear to start creating.

The removal of those gates, both financial and the gate of years of practice, is what allows for that creation. I think we should celebrate that, not regulate it. Why does not "paying your dues" make your contribution invalid?

If your art, regardless of how it's made, helps you connect to others, I say that's a good thing.

I have a lot of other ideas in my head I'm trying to get out and make into a coherent argument, but I'm having trouble. For instance, we've had computer spell-checkers for decades now and no one blinks at that. We've had simple grammar checkers for a while as well, predating the AI everywhere thing we're seeing now, and no one really cares about that. It's only a matter of time before AI does more and people don't care. The only thing that really frustrates me is when people give up their agency to AI and stop thinking for themselves. AI is an amplifier of a human's ideas and abilities. But humans are still important.

A good way to think about artists, writers and the sort is to recognize they're performers who use their specialized mediums as a basis for their performances. This is to say the performers themselves don't fade into the background. Their work isn't judged alone and by itself as technically evolved curiousities. The works are frozen in time like amber, yes, but like in any other performance art, their personalities are fully on display everytime those works are perused, every time they make a public appearance or remark to boost their own publicity. It's those parasocial bonds they create which engender a following. AI would need to become not only competent at manufacturing the underlying goods but fulfilling all the other roles these freelancers take on. For freelancers, the whole operation is a one man show where they must become not only technical experts, but also frontmen who can appeal to crowds, manage public images, deal with other professionals. So an AI would need to be able to cover all of those bases in order to directly replace the individual doing them, or at least to turn them into proxies the AIs can micromanage.

That was already the status quo for artists, though- they were already waiting tables.

Centaurs (humans using AI tools) don't last forever. They are a transitional stage until the AI gets good enough to replace the humans entirely.

You could say that about photography as well, that the humans controlling cameras are a transitional stage until the cameras get so good that they don't need people. Yet it's been over a century, there's uncountable hours of photography being captured at all times all over the planet all the time and yet almost none of it is interesting outside of like the nature wildlife automated photography that some people like to look at. I could probably find beauty in automated traffic cameras if I felt like it but I'm in the minority and most people would rather see a human produced and curated series of photographs than a randomly generated or algorithmically curated series of photographs

Again I am an artist and designer and no one is telling me I'll have to flip burgers right now, I can think of 50 other more creative tasks I could monetize before I'm going to sweep a floor. lol. Besides that the bottle neck in the economics of art has never been supply side but rather demand side and AI has NO answer for the demand problem. I could already produce thousands of t-shirt designs a day if I wanted ten years ago, the technology was already there, but the demand for this variety of t-shirts does not exist.

I can sort of understand, because I had a skill that was scarce and valuable, and now it's, well, not. It's a big hit to my self-esteem. Artists are probably feeling the same way. Human exceptionalism is at its end; I've been watching it coming for years, so I could mentally gird myself. But for a lot of others, this just came out of nowhere (especially if they listened to the wrong brand of AI "expert", ahem), and it's hitting hard.

It doesn't make protectionism and knowledge-hoarding the right response, but it does make it an understandable one.

This all goes back to my hobby horse of arrogance and how bad it is. You alone have to be the scarce valuable thing and now there's someone else so you are sad? Oh my god. If cats evolved to suddenly be amazing artists I'd be so stoked to see the cat art. Computers evolved to be amazing artists so I want to see the computer art. If cats evolved to be amazing at math I'd want to see the cat math. Computers are good at math so now I want to see the computer math. There is no other voice or force in western society telling people to stop being so arrogant but it would solve so many problems if we had this as a cultural idea to even a tiny degree.

Artists are probably feeling the same way.

Did you read my comment you replied to I am an artist and not feeling a big hit to my self esteem if anything I am excited to have new tools available to express things that were completely out of my reach in the past.

Did you read my comment you replied to I am an artist and not feeling a big hit to my self esteem if anything I am excited to have new tools available to express things that were completely out of my reach in the past.

You are very much an outlier. The vocal parts of the art community are united against AI.

That's not what "doomer" means. Doomers are people who think there's a chance that a sufficiently smart AI will wipe out humanity.

Doomer is a far more general term which encompasses pessimists of all stripes e.g. climate doomers.

Can't mathematicians who don't make it in academia just become actuarials or high school teachers, depending how they weight work-life balance vs compensation? I literally don't understand why they think it's completely hopeless. You didn't get what you wanted, in a longshot career path. No need to rope. Just take the backup plan.

They can become janitors too. Knowing complex analysis doesn't make one unable to clean toilets.

I mean they can do that too, but it pays much worse and requires retraining.

Good Will Mopping will be a fun movie.

You know, this makes me nostalgic for when I'd effort post on AI advances in this thread. And they'd usually attract people arguing that it wasn't CW material.

Simpler days. Nobody seems to be disputing it now.

I'd dispute it but I favor a loose CW definition to avoid people posting the same boring topics all the time, though AI shit is getting there.

In a post-AGI world, all topics are AI topics.

I'm looking forward to the AIs shitposting hilariously (and not the glowie disinfo and suppression kind of shit posting)

You're confusing a lack of re-litigation with agreement. I thought, and still think, that AI discussion doesn't belong in the CW thread. But we already had the discussion and the community decided it was fine, so why would I continue to argue the point?

I still think we need a dedicated AI thread ("Skynet Saturday" or "Singularity Saturday", per Claude) to both discuss news and post generations.

Be the change you want to see.

I feel like this is almost inevitable. If people think there's too much AI content on the CWR now, next year is going to be rough.

Much like blue-collar workers turned out to be much less exposed to being made obsolete by AI than the white-collars who sneered at them, the 'hard' sciences are on life support after decades of sneering at biologists. Who knew our idiotic nomenclature, awful qualitative methods and black-box approach to running clinical trials were actually the greatest moat against superintelligence a group of mathematically incompetent wordcels could devise.

Interesting to think that AI is a 1000x coder, 10,000x mathematician, but a 2-3x biologist on a good day. At least I'll be afforded the joy of pointlessly pipetting liquids right up until the paperclippers come for me. On the downside, that's all you need to make turboCOVID while falling far short of curing cancer or aging.

the 'hard' sciences are on life support after decades of sneering at biologists

LMFAO yes absolutely. Even within physics the areas being dismantled are theoretical physics while the experimental physicists are carrying on doing their jobs just as well or even better before AI, because to them it's literally a complementary assistant that boosts them rather than replacing them. And even within theoretical physics the areas getting blasted the most are the most "rigorous" ones like QFT and GR while the ones that are still sort of saved from the worst of the shelling so far are things like lattice field theory etc. where people randomly go around doing approximations because otherwise nothing works...

I am in experimental nuclear/particle physics, more on the making-stuff-run than paper-writing side of things. For me personally LLMs are a huge boon.

I mean, on the one hand, they have made huge parts of my skillset obsolete. Before I was one of the few people in my department who knew how to run gdb to hunt down a segfault in some terrible ROOT program. This is not an employable skill any more -- Claude can do this far better and more reliably than I ever could at a tiny fraction of my cost.

But as someone lucky enough to have a job, Claude is like having a PhD student who happens to have read every book on every programming language ever working for you. Tasks which would have taken me half a year (because I lacked the specific skills) can now be done in weeks. I have even taken a liking to systemd of all things, because I don't need to learn the syntax myself. Going down some Linux rabbit hole (how do I figure out why an interrupt handler runs at 100% cpu?), which would have taken me days before is now a matter of an hour.

Still, it used to be that in my field, a lot of students were employed in data analysis. Generally, this involved writing terrible ROOT code to get some physics channel out of a dataset, then turning this into a publication and a PhD thesis. In the short term, they also benefit from LLMs -- at least for the code writing part, which used to be a majority of the work.

But solving for equilibrium, it seems unlikely that anyone will employ PhD students for data analysis in the future. They are expensive (compared to LLMs), they suck at programming, they start out with a lot less domain knowledge than the bots, they come with interpersonal conflicts which require management. In the past, there was no way around them (short of employing programmers, which would take larger salaries and require physicists to work interdisciplinarily). These days, a senior physicist can probably use an LLM agent to do some analysis without spending more time mentoring than she would for a PhD student.

Data analysis aside, experimental physics still contain plenty of manual tasks which are not easily automated. We have used PhD students to glue photomultipliers to scintillators before, and these jobs will still require doing.

The other question is the long term career path of physics PhDs. Before, most of them (in my field) ended up in the software industry, which is what kept our system running. Prospects there are probably not so great at the moment.

I’m trying to think whether I would continue my interest in politics and culture if AI effectively solved it. Probably not. I would likely just admire the results of the AI, rather than researching or exploring anything myself.

I think politics will be with us forever: Aristotle was right when he said that man is a political animal. Possibly in any kind of AI post-scarcity or even superabundance we'll have even more time to be political. Solve the material concerns and social problems are all that remains.

There will always be hierarchical status/face games. They will wear different faces, but they won't go away.

Edit: I should have said even if you want to ghost out post-scarcity and just do your own thing without harming anyone else, there will always be Someone That Thinks You Aren't Doing The Right Thing. They will come in and start interfering in your life. And then you're right back here.

ML also doesn't have a meaningful bank of rigorous conjectures

Sort of true. (I am a ML professor.)

COLT is the main conference for ML theory, and to ML researchers is considered more prestigious than ICML/NeurIPS/etc. It has a much lower impact factor, however, because it is truly a math venue about proving statistical theorems and does not accept any applied/experimental work. It has always had a track for papers presenting open problems. You can find the latest year's CFP at: https://learningtheory.org/colt2026/openproblems.html#cfp

I expect most of these problems to be "much easier" than standard math open problems and that these AI systems could prove most of them. The reason is that they don't receive the same amount of attention as "traditional" math problems. The number of researchers who can meaningfully even understand any one of these problems averages <100. That's similar to many math problems, but the difference here is that the ML researchers do not actually spend time thinking about the ML open problems because they are spending most of their time thinking about realworld ML applications and chasing $$$. Traditional mathematicians don't have either of these distractions.

There are a handful of classes of open problems where a resolution could meaningfully improve user experience with LLMs somehow. For example, there are open problems about improving the sample efficiency of reinforcement learning, automatic hyperparameter search, and A/B testing; and all of these are subproblems that the major labs have to implement to train their models. The actual math is abstract enough, however, that I don't see a lab keeping a result as a trade secret if they do resolve any of these problems.

difference here is that the ML researchers do not actually spend time thinking about the ML open problems because they are spending most of their time thinking about realworld ML applications

Guilty. Though I do think about sample efficiency, OOD performance, and methods for encoding expert knowledge as inductive biases, which I think would fall under COLT questions. I just have a practical real-world use case in mind...

My question to you is how many of these are represented as formal mathematical formulations that are solvable without experimental testing? It's one thing to write a tight math proof that can be checked in a Lean Solver, quite another to prove that your new formulation of GNNs functions broadly under distribution shifts in multiple domains.

You will find more greek letters than English ones in a COLT open problem, and never anything empirical. Here's a representative example I just googled: http://proceedings.mlr.press/v125/van-erven20a/van-erven20a.pdf


When googling for the above, I also stumbled on a 3mo old repo that has some AI resolutions to some COLT open problems: https://github.com/Pengbinghui/pipeline-math. The problems they resolve are:

  • Shuffled SGD — the SS–RS–GD inequalities (Yun, Sra, Jadbabaie, COLT 2021)
  • Learning measured-output quantum circuits (Kun and Reyzin, COLT 2015) (Partial solution).
  • Unweighted data selection for linear regression (Hanneke, Moran, Shlimovich, Yehudayoff, COLT 2025) (Problem 3)
  • Robust conditional probability estimation (Langford, COLT 2010)
  • Fixed-parameter tractability of zonotope problems (Froese, Grillo, Hertrich, Skutella, COLT 2025)

Of these, I've personally spent some time thinking about shuffled SGD and robust conditional probability estimation. The SGD problem is vaguely related to LLM training in a "that's interesting" kind of way, but not something that has any real world impact. The robust estimation result I could see being actually implemented in the backend of FAANG companies to improve performance of various algorithmic feeds and help them make a few more million/year, but no impact on LLM training.

You probably know more than me here, but it seems like the most interesting results would be bounding "intelligence" (which remains ill-defined) as a function of model size and training data size. There are some results in that direction, but no Shannon Information Theory class results.

It's be really nice to affirmatively prove "three sentences of training data and an arbitrarily large model cannot create the Omniscient Machine God." This seems an obvious result, IMO, but even "maximal intelligence scales sub-linearly with both training dataset and model size" would seem to dispell "FOOM" singularity concerns.

I teach basic NP hardness results in intro CS classes now specifically so that students can make arguments like this. I feel like everyone these days should know that AI will never solve a traveling salesman problem[*], even approximately! Of course, neither will any human :(

[*] the usual caveats of P!=NP and general graphs (not euclidean/metric) apply.

Fun results:

Hilbert's tenth problem over (\mathbb{Q}) is undecidable
The quasi-Riemann hypothesis
The rational Hodge conjecture holds for every CM abelian variety
Integer multiplication can be done in sub-log-linear time. Oh, there's also a sub-log linear DFT.

What can I say: Man Proposes, AI Disposes...

I do wonder how the reaction will be when, almost inevitably, some PhD student or postdoc gets scooped on a problem they had gone all in on and commits suicide. It's not like OpenAI seems to be following the standard process of extending feelers through the field to get in touch with any existing noosphere-homesteaders before solving these.

People still play chess even though computers have long been better at it than use humans.

I'm in programming, one of the fields that has already been most affected by AI. I can also report that I'm having some of the most fun and being more productive than I've ever been. My key has been to just embrace AI hard and use it as a tool the best that I can.

You also see the culture war starting up in that space. There are several Linux distros that have people pledge that they haven't used AI in the process of doing their contributions. My prediction is either: a) people are just lying about not using AI or b) they'll just lose and people who are using AI will simply outperform them in the marketplace of ideas.

So math people simply need to get on board. The train is leaving no matter their opinion of it.

The difference is that people largely only ever played chess for the love of the game. There are nontrivial numbers of people who do programming for the love of the game; in my experience, though, there are very few mathematicians who derive sufficient meaning from the "game"/process alone, and those who are are unlikely to find themselves in such high echelons of academia, instead gradually filtering out into other fields where you can continue pursuing it as a game (ex.: Martin Gardner, Greg Egan...).

I'm certainly in the "love of the game" category of programmer. I honestly think that you have as many in math. And unlike computers, few people ever went into math to get rich -- so my internal estimate is even higher for the math side.

Side-topic: I think the programmers who are in it for just money are the ones that are going to get selected out. I've been at this for a long enough that when I started a programmer was a good white-collar job, but not one to get you rich. My experience, and frustration, with how the field has changed over the decades is that the quality of people has become absolutely dreadful outside of a few highly selective places. And even places like Amazon (where I spent some time) are compromised.

I get joy from many things. From tweaking things big and small. I don't think I ever really got joy from the act of writing code. It was always a means to an end. And the end doesn't have to be strictly useful. Something I've been kicking around for the past 35 years (also: crap... I'm feeling old) is an implementation of the Game of Life from John Conway who would fit in with your list perfectly. I came up with what I still think is the fastest way to play it with "normal" code in the general case. Using Claude I was able to get this implemented with a front-end that I can interact with. I get joy from the algorithm. Joy from thinking of low-level optimizations. Joy from learning more about the innards of how the processor I'm on handles cache lines.

I don't get joy from setting up a build system. No joy from setting up a front-end that runs on my Mac. Very little joy from tweaking an unrolled inner loop. Less joy from setting up unit tests.

Having AI lets me get to that joy. And it makes me incredibly happy. And it lets me paper over the parts I consider ditch digging.

My gut says that the folks in math want the same thing. Come up with an idea, and have something else handle the proof. There's joy in simply making something better or pushing the frontier a bit more.

And with AI, you can make bigger leaps. Ask bigger questions. Get more joy from discovery.

In case you're interested: https://github.com/gburgyan/vlife and a technical paper in docs that's more presentable.

There are absolutely people in math for other reasons than money, but among the countless mathematicians I've known very few have been motivated by anything that did not factor through being the first to prove/discover/define something that others care about.

That's the thing exactly. If, through using AI tools, you can be the first to prove something, isn't that a good thing? Only through discovery can we reach higher to ask bigger question. It took a lot of work as a civilization to get to the point where we had the language to construct things like the Navier–Stokes equation in the first place; they don't just get written down in a vacuum.

So we've gotten to a step change in the system. The real challenge will be to be the first the ask the question in the first place. Even if that is automated away, you're still left with the pure joy of learning and discovery which is rewarding unto itself.

I can also report that I'm having some of the most fun and being more productive than I've ever been.

Same here.

Culture war would happen. The two main sides would be 1) "This is terrible. We should control AI and put my preferred political ideology in charge of society, even the parts of my ideology that have nothing to do with AI." and 2) "But you didn't care when the coal miners were committing suicide. We should put my preferred political ideology in charge of society, even the parts of my ideology that have nothing to do with job loss."

I think we’ll invent a new entertainment form of ‘challenge math’ or something where the idea is you try to figure out a proof you haven’t been exposed to competitively even if AI has the answer, and it’ll be like a game. Either way, I really don’t think this is reason to end it.

OTOH, when manuscript traditions and horses went obsolete, we didn't make this stuff major competitions. Calligraphy is now an art form that most people never encounter and horses are a luxury despite the government giving them away for free.

Horse breeding was and is controlled by humans. The dial was simply turned down as demand fell over 50+ years between around 1900 and 1950. Working horse lifespans are maybe 20 years. The human equivalent would, at worst, be something like a one child policy for successive generations and/or cratering birthrates which we already have.

despite the government giving them away for free.

Trump's trying to swing midterms with the "free pony" program again?

No, there’s been a free pony program for a while. The bureau of land management has free ponies if you want one.

You have to go find it yourself for that one though, don't you? Kind of like how there are 'free fish' in Central Park?

(ED: or ducks even)

No, they will literally ship you a free pony that they rounded up in the desert. It causes a public outcry among the sorts of people who keep track of government programs to cry about when they euthanize them or sell them to quebecois butchers.

Huh -- I had heard of permits or whatnot but thought you had to go get the horse yourself -- surely shipping is not included!?

What a country!

(that said I've heard that such horses are pretty impractical to train most of the time, so yeah people are probably eating them or something...)

Plenty of horse-involving sports still

We still use calligraphic scripts ornamentally, right? On books and signage and such.

Though I’m more exposed to East Asia where there is probably more of a calligraphy scene and more calligraphic script (even if printed) as decoration or for style.

There’s cursive script(not the same thing) used ornamentally very heavily, but most people in the west never run into calligraphy. Ever. You can hire an artist to do calligraphy for you if you’d like, but it’s quite rare.

I suppose my background is colouring this, then, especially since I dabble in calligraphy.

My immediate question was about whether this closer to the current peak of performance, or just a lazy demonstration of OpenAI's power. So, I took one of the particularly interesting preprints to me (memory and precision in Gaussian models) and tested whether it's at the edge of capabilities or not. 30 minutes of back and forth with Astra (itself behind OAI's internal model) resulted in a significantly stronger result, on multiple dimensions. While I finished my first bottle, I spent a fair amount of time convincing myself of the result; I was convinced the strengthened results were plausible. Write this off as AI psychosis if you want, but try it yourself; I'm genuinely curious for what you get. (The back and forth, here, was entirely me saying "you can do it!", "keep at it, I believe in you", and "you've got this, finish it!")

Did you have it make a proof in Lean, or just spew plausible sounding math?

The spew, of course; can't waste too many tokens.

But since the original proof came with Lean code, I'll kick off a run to generate the Lean code for the extension. Sit tight. (I wonder if Lean code has enough degrees of freedom to dox me via watermark. Oh well, the omniscient machine god will be able to infer who I am anyway.)

This isn't a case of OAI spending millions for a marketing bump.

"Hours of ChatGPT Pro thinking time" is a very fun unit of measurement, because OpenAI recently added an "Ultrafast" tier of Astra, which runs at 300 tokens/second, to the highest Pro subscription tier.

At API pricing this model costs $300 per million output tokens (assuming, charitably, that they're not using long context mode). Conveniently, this works out to about $300 per hour of runtime, if you don't count input tokens or cache usage.

So each individual attempt at a problem cost about a thousand dollars. Even assuming OpenAI only tried their 4,000 attempted problems once each, that's at least four millions they spent on this particular marketing bump.

In other news, this morning someone used ChatGPT to improve the new bound on multiplication by a factor of 2^104.

T(n) = O(n (log n)^(1 − κ)),

With κ = 2⁻⁷⁸ (tightened from κ = 2⁻¹⁸²)

Conspiratorially, my inclination was to think they filtered them out for competitive advantage. I can't find any evidence of that in the GitHub repo or any of the preprints (equally plausible: deduping), so maybe they judiciously decided not to point their mathematical ballista at ML.

They have absolutely pointed the ballista at machine learning. What you're seeing is them briefly pointing the ballista anywhere else to produce something to share, because there is no way they are sharing their trade secrets and competitive advantage publicly.

I'm quite sure they have. These math results are likely almost an afterthought; I wouldn't be surprised if they've already put 10x of the compute they used for these math results into researching ML.

10x? More like 100-10,000x I'd wager. See the piddling $4 million estimate above. It makes little difference even if it cost them $40 million instead. >>$100m is the ballpark of where it might make sense for them to not try, because the reputational aura and PR is worth it.

What did people think RSI meant? Vibes? Papers?

"sub-incremental counterexamples to previous conjectures with no practical applications"?

No doubt. I'm sure you could describe the majority of human-derived mathematical findings in the exact same words. I'm sure figuring out optimal square-packing has made billions of dollars for someone.

Nobody is making any money on a 1-10^xx exponent on integer multiplication's big O -- the fact that you think this shows that you are relying on an unreliable interlocuter to form your opinions.

I don't dispute that some of the other results may have some useful application, but the ones of this form are only interesting in the sense that humans tend to expect round numbers, and have been mostly correct in this assumption to date. (I do share this expectation and personally think it's more likely that there's some kind of mistake in these proofs -- but if I'm wrong, that's interesting!)

Your bot can probably explain this to you if you ask, but briefly: Big-O notation typically disregards the portion of algorithmic complexity that's constant doesn't and depend on the size of the dataset; ie. whether a loop takes a millisecond or a minute to execute once is not considered. Something that takes a minute per cycle but is O(1) may be better than something that does it in a millisecond but is O(n^2) -- it depends on how many "n" you are interested in though.

TBH I haven't looked into the actual algorithms involved here, but given the length of the proof I'd expect the constant term (call it overhead) to be quite high compared to existing methods -- considering the tiny difference in the exponents, you would need n to be literally approaching infinity for the proposed algo to make any difference at all. In the case of the integer math, this would mean that multiplying integers of ~infinite length could be somewhat faster -- but we can't even test it, because computers don't have infinite memory.

TBH I haven't looked into the actual algorithms involved here, but given the length of the proof I'd expect the constant term (call it overhead) to be quite high compared to existing methods -- considering the tiny difference in the exponents, you would need n to be literally approaching infinity for the proposed algo to make any difference at all. In the case of the integer math, this would mean that multiplying integers of ~infinite length could be somewhat faster -- but we can't even test it, because computers don't have infinite memory.

Per wikipedia, the existing O(n log n) algorithm for integer multiplication that the AI improved on is already considered "galactic" - i.e. only optimal for impossibly large data sets. The practical algorithm for the largest numbers we can realistically multiply is a FFT-based approach which is O(n log n log log n)

To my non-expert eye, using the new slightly-faster DFT in the AI papers in the bog standard FFT-based multiplication algorithm should give you a faster multiplier than the new multiplier the AI found explicitly.

The FFT in the papers looks pretty galactic itself though; something like a 1- a trillionth exponent?

Nobody is making any money on a 1-10^xx exponent on integer multiplication's big O -- the fact that you think this shows that you are relying on an unreliable interlocuter to form your opinions.

Do I think this? Nobody told me. It's not true. I've never said it, I've never thought it, and you can go look.

I'd prefer you stick to what I've actually said. Or if that's what you thought I implied, I ask that you at least acknowledge when I clarify otherwise.

I also know how Big O works, for the record. I took CS courses in my free time once. I am a nerd.

You just said it though! (well, maybe your LLM did)

I also know how Big O works, for the record.

It doesn't really seem like you do; an optimal square packing algorithm that is better only in the case of infinite squares is not making anybody any money either.

More comments

I'm sure they have pointed it at ML, but ML isn't exactly the same as formal math. I'm probably describing this poorly but ML is very experimental/empirical, not theoretical like formal math. The math heavy side of ML would be finding better optimizers, regularizers, improving backpropagation or finding a better method than SGD. The question becomes how much of an edge do these provide? Because better algorithmic components still have to deal less-better data or compute components and how well all those now perform better is not theoretically provable.

What most of these breakthroughs tell me is that supervised learning is highly effective at formal math, and the benefit of something like Lean + Solvers to allow for computational checking of math proofs has been a fundamental driver. Whether or not strategies learned on this frontier are applicable to broader fields that are less formalizable is unknown to me.

Much of the empirical part of ML, though, is also something the LLM can do autonomously. The agent can write up some pytorch or jax to train a model on an existing dataset, observe whatever quantities you want, repeat. It does have a longer feedback loop and more compute requirements than pure math, but quite automatable.

The pure math part of ML is unfortunately quite weak right now and hasn't played a huge role in the current boom; pure math research could provide a much stronger basis for understanding learning and creating new approaches, and I think it will, but that's speculative.

We'll find out soon.

The longer feedback loop is unfortunately the expensive part, as we talked about last time. The real game changer would be to reduce the compute cost and training time for LLMs and/or figure out a more efficient, non-quadratic self-attention mechanism. However, at some level the cost is the moat frontier labs have, so there is almost perverse-conflicting incentives not to improve it but also requiring it for RSI to happen.

The pure math part of ML is unfortunately quite weak right now and hasn't played a huge role in the current boom;

Agreed.

pure math research could provide a much stronger basis for understanding learning and creating new approaches, and I think it will, but that's speculative.

I'm sure the mathematicians left out in the cold by the big data revolution of ML will rejoice that there was indeed a mathematical approach to LLMs over the pesky engineering approaches that were developed by the peons from MIT.

I'm sure the mathematicians left out in the cold by the big data revolution of ML will rejoice that there was indeed a mathematical approach to LLMs over the pesky engineering approaches that were developed by the peons from MIT.

Damn this reminds me of the ML class I took back in my maths degree. We spent the whole term proving theorems about perceptrons rather than being taught any sort of ML that was being used in the world.

Rite of Passage honestly. I had one similar about the fundamental theory of AI, lots of proofs. While fascinating I actually wanted a class of how to train neural nets back when Torch was still in Lua and how to use Cuda.

I'm not able to appreciate these results because I don't understand research level math. It seems that my inability to appreciate these results is the only real consequence of my ignorance, as it doesn't look like any of these discoveries will materially change any of our lives (except maybe mathematicians' lives).

I'll be impressed when cryptography is broken.

I'll be impressed when cryptography is broken.

Cryptographic results are strikingly underrepresented in this batch. But there are a couple possibilities: 1) OAI didn't attempt any; 2) OAI attempted many and failed on all of them; or 3) OAI attempted and succeeded on many of them, but some force incentivized them not to publish the result.

You might be waiting for something that is being intentionally withheld. Atomic energy research didn't suddenly hit a wall in 1942, even if no papers on it were published.

This is my problem. My lack of math expertise means that a scenario where AI just unlocked the secrets of the universe and a scenario where they spent millions of dollars to solve some pointless toy math puzzles that we shouldn't have been funding mathematicians to work on in the first place look exactly the same to me. It's frustrating how no one impressed by these accomplishments seems to have made any effort whatsoever to explain why any of this should matter to a layman. Does Open AI not have any PR people on their team who know how to draft a proper press release?

It's frustrating how no one impressed by these accomplishments seems to have made any effort whatsoever to explain why any of this should matter to a layman.

It's been <24 hours. If you want discussion from older results, see any of the famous math youtubers like 3blue1brown for layman-accessible intros: https://youtube.com/watch?v=TfyPshgMbug.

After discussing with GPT, my strong suspicion is that this is because there isn't much to say. The results broadly fall into:

  1. Tiny improvements (1^1.128) to the effeciency of things we already knew how to do.
  2. Solutions to problems for which we already have far more efficient approximators, or which only apply under conditions which will never apply.
  3. Proof of things that we already knew with great certainty but couldn't prove.
  4. Insider baseball.

What was the most recent mathematics result that didn’t fall into one of those categories?

They benefit from maintaining mystique if the actual successes are less than what the hype suggests (as you hypothesize). You're not giving their marketing departments enough credit if you'd think they'd fumble not trumpeting something of serious consequence where investment inflows are concerned, especially given a pattern of relentless, and indeed unprecedented, hype mongering.

They have PR people but they don't understand the results either.

I believe one of the AI companies published a break in one of the post-quantum crypto finalists (HAWK). It wasn't the one NIST settled on, but I bet they looked at those. Note that NIST published crypto standards have at least once in the past been resistant to attacks not yet published when they were chosen.

But also, the P=NP Millenium Problem, which hasn't been claimed yet, would have very strong cryptography implications depending on the direction that goes.

But also, the P=NP Millenium Problem, which hasn't been claimed yet, would have very strong cryptography implications depending on the direction that goes.

If a constructive proof of P=NP is found, smart money would make an easy trillion+ dollars mining bitcoin and shorting the cryptocurrency markets (and all digital infrastructure stocks, for that matter) before claiming the piddling little Clay Millennium Prize.

That's for amateurs. I'd open a travelling salesmen route-planning company!

Depends on if the construction is galactic or not. We might have NP=P formally, but never be able to use it on any remotely practical problem. The world would mostly continue running as is.

I'll be impressed when cryptography is broken.

A lot of people give lambs to the pure math status cult, so those people need to be impressed since AI just automated pure math completely.

AI became the best mathematician alive when it solved the Navier-Stokes problem. How does this change anything?

Well, AI skeptics questioned the NS result at the time of release, with variations of "AI stole it"; "it was just providing a solution that blows up instead of doing real math, so it doesn't count"; "OAI spent millions of dollars doing it, I could probably have done the same if you gave me an eight figure check." I think that this new release should put those concerns to rest.

Though, who knows: Wired (the onetime home of Kevin Kelly!) has decided to cover the most momentous day in mathematical history with "OpenAI is pissing off mathematicians".

The cost is what should drive any updates. These took three hours each; that's cheap enough for pretty much anyone to be able to produce these kinds of results, when they get access to models of this class. And that cost is rapidly falling. This has substantial relevance for both how quickly we should expect AI to diffuse as well as how practical RSI is.

Wired (the onetime home of Kevin Kelly!) has decided to cover the most momentous day in mathematical history with "OpenAI is pissing off mathematicians".

The article:

The group, she tells WIRED, asked the firm not to just publish them in a blog or tweet, like they had done with 10 problems earlier that month. Instead, it would be important for OpenAI to publish papers explaining the work so that mathematicians could absorb, digest, and use the results, according to Kra. “Apparently, that input was ignored,” she says.

But the GitHub repo contains hundreds of papers. Will this satisfy the mathematicians? Somehow I doubt it.

The group, she tells WIRED, asked the firm not to just publish them in a blog or tweet, like they had done with 10 problems earlier that month. Instead, it would be important for OpenAI to publish papers explaining the work so that mathematicians could absorb, digest, and use the results, according to Kra. “Apparently, that input was ignored,” she says.

But ... at the top of the linked blog post is a "Read the Paper" button, and it wasn't a broken link; the corresponding paper explaining the work is in the Wayback Machine that afternoon.

I'm going to give her the benefit of the doubt and assume that the reporters (who aren't directly quoting her here) are mis-paraphrasing or mis-summarizing an argument that was just questionable (something like "the paper should come first so independent experts can look it over before it's hyped up") and it was they who turned it into outright libel like "just publish them in a blog or tweet, like they had done". I still remember a coworker's wince when he caught me talking up our work to a PR person in merely technically-accurate language rather than accurate-and-carefully-worded-to-be-nearly-impossible-to-misinterpret language. Even if Kra had been responsible for the libel, Wired would still have skipped basic fact-checking here, so Occam's razor is pointing at them.

But regardless of the details, at least somebody involved in this article is failing, ironically at a level which I would have considered to be evidence for "clearly not general intelligence yet!" in an AI.

Nah. If you were human and that were your only achievement, then sure, the attention-deficient celebrity chasing normie press might lionise you as such, but those in the know would not rank you above people who have demonstrated a lifetime of consistent progress on hard problems and good intuitions (such as Terry himself, and several of his associates).

There is only one Terence Tao, but there are 47 living Fields medalists, and solving Navier-Stokes is going to get someone a Fields medal. If "like Tao" is your standard for a successful mathematician, it is too high.

The parent poster literally said "AI became the best mathematician alive". I was therefore not talking about my standard "for a successful mathematician". My standard for being the best mathematician alive does involve being as good as any other mathematician alive, which does in particular mean being >=Terry.

How is Tao better than AI? In a year AI will easily have a bigger and more important publication record. The AI that exists right now, that is, run on math for 1 year.

I wasn't saying Tao was better than AI - I was saying that he is better than the average Fields medalist.

I think Navier-Stokes counts as a human-AI collaboration, possibly the last one we shall see.

There's probably a double digit number of grad students staring into a glass of whiskey tonight and thinking of hanging themselves.

I mean, maybe, but only because the news hasn't hit everyone yet.

I'd expect four digits in a week.

I'd bet on <5 in a week. At least using the standard where they explicitly blame their obsolescence in writing before killing themselves.

To be pedantic:

  • Completed suicides
  • Mathematics majors
  • Explicit reference to AI
  • I'm willing to allow for other methods of suicide

Yeah, if we're taking about actual suicides rather than mere ideation that drops the numbers a lot. I'm not sure we have a disagreement here.

I’m not generally a fan of reading the tea leaves of suicide notes but I think it’s best to interpret these - even if we literally accept their reasons for suicide as true (and I think generally this is the case when intelligent people kill themselves, that is to say that their stated reasons are usually their actual reasons even if they are bad reasons) - as borne of fear rather than existential crisis, and that fear is ultimately tied to material circumstances. There is a lot society can and will need to do, if it survives, to repair our sense of meaning as it’s redirected away from work.

If a four digit number of graduate students are thinking about hanging themselves in a week's time, that would be a dramatic improvement in graduate student mental health. (The US graduates about 60,000 PhDs a year, multiply by the number of years the average PhD-bound graduate student spends in a suicidal state).

Not everybody drinks whiskey or picks hanging.

A rigorous mathematician would atleast explore the potential of hanging

Although Dr D'Eath (an actual maths lecturer at Cambridge) eschewed all forms of ligature maths, preferring a more direct approach - he lectured on Killing's equations.

OpenAI's numbering goes upto 377, but they claim to have only 372 proof families. Thus some numbers are missing. These are 45, 61, (Algebraic and Complex Geometry), 70 (ACG or Real and Complex Analysis), 123 (Theoritical Computer Science), 163 (Combinatorics).

Yeah. This made me suspicious, and it seems like not reindexing is a mistake a human would make, not a model. This seems like a last minute redaction to me. I can't find any clues in the repo about what they were, though, aside from the general fields you mention.

I don’t see why mathematicians would be so upset. 99% of professional mathematicians have never contributed to the field of mathematics in any real way.

There are two arguments about AI and employment.

The first and primary one is about mass unemployment. This is a real issue but people obsessed by it also need to understand, as I think @self_made_human increasingly does, that basically all of the surge in white collar labor since 1970 consists of fake make-work very thinly disguised by the laws of classical economics. The invention of the computer literally automated almost all white collar labor as it existed in, say, 1920, and yet instead of an unemployment crisis we just gave ourselves fake titles and decided to spend all day having meetings and making PowerPoint slides. The format in engineering and medicine was very slightly different but vast amounts of both professions, their demand and their provision were and are make-work too. I go to the doctors office and a GP diagnoses the way AI can in 20 seconds or a competent and cheap nobody who could use a search engine could do in 5 minutes even 25 years ago. If society survives the AI X-risk, we will emerge into a world of such unfathomable abundance that it will be trivial to provide every human, certainly in the rich world but eventually the whole world, with a comfortable material life. Dying from an AI engineered virus or a drone war or something is far more likely than starving because AI replaced all jobs and the owners of capital are too tightfisted to share.

The second and more interesting one is about meaning. The mathematician, even though he never meaningfully contributes to the field, imagines that he might. This is why AI makes him depressed. It’s the lack of hope that kills you, but is it? I mean, is it really? Meaning is invented, we’ve spent thousands of years inventing rather arbitrary forms of meaning. We can invent more. We can have kids. We can touch grass. We can compete against each other, if we want to, in games and competitions. Academia is already one such game. If mathematicians still want to play it, all they have to do is change the rules.

I guess I’m more sanguine on this than most. I increasingly converge on one observation, which is that it’s extinction or utopia. Other scenarios (weird AI hell) are plausible but increasingly unlikely. Compute always wins, the bitter lesson is always true, so the only question is whether it destroys us entirely (or someone uses it to) - accidentally or deliberately are really irrelevant - or whether it liberates us from material want entirely.

basically all of the surge in white collar labor since 1970 consists of fake make-work very thinly disguised by the laws of classical economics.

Eh? I'm skeptical of the majority of white collar labor being "bullshit jobs", for reasonably standard arguments. Especially before 2023.

Which laws of classical economics?

If anything, the classical tradition predicts the opposite: firms carrying dead weight get undercut by firms that don't.

Computers didn't automate jobs so much as they automated tasks, and the freed-up labor went into tasks that had previously been too expensive to bother with.

IMO, a lot of nominally "bullshit" jobs aren't. Plenty of HR and compliance roles are a rational response to entirely real regulatory and legal burdens. If firing your DEI manager means you're exposed to more expected liability in litigation than you spend on their salary, you pay them. You may think the regulation itself is dumb, and you may well be right, but then your complaint is with the legislature.

The firm is responding correctly to the incentives in front of it! A fire extinguisher that never gets used hasn't failed at its job!

I am not going to posit a precise fraction, but I will die on the hill that before LLMs, truly net-negative employees were not the majority. "Less productive than they could be" is a long way from "fake." Imagining some alternative universe where companies and economies run ruthlessly rational as the baseline, and then measuring reality against it, is bonkers.

I go to the doctors office and a GP diagnoses the way AI can in 20 seconds or a competent and cheap nobody who could use a search engine could do in 5 minutes even 25 years ago.

Please note that I am not the kind of doctor who puts the profession on a pedestal. I've been loudly predicting medicine's eventual obsolescence for years (and saying it's late in the queue, after all the other poor bastards), and I think the Nolla Health acne approval is the canary in the coal mine. But this claim is remarkably incorrect.

Any donkey with WebMD could diagnose the common cold. Similarly simple issues might account for >80% of the problems attended to by all doctors. It's the other 20% where the additional expertise and clinical experience draws a premium. The numbers are approximate, of course, but the distribution has a long and expensive tail, and that tail is where people die.

We don't even have to speculate about the "competent nobody with a search engine." Someone measured it, and not 25 years ago either. A 2015 BMJ audit ran standardized patient vignettes through 23 online symptom checkers, which is about as good as "algorithm plus internet" got in the pre-LLM era. They listed the correct diagnosis first in 34% of cases, and gave appropriate triage advice in 57%. The follow-up in JAMA Internal Medicine pitted physicians against the same vignettes, and physicians listed the correct diagnosis first 72.1% of the time versus 34% for the checkers.

Breaking it down further: physicians did better on high-acuity and uncommon cases, while the symptom checkers did better on low-acuity, common ones (when compared to how they did on the difficult ones, the doctors beat them in all scenarios). That's the 80/20 split, showing up in the data. And those vignettes included medical history but no physical exam, lab results or blood tests, which strips out a good chunk of what a clinician actually brings, so if anything the gap is understated.

Also:

A 2021 JAMA Network Open study⁠ asked 5,000 US adults to assess clinical vignettes before and after internet searching. Diagnostic accuracy improved from 49.8% to 54.0%, with an average search time of 12.1 minutes. Triage accuracy did not improve significantly.

Of the 5000 participants, 2549 were female (51.0%), 3819 were White (76.4%), and the mean (SD) age was 45.0 (16.9) years. Mean internet search time was 12.1 (95% CI, 10.7-13.5) minutes per case. No difference in triage accuracy was found before and after search (74.5% vs 74.1%; difference, −0.4 [95% CI, −1.4 to 0.6]; P = .06), but improved diagnostic accuracy was found (49.8% vs 54.0%; difference, 4.2% [95% CI, 3.1%-5.3%]; P < .001). Most participants (4254 [85.1%]) were anchored on their diagnosis. Of the 14.9% of participants (n = 746) who flipped their diagnosis, 9.6% (n = 478) flipped from incorrect to correct and 5.4% (n = 268) flipped from correct to incorrect. The following groups had an increased rate of correct diagnosis: adults 40 years or older (eg, 40-49 years: 5.1 [95% CI, 0.8-9.4] percentage points better than those aged <30 years; P = .02), women (9.4 [95% CI, 6.8-12.0] percentage points better than men; P < .001), and those with perceived poor health status (16.3 [95% CI, 6.9-25.6] percentage points better than those with excellent status; P = .001) and with more than 2 chronic diseases (6.8 [95% CI, 1.5-12.1] percentage points better than those with 0 conditions; P = .01).

I gain real value as a doctor by referring a patient to another doctor of a different specialty.

Because they know something I don't.

If medicine were mostly lookup, referrals would be mostly theatre. They aren't. Specialists rely on knowledge built from thousands of patients, and before 2023 or so, nothing you could type into Google came close to replicating it.

This is still not entirely wrong today, because AI suffers from a lack of embodiment. A model can't palpate an abdomen, notice that the patient who says he's fine is grey and clammy, or smell the ketones. What has changed is that you can now absolutely use ChatGPT to let less-trained medical staff sub in for doctors, with the nurse or PA acting as hands and eyes and the model supplying much of the knowledge. I expect that to eat a large share of what I do. But that is a 2024-onwards development, and it's the reverse of your claim: the reason LLMs threaten doctors is that they finally replicate expertise that search engines never could.

(You can find me on record saying all of this years ago. It represents a catastrophic risk to me even without 100% unemployment. 90% of us ending up unemployed is about 90% as bad. Even paycuts would hurt too.)

While AI continuing to be exceptional at mathematics doesn't particularly make me think that it's more likely that most/all labour is imminently automatable, I do agree with your point that if this is truly the case, there's really no point in trying to get rich to avoid dying on the streets when the 0.01% decide to mass unemploy the entire population or whatever. Coordination issues amongst the powers that be aside, it is obviously just not a stable equilibrium for the middle class to be unemployed and starving, while some dude who happened to own the right stocks at the right time is allowed to enjoy a nice lifestyle in perpetuity.

Either most everyone manages to secure material abundance via some combination of make-work and redistribution, which honestly already describes significant swathes of modern economies as you say, and having been rich perhaps confers some marginal status advantages and nothing more, or Something Happens that having a bunch of cash doesn't meaningfully insulate you from: e.g communist overthrow of the state (for real this time), violent state repression and expropriation of most the population, or AI loss of control where everyone dies.

I don't think I'm really all that sanguine re: the question of meaning though. The middle-class in developed countries are already mostly post-scarcity relative to a pre-industrial and a pre-computing basis in nearly all ways (the big exceptions would probably be in zero-sum usage of land and the decay of the body). Yet despite this, most people in the middle class are various flavors of miserable, generally aren't able to generate their own meaning ex nihilo, and mostly drive themselves mad via status comparison and/or the culture war.

If you thought hyper-partisan politics and status competition was bad in 2026, just wait until it's impossible to better your position via actual economic contribution, and the only way to access more resources, status and positional goods is to take it via the political process out of the arms of another. Even if the "utopian" world pans out, I frankly can't see this leading to anything but mass misery even if material needs are met in abundance.

Yeah, this whole "permanent underclass" thing is a load of baloney. You aren't being saved from the permanent underclass if your net worth today is $10m or even $25m vs $1m or $10k. If that very specific scenario plays out, having a moderate amount of assets doesn't insulate you, instead it makes you juicy food worth concentrating on for the real big multi billion dollar players.

I'm decently well off and I've come to the (very sensible, I think) conclusion that in the unlikely event of rich oligarch AI dystopia, my lot lies with the ordinary man rather than the centimillionaire+ class.

Yeah, rich people also aren’t coordinated, most don’t have bunkers, they’re highly divided by ideology, and most have some loyalty to some non-super-rich groups (for reasons of tribe, faith, nation, ethnicity, political belief and so on). This plus the speed of the true singularity will make the requisite coordination almost impossible.

The ‘permanent underclass’ thing is a legacy of Great Recession (‘Humans Need Not Apply’; Elysium) or even dotcom bust era (as in the case of Marshall Brain’s Manna which I think was 2002) AI fiction where self driving cars and self checkouts and robotic nurses and cashier screens at McDonalds automate low paying service jobs while the rich and upper middle class prosper, without overall productivity seemingly increasing much, which means the poor basically live off worse-than-section-8 level welfare forever. This was nonsensical even if one agreed with those predictions (automating a lot of these low level service jobs would be a huge productivity uplift that would make things like welfare much cheaper to provide and afford those on it a much higher quality of life), but of course it turned out that actually white collar work was the easiest thing to automate (before truckers in many cases, even!).

It also features impossibly misaligned incentives, since the ‘bottom of the rich’ would know they were next in line for the permanent underclass at any time, something that eventually cascades until there’s what, only one man left out of the permanent underclass? That doesn’t really make sense.

I don’t think democracy or what now passes for it can really survive AI, and I question whether that’s a particularly bad thing anyway, but I expect that any non-extinction scenario eventually plays out as abundance almost impossible for us to imagine today (like a wild animal subject to constant hunger, thirst, sickness, fear, predation, injury, etc ‘imagining’ life in a nice Northern European zoo), such that twenty first century distinctions become irrelevant. I also think AI plays the role of the zookeeper in this scenario, and so the idea that it lets Elon Musk control star systems just because he happened to be very rich and powerful in the years before it took control seems about as likely as the zookeepers letting the big elephant eat all the others’ food just because it happens sometimes in the wild.

The invention of the computer literally automated almost all white collar labor as it existed in, say, 1920, and yet instead of an unemployment crisis we just gave ourselves fake titles and decided to spend all day having meetings and making PowerPoint slides.

Wouldn't we expect competition to have optimized this away? This is at the heart of the worldview disagreement between "capitalism is just the invisible hand in action" versus "capitalism is a conspiracy by the powerful to remain in power."

the argument is that markets are more efficient at resource allocation than central planning, not that they’re very efficient.

When Musk bought Twitter and now operates it with 20% of the staff it had, or when Google increases headcount by 1000% in a decade despite operating primarily the same product, this isn’t maximum efficiency. Managers establish fiefdoms. Individual employees care more about themselves than maximizing shareholder returns. Incentives are not fully aligned.

Twitter and Google both had relatively low labor-to-capital ratios to begin with. So, if true, they were increasing headcount by x% from a low base, meaning it might not have been very consequential for their bottom lines. The examples could be cherrypicked anyways, and do not necessarily have bearing on overall patterns. Overall, we've seen the economy restructure itself from a primarily agrarian one, towards a manufacturing one, towards a retail and service based one across the past century or so, and all of that indicates a responsiveness to actual economic demand.

In fact, we've not only seen a shift in employment patterns towards human-necessary roles, but we've also see a rise in economic dominance by companies with relatively small labor footprints. Tech companies, in general, as mentioned, do not employ many people in relation to how much money they rake in, and they've been dominating economic growth and market indices for the past few decades. If they were insensitive to the costs of employing unnecessary labor, and if this was the basis by which the entire economy was organized, we might see them gradually becoming as labor-heavy as other, previous market leaders.

But we don't see that. We see them maintaining unprecedentedly lean workforces. That's consummate with a representation of the economy where humans in data-organizing roles are increasingly automated away. These humans don't disappear, but fill in gaps elsewhere, such as in retail and healthcare, providing tangible physical services. Pet grooming has seen a boom in employment.

Some of them do apparently move into roles which are of dubious practical value, but which have been necessitated in large part because of regulation and bureaucracy, such as consultants. However, these regulation regimes have led to decreases in pollution and other harmful things, so the consultants and such can be said to be necessary components of these overall systems' functioning and thus their benefits. Furthermore, they do plausibly provide value in the expertise they are able to provide as well.

Overall, I do not see a basis for thinking the economy is unresponsive to actual supply and demand inputs when it comes to labor organization. The level of its efficiency may be up for debate, but it does appear to shift over time in relation to observable developments.

Also even maximizing shareholder returns =/= maximizing productivity in every case. A company can make a very share-price friendly move to invest in something unlikely to succeed that is the current recipient of market hype, even if it's unlikely to actually work out.

it’s extinction or utopia

We shouldn't have mocked the people from the "What's your job on the leftist commune?" Twitter thread. Seems like they might've been right after all. If we survive as a civilization and do not get made redundant via a BNW-like transition to a tiny humanity.

There's no "we" when you and all your children didn't make it through the cull.

I'm sorry, I don't get it.

Wrong reply, my mistake.

I don’t think mass automation AI future proves communists correct. Not only because the timely invention of AI required capitalism (yes we don’t have the counterfactual, but I’d say it seems likely) but even more importantly because the only way any kind of ‘from each according to ability, to each according to need’ economy can work is if unlimited ‘ability’ is provided by AI and unlimited ‘need’ can be absorbed by it.

Well, communists did often say that capitalism would bring about its own destruction. Communists also often said that far from capitalism being useless, it is in fact the necessary precursor for communism.

It's hard to give them full props for these ideas given that originally these ideas did not envision something like AI being the mechanism.

However, communists did think that increases in production would play a major role in the transition, and they were aware that technology is key in increases in production.

Marx: "This “alienation” (to use a term which will be comprehensible to the philosophers) can, of course, only be abolished given two practical premises. For it to become an “intolerable” power, i.e. a power against which men make a revolution, it must necessarily have rendered the great mass of humanity “propertyless,” and produced, at the same time, the contradiction of an existing world of wealth and culture, both of which conditions presuppose a great increase in productive power, a high degree of its development. And, on the other hand, this development of productive forces (which itself implies the actual empirical existence of men in their world-historical, instead of local, being) is an absolutely necessary practical premise because without it want is merely made general, and with destitution the struggle for necessities and all the old filthy business would necessarily be reproduced;".

Engels: "Only the immense increase of the productive forces attained by modern industry has made it possible to distribute labour among all members of society without exception".

Right, they envisioned it being because of the tendency of the rate of profit to fall - ie capitalism being so maximally efficient that it turns in on itself - which is actually untrue (capitalism is more efficient than the alternatives but not maximally so; rates of profit ie margins are actually quite high in historical terms right now).

I'm not saying it proves communists correct. I'm saying the jobs the kumbaya-adjacent leftists envisioned themselves having might be the only source of meaning left in a post-singularity utopia.

I don't think that any kind of work is necessary for meaning. There are other sources of meaning, such as curiosity, philosophical exploration, and social connection.

The supposed connection between work and meaning is, I think, largely just humanity making the best (meaning) of a bad thing (work).

People of the ruling class in societies like Classical Athens would have been shocked by the idea that productive work is necessary for meaning. They thought of work as something for defeated people, slaves, inferiors. Yet they did not lack a sense of meaning in life. But AI is coming for some of their sources of meaning, too, such as the feeling of being on top compared to other Earthly beings.

There is meaning in a sense of pride and accomplishment, but as that infamous post suggests, that can be fully simulated.

'Work has meaning' is a coverup for simian greed. People who say they get meaning from their work really mean they enjoy the power it affords them as measured in wealth and status, and that it would sadden them to see their dreams torn to bits owing to a turning of Fortuna's wheel. That is at the heart of what I see in all these salarymen, men who boast about their portfolios and talk in big, drunken voices about their vaunted careers. The reddit normies. Underlying all their grown-up chatter is nothing more than the humblebragging of narcissists who thought they were winning the darwinian game.

If you have a device which allows you to read the minds of other men, by all means, share it with us so we can all use such a useful tool. But for my part at least, you are wrong. I derive meaning from my work because of the satisfaction of a job well done and making the world around me a tiny bit better, not because I'm a narcissist who secretly glories in being better than others.

What do you draw meaning from that cannot be sneeringly explained away in unflattering terms?

Nothing. I've evaluated my own goals, values, and notions of what I'd like to accomplish and realized they all came down to either direct egotistical buttressing or else buttressing for secondary layers. @SubstantialFrivolity above notes he gains satisfaction from making the world around him a tiny bit better, but that reeks of a secondary layer of buttressing, while even his talk of jobs well done sounds like he's glorifying perfectionist urges latent in certain 'engineering-minded' individuals rather than, necessarily, true results.

True results would reveal that, for instance, slop-programmers earn more not only in commercial terms but also in terms of social praise in many cases than perfectionists, so that external signals would seem to validate their proclivities to a greater degree. Many of the world's most successful people operate on a haphazard basis. The proclaimed principle of doing a job well done would seem to be entirely removed from objective measurements, and so its validity can only be attested by its own proponent, whose benefits in promoting social acclaim for urges he inherently possesses cannot escape notice.

Unflattering terms cannot be avoided because the underlying human animal is unflatteringly conceived. A species born of darwinian struggle, embedded in social structures encouraging egotistical self delusions and self-promotion of their own worth and validity, all the while chasing means to further darwinian success. The instincts for biological selflessness and charity, in comparison, are very little.

More comments

I suspect that the main sources of meaning will continue to be things like religion and family / childrearing / marriage that code as trad.

I'm less sanguine than you. The utopia AI will offer us does have advantages, the biggest being significantly healthier throughout a (neverending?) lifespan. And I, personally, am very excited for the knowledge and discoveries it will have about the universe.

But beyond that, the utopia won't fundamentally change that much. Even today, no one (at least in the USA) is starving for lack of food. The deprivations most people feel are positional and status-related: most people can get a house somewhere, but when people complain about not having access to housing, what they mean is they can't afford a brownstone in Park Slope.

Before AI, society could, at least in part, use merit and economic contribution as a way to allocate those inherently exclusionary goods; that's going away in the utopia. I'm lucky enough to have a almost two-decade-long career during a time where I could convert my intellectual capabilities into something that lets me acquire those resources, but that path won't exist for people like me. I don't have faith that whatever way society decides to allocate status resources will allocate them in a way I like (though I expect I personally am positioned in a place to have the level of exclusionary goods I want; it's more what I'm imagining for future generations).

That's following the optimistic utopia branch. The extinction branch is also quite likely (though I'd guess on the order of decades, not months or years) and much worse.

most people can get a house somewhere, but when people complain about not having access to housing, what they mean is they can't afford a brownstone in Park Slope.

One reason that matters, though, is because opportunity clusters. If you can't afford a brownstone in Park Slope (wherever that is), that's a problem because you then can't access the high-paying job/networking/cultural opportunities in Park Slope. There is a plausible downward spiral of

couldn't quite afford Park Slope -> moved to the boonies -> could only find low-paying local jobs -> could never even think of moving to Park Slope -> can't even afford the boonies any more

which doesn't quite hold in a utopian society.

Park Slope is a nice neighborhood in Brooklyn. A brownstone in Park Slope is roughly a Victorian in Noe Valley.

In the utopia, how does someone not in Park Slope get their brownstone in Park Slope?

I think the idea is that while some people value a brownstone in Park Slope for its own sake, some value it because it puts them near the job opportunities of NYC and otherwise wouldn't want one.

If the job opportunities of NYC don't matter anymore, that second group no longer wants to be in Park Slope, so while there probably still won't be enough brownstones to go around, there may be less of a gap than there currently is.

We're also seeing some significant declines in fertility rates, and I think more so in the demographics that want to be in Park Slope. So if that trend continues in utopia (which seems reasonable, material abundance if anything seems to get people less interested in children if we look at the modern era), over time there will be fewer people looking for them and the gap might shrink further.

If we get infinite life extension that might be a different matter, but I expect eternal youth will also lead to "eh, I'll get around to making kids later" for some very long definitions of later, and the city-seeking crowd also seems to have a problem with depression that might lead to disinterest in infinite lifespans.

Thanks @wraelk!

@ThenElection, this is what I was trying to say, but much better put. Currently, the area you live in affects the trajectory of your life: moving to... Idaho? (forgive the Brit please) doesn't just mean you have shelter in a less prestigious place, it means that your income and future life opportunities are going to shrink.

That means that the number of people who want a brownstone in Park Slope is currently

people with a strong attachment to having their shelter in Park Slope specifically + people who want a good future for themselves = a really big number

whereas in utopia it's just

people with a strong attachment to Park Slope = hopefully a smaller number

Idaho? (forgive the Brit please)

Idaho is actually very expensive, like many western rural states.

'Move to Ohio' is a reasonable answer to those who can't afford New York, though- the purchasing power will actually be higher, even given lower average salaries.

moving to... Idaho? (forgive the Brit please)

That works for what you're trying to say, Idaho is a super empty and rural place. 44th in population density.

This was something I was really hoping would happen with all the remote work from COVID, businesses would realize they could have employees from all over the states and people could move wherever made the most sense for them independent of their job, reducing a lot of the housing pressure in major cities and letting smaller towns build up more. Unfortunately, while there are still more remote jobs than there used to be, a lot of that dried up.

The only reason I'm not saying it's over is because I've said it before. It's been over for a while. And we've barely begun.

The good news, for mathematicians, is that their employability (in academia) did not hinge very strongly on tangible economic output. Nobody funds an algebraic geometer expecting a return on investment, which means nobody can defund one on the grounds that a datacenter in Texas now offers a better one.

The bad news? Everything else.

Just look at this fucking sweep. Just look at it. Hilbert's tenth over the rationals. Rational Hodge for CM abelian varieties. Integer multiplication below n log n, which I had mentally filed under "the floor, go home." A quasi-Riemann hypothesis thrown in like a free tote bag. Any one of these would have been the defining result of a human career, and they were released as a batch, in a GitHub repo, on a Tuesday, alongside hundreds of others.

And it wasn't brute force at absurd cost. OpenAI says the average result used roughly three hours of ChatGPT Pro-equivalent compute. I agree with your Pareto intuition, but the deeper point is that three hours is a price, and prices in this industry fall monotonically. Gwern spent 2020 arguing that neural nets would keep absorbing compute and sprouting new abilities, while noting the idea was so unpopular that it would only be accepted as a fait accompli. Well, here is the fait, and it is quite accompli. Next year the same theorem costs twenty minutes. The year after, it's a rounding error on someone's API bill.

Your own experiment deserves more attention than you gave it. As others have already done to good effect, your research methodology consisted of being a motivational poster ("you can do it!", "I believe in you") and you got a materially stronger result in half an hour. The binding constraint has shifted to elicitation and persistence.

Some might claim that human mathematicians are necessary to humanize and understand the proofs, or to come up with new frameworks and fields of research.

Maybe. I don't know. The strongest version of the argument is that mathematics is a conversation between humans about what humans find interesting, and a proof nobody understands is a tree falling in an empty forest. I'm sympathetic to that! I also notice that this position has been retreating to progressively smaller hills since the first computer-assisted proofs, and each hill gets defended with the same conviction as the last.

I do know that the current state of affairs is unlikely to last for more than six months anyway. Just get the models to come up with their own novel conjectures, or to attempt an overarching synthesis. I don't think turning abstruse Lean proofs into something human-readable will be particularly difficult, though human readability has long ceased to be a major concern for anyone. Note that OpenAI is shipping Lean formalizations for many of these proofs, so "trust me bro" is off the table; a type checker cares nothing for anyone's feelings. We're at the point where people are coping that the tsunami of new discoveries hinge on hither-to undiscovered bugs in Lean. You wish.

Every day, I feel the pace of progress accelerate, and then the practical ramifications show up in my life a few months later. I used to think of this as a lag. These days it feels more like the gap between seeing lightning and hearing thunder, and the gap keeps shrinking.

We've just had Scott Aaronson announce a new UT Austin course, CS395T: AI Alignment Theory, with Yudkowsky's AGI Ruin as the first assigned reading. To quote:

I vividly remember encountering Eliezer Yudkowsky and his Sequences 20 years ago. I remember thinking: even if these people talk and act like crazy cultists, still, let me bend over backwards to be epistemically virtuous, and entertain their ideas on their merits, as very few academics would. Even if, of course, I ultimately end up rejecting the ideas, on the simple ground that powerful AI is such an absurdly remote prospect that it's almost impossible to say anything useful about it today, outside the realm of speculative fiction.

For my failure to see what was coming, it seems like an appropriate punishment that I'm now, in 2026, effectively teaching a course on Yudkowsky Studies. And it's the most important course I can teach.

I'll give credit where it's due, but also: when the designated Reasonable Skeptic of the rationalist diaspora concedes the point in writing, the Overton window has relocated.

We've also just had an American company, Nolla Health, receive regulatory approval in Utah for AI to issue initial prescriptions. Yes, it's acne. Yes, the model can only choose from a short list of physician-approved topicals, and clinicians reportedly agreed with its recommendations in over 96% of cases. Everything starts as acne cream. Utah let Doctronic's AI handle prescription renewals back in January; nine months later, it's writing first-line scripts. You can extrapolate the line yourself. That's the harbinger of medicine's fall, as far as I'm concerned. I'm impressed we held the moat this long.

At least I've got a moat. The NHS is famously a slow ship to steer, and sclerotic at that. The future arrives unevenly distributed, and we can count on the service the last postcode on the delivery route. For once, I find that reassuring <3

Good career choice, past self_made_human. Turns out there are concrete benefits to taking theoretical but plausible concerns seriously, and preparing accordingly.

Integer multiplication below n log n, which I had mentally filed under "the floor, go home."

If I understand your comments below, this is actually false, you didn't have any expectations about the bounds for this problem. Presumably this claim was just AI generated.

The improvement in the paper goes to n(log n)^(1-(2^-182)), which while kind of neat is essentially the same as n log n (unless we're multiplying numbers on a computer larger than the universe). It's a toy result. Your amazements suggests you probably didn't even look at the paper.

In other words, I consider your thoughts on this totally untrustworthy. Please don't use AI to write your posts.

The improvement in the paper goes to n(log n)^(1-(2^-182)), which while kind of neat is essentially the same as n log n (unless we're multiplying numbers on a computer larger than the universe). It's a toy result. Your amazements suggests you probably didn't even look at the paper.

Ahem. Please look at this comment I had left a scant few minutes back.

https://www.themotte.org/post/3960/culture-war-roundup-for-the-week/485779?context=8#context

I am well aware that this is practically indistinguishable from O(n log n). I know how exponents work. I haven't checked yet, but a sufficiently ridiculous constant factor would make it even more impractical. That's been known to happen.

If I understand your comments below, this is actually false, you didn't have any expectations about the bounds for this problem. Presumably this claim was just AI generated.

I had read mathematicians, on Twitter, expressing surprise that we went below the previous SOTA by any margin, no matter how minuscule. The model rephrased my notice of secondhand surprise into a first person version.

That is such an innocuous, irrelevant change that I wouldn't have bothered to correct it for it's own sake. I hadn't even told the model to only wrap my references in links. It was entirely at liberty to add minor context. The only reason I point it out is because I respect @2rafa, and wanted to declare it as a concrete example of a phrase I hadn't hand-typed.

In other words, I consider your thoughts on this totally untrustworthy. Please don't use AI to write your posts.

Totally untrustworthy? What a massive overreaction. A totally untrustworthy person would deny the whole thing. Instead, I'm punished for admitting any use. You're lucky I think the price of honesty is acceptable.

If you write me off, that's your problem rather than mine. Particularly since I didn't use AI to "write" my post, I had already written a post, and I threw into it for minor improvements and explicit citations I didn't have the time to source. By my standards, the changes I've quoted from the original are positive or at least benign.

If for the longest time the best exponent was some nice number like an integer or rational with small denominator, and the algorithm reaching it relatively simple, then even the tiniest improvement is notable. That it never performs better in cases small enough to be encourted in the real world, is not what mathematicians care about. Were it so, Big O notation would not be as common, and ultrafinistism would be the majority position.

I don't think I agree with that perspective on Big O notation. It's common because generally people don't intuitively understand the difference between, say, quadratic and exponential growth. But of course in computing it's crucial, so people need to be taught it over and over so that they will understand it intuitively. It has a huge impact on how fast an algorithm performs in real world implementations.

The difference here is that the smallest example where log(n)^(1-2^-182) is noticeably different is if you were multiplying two numbers whose physical representation in bits is larger than the universe. It's totally meaningless.

The best exponent was some nice number because those are easier for humans to develop proofs for.

I see no reason why you couldn't replace the hype man with a cheap AI either. One LLM tries to find the proof. The other is instructed to hype it up. No human involved, other than in deciding what proof to look for next.

They do, I've heard of it being done for at least 2 years.

I consider that so obvious I didn't find it worth saying. Unless, for some stupid reason, the models learn that they're being prompted by print.ln instead of a human.

Given that reasoning effort is something that can be adjusted (usually by simply a privileged part of the prompt, rather than some fancy dial), it's moot.

when the designated Reasonable Skeptic of the rationalist diaspora

Who? Aaronson? The "took his family to Disneyland and aborted the trip after having a meltdown over seeing people not wearing masks" guy? He's many things, but nobody is designating him any combination of "reasonable" and "skeptic".

(By the way, concurring with @2rafa's suspicion below, after you have acquired your reputation as an unapologetic meat proxy, it does feel extra suspicious when you make such catchy irresistible logit-attracting utterances whose substance seems bizarre when you actually think about it.)

Some people have forgotten that I'm capable of sarcasm.

I'm more than aware of Aaronson's quirks. In some regards, he's a pathological quokka, and in other cases, he's had a rather overblown threat response and reacted out of proportion. See his posts about his fear of anti-semitism for the latter, or the fact that he inspired the other Scott to write "Radicalizing the Romanceless."

In the Rationalist sphere, he's the one who's always painted him as the reasonable person, not prone to theoretical flights of fancy.

Anyway, unapologetic? Not at all. I just think it's a tradeoff I'm willing to make, even if it annoys you.

the fact that he inspired the other Scott to write "Radicalizing the Romanceless."

pushes up glasses

Well actually, he inspired Scott (PBUH) to write “Untitled”, which came out a year or so after “Radicalizing the Romanceless”. Both absolute bangers, though.

I am ashamed of my error, and will commit sudoku out of repentance.

(That's what I get for not running a throw away comment through an LLM for fact checks. I'm more likely to 'hallucinate' than they are.)

Post the sudoku when done so we can confirm you are a man of honor and good repute.

commit sudoku

I lold.

git push sudoku

Have you heard about Poe's Law? Sarcasm doesn't work when it's equally plausible that you would say the same thing in earnest, and attaching a clichéd epithet like that without any deeper reasoning behind it is the sort of thing an LLM would do in earnest.

Anyway, unapologetic? Not at all. I just think it's a tradeoff I'm willing to make, even if it annoys you.

All's well. You're willing to annoy me with AI content, and I'm willing to annoy you by derailing your threads with speculations about AI content :)

All's well. You're willing to annoy me with AI content, and I'm willing to annoy you by derailing your threads with speculations about AI content :)

I consider that a reasonable compromise for all involved. It's the people who think they have the right to dictate how I spend my free time that I take umbrage with. I otherwise aim to be honest if anyone explicitly asks.

Sure, AI may be embarrassing humanity's proudest intellectual achievements. But there's a bright side. We can finally stop arguing about how AI is an overhyped bubble that's about to burst! Haven't even seen the term "stochastic parrot" in a while.

Another plus: AI can now compose its own Disney villain song. I'm not exactly sure where "writing a catchy diss track against people who think you're not intelligent" lies on the spectrum of intelligence, but it's probably a little ways past the Turing Test.

We aren't arguing about it because there's no point. At this point, both sides have tired of trying to convince each other, each still thinks they are right, and each still thinks the other is willfully ignoring evidence that disproves their POV. What's there to argue about any more? Life's too short.

We can finally stop arguing about how AI is an overhyped bubble that's about to burst!

I'm sure we can do that, but maybe we shouldn't.

There's been some pretty not-great signs for the financial state of the industry lately: Microsoft looking to wean off of Anthropic, Harvey and Thompson Reuters moving to their own models, the force majeure notice by Oracle in New Mexico, the failure of SB Energy, Holtec, and Aggreko to IPO (not to mention OpenAI and Anthropic!), Firmus missing its rental payment, and generally increased skepticism on the part of investors.

I think it's been a huge problem for thinking clearly about the situation that the financial concern regarding the AI industry has been latched onto by the worst AI skeptics, who tend to flatly deny the capabilities of the models.

Meanwhile on the flip side, I suspect there may have been a parallel problem that the people who are the biggest boosters of the technology are the ones most likely to be suffering from mild AI psychosis from talking with them all the time.

Thus, the AI debate has been, somehow, between "AIs suck and there is a massive bubble" and "AIs are the best thing since sliced bread and nuh-uh," which excludes two entire quadrants of possibility from the conversation, "AIs suck and will be profitable" and "AIs are good and there is a bubble."

This state of discourse creates epistemic closure on the topic in a way that I think is preventing a lot of people from assessing where, exactly, AI is sitting.

The word "bubble" is hurting more than it's helping here. It's beyond question that the money spent on training frontier models will produce an absolutely massive amount of value for the world. But it's still to be seen how much of that value can be captured by the labs themselves. This isn't really the same shaped as tulip mania or people paying absurd amounts of money for pets.com. If the labs fail economically(Which I think could only happen if scaling laws collapse roughly tomorrow) then the weights will still be around and still worth quite a bit to serve on the hardware which it will still be around and be worth it to keep running inference on. The labs could fail but they wouldn't be failing because AI is overhyped or a scam, just that running a research company in a competitive environment without state granted monopolies on the produce of your research is a brutal business to be in.

If the labs fail economically(Which I think could only happen if scaling laws collapse roughly tomorrow)

It seems pretty plausible that the frontier labs could fail simply because demand doesn't meet their projections. Anthropic, for instance, has something like $200 billion in commitments to Google and Amazon out to 2036 where, according to the terms of the deal, they pay regardless of demand.

There's already signs that demand may soften for Anthropic specifically (the OpenRouter trendline is switching away from Anthropic towards open source models iirc, Microsoft is looking to cut their spend with them, ditto (we can infer) Harvey and Thompson Reuters, Astra is apparently universally beloved by coders). If people start switching from Anthropic (and it doesn't take many: keep in mind that 80% of Anthropic's revenue (like OpenAI's) comes from 1% of their users, and 2 of Anthropic's customers generated about 25% of their 2025 revenue), they can either

  1. produce a superior product
  2. cut spend
  3. raise prices
  4. die

Except they can't cut a lot of their spend (as per above), if they raise prices, OpenAI eats them anyway, and perhaps Astra or OpenAI engineers are good enough that they are simply locked out of producing a superior product. That leaves #4, die.

Maybe this seems good for OpenAI, except that if the timing is bad it sends the market into an AI panic and could spoil their IPO, and OpenAI is also on the hook for a bunch of infrastructure bills (although they may have structured them more flexibly, I'm not sure offhand). And of course there are a lot of other things that could also ruin an IPO: another pandemic, major war breaking out in Europe or the Pacific or Middle East, political unrest: pretty much any little thing that goes wrong and tightens the belt could crack up the revenue stream, OpenAI is not a profitable company, and the open-source models are nipping at its heels.

then the weights will still be around and still worth quite a bit to serve on the hardware which it will still be around and be worth it to keep running inference on.

Yes, I agree (and have said before) that "AI is not going anywhere." I agree this isn't the tulip mania. But we're in a subsidized era of AI right now, and it may look very different once that subsidy ends.

they wouldn't be failing because AI is overhyped

Perhaps the frontier labs fail for some other reason (or don't fail at all) but I think it's pretty fair to say that AI has been overhyped. OpenAI said it was going to spend $1.4 trillion on infrastructure by 2030. That's hyping. Then they slashed their public infrastructure commitments by more than half, to $600 billion (because their CFO was worried that they were overhyping), although they've brought the number back up since to $750 billion.

If they actually revise their numbers back up to $1.4 trillion in 2030 and meet that infrastructure goal, feel free to ping me and I will agree that this was a bad example. Then I will point you to the 2027 Project where it postulates that the robots would kill us all by now, as my fallback example.

Then I will point you to the 2027 Project where it postulates that the robots would kill us all by now, as my fallback example.

AI 2027 said superintellgent AI would be created in 2027, not that it would kill us all that year or even that it would be publicly known. In the scenario, everything appears to be fine until the AI suddenly kills us all in 2030.

Yes, and I am referring to a conversation I may have in 2030 :)

It seems pretty plausible that the frontier labs could fail simply because demand doesn't meet their projections...

This is basically describing what I mean by if scaling stops tomorrow. It can certainly be argued but it's my contention that if things keep scaling like this and all else holds steady the labs could just start directly eating segments of the economy. The "application layer" guys seem to expect the labs to just lay down and let them pay commodity prices to stand up a thin wrapper. But at a certain point anthropic can just instruct opus 7 "go ahead and stand up any product that would be profitable to run by a wrapper company, make no mistakes". That's a bit of a tongue in cheek example but what is stopping anthropic with a sufficiently advanced unreleased model from just becoming a hedge fund and doing to the financial industry what they're doing to mathematics? They have a wetlab, there's a lot of money in drug discovery.

I'm not making a strong claim here, I don't know what even the near future will bring. I don't think the labs failing is even all that unlikely. But I also don't think they'll have trouble finding investors money if they ask for it and it's not obvious that scaling laws have failed. I've personally set aside like a quarter of my net worth to be ready to invest in ant and oai at IPO.

But we're in a subsidized era of AI right now, and it may look very different once that subsidy ends.

I'll poke at the term subsidy just like the term bubble. Research is being subsidized, the inference is not. And it's hard to tell the degree of the research subsidy.

If they actually revise their numbers back up to $1.4 trillion in 2030 and meet that infrastructure goal, feel free to ping me and I will agree that this was a bad example. Then I will point you to the 2027 Project where it postulates that the robots would kill us all by now, as my fallback example.

If you've been keeping score the 2027 project is doing much better than its detractors and pretty damn good in absolute terms.

That's a bit of a tongue in cheek example but what is stopping anthropic with a sufficiently advanced unreleased model from just becoming a hedge fund and doing to the financial industry what they're doing to mathematics?

Why haven't they done this already?

At least part of the answer to this question is probably that models often perform worse in real-world applications than their evaluation benchmarks would suggest. Thus, while the models are improving and becoming more useful for real world work, Anthropic can't just throw them at a problem and say "make no mistakes" without staffing up humans to oversee them.

Even in coding, which they have been optimized for, the data that I've seen indicates that trust by developers has fallen as adoption has risen, and that as LLMs have increasingly been adopted in coding, the amount of insecure code has risen substantially while deployments have actually slightly dropped.

Just as a personal example, the other day I asked Sonnet 5 to implement a design and it decided to simplify the product on its own recognizance. For me, it's a little annoying. But obviously at scale for a real-world product that's disastrous.

This is not an "LLMs are useless" post (I'm very very interested to see what comes out of the wetlab!), but it is a "LLMs are not ready to go to the moon on their own" post. It wouldn't surprise me if the primes try something like this, and they definitely might succeed. But it wouldn't be a "we can just let the AI do everything" sort of thing.

I also don't think they'll have trouble finding investors money if they ask for it and it's not obvious that scaling laws have failed

It seems like Anthropic and OpenAI think they will have at least momentary trouble, hence the delayed IPOs, and I noted the trouble the hardware suppliers were having upstream. If investors are becoming skeptical of datacenter investment, there's reason to think they are going to be skeptical of the primes. If they are skeptical of datacenter investment but not of the primes, they are just shooting themselves in the foot very dramatically, since the primes need the datacenters to scale. I suppose it's possible that they are all-in on orbital ones after the Space-X IPO, though.

This isn't a hard prediction here (nor do I think Anthropic and OpenAI delaying their IPOs is necessarily a bad idea or indicates terminal distress) but I think it's important to keep your eyes open. Particularly if you're intending to throw your net worth at them.

Research is being subsidized, the inference is not.

I'm getting inference for free, it's being subsidized.

If you've been keeping score the 2027 project is doing much better than its detractors and pretty damn good in absolute terms.

The area where I think it did the worst was in the economic projection of the stock market as a whole, which I think is pretty relevant to the fundamental question here.

Why haven't they done this already?

As I said, if the scaling stops tomorrow then maybe the models simply are not good enough to do this autonomously and other human labor is required to guide them along. If it continues at this rate for another year or two things change dramatically.

Even in coding, which they have been optimized for, the data that I've seen indicates that trust by developers has fallen as adoption has risen, and that as LLMs have increasingly been adopted in coding, the amount of insecure code has risen substantially while deployments have actually slightly dropped.

Just as a personal example, the other day I asked Sonnet 5 to implement a design and it decided to simplify the product on its own recognizance. For me, it's a little annoying. But obviously at scale for a real-world product that's disastrous.

I don't have time right now to rehash this debate again. using sonnet instead of opus, not having a proper harness and planning/review cycle. I've seen the thing go, I've put out projects in weeks that would have taken months. If you don't believe it has the juice then so be it.

I'm getting inference for free, it's being subsidized.

You're getting pennies of sonnet for free, I'm burning through $3k+ in tokens a month. Although the psychology here is confusing to me. You'd have to consider your time basically worthless to avoid playing $20/month for some opus 5.5 usage instead of sonnet 5.

More comments

There are still holdouts, but at this point I regard them with more psychiatric curiosity and pity than I do as entities to reason with.

As an analogy: people who ignore the insurance on their beach side property shooting up vs those ignoring an active hurricane evacuation warning. Both could be doing better. One side is past my ability to help.

Even Gary Marcus has shifted to claiming that the modern models are not "pure LLMs".

Even Gary Marcus has shifted to claiming that the modern models are not "pure LLMs".

Modern models use tool calls, so this seems straightforwardly correct, at least in a certain technical sense.

The best kind of correct, perhaps, but that interpretation really takes the wind out of the sails of his conclusions. He wanted to "start easy: 5 digit multiplication problems". GPT4 had a 6.67% success rate, ha ha! But what do you think Gary Marcus's success rate would be? Modern LLMs might still need chain-of-thought to answer such a question reliably without an external tool, but I'd bet that Gary needs the same and I'd even bet he needs external pencil-and-paper to keep track of the intermediate steps. He says "The LLM never induces such an algorithm", but when I ask for a tool-less multiplication I get a reasonable digit-by-digit algorithm executed in the output text, which is pretty weird if "the LLM based system is generalizing by similarity, doing better on cases that are in or near the training set, never, ever getting to a complete, abstract, reliable representation of what multiplication is."

This guy still gets quoted as an "expert".

I don't think using tool calls reflects poorly on LLMs are products at all - if anything it enhances their value. This isn't an "LLMs suck" post.

However a lot of people view intelligence as a "unified" property. I've been pushing back on that idea on here for a while because I doubt that will be correct; my guess (and so far I've been proven correct; see LLMs getting worse at writing as they specialize for coding) is actually that there are benefits to intelligence specialization. This doesn't mean you cannot wrap those specialized compartments together into a unified process, of course. The human brain, for instance, is "unified" but it has specialized regions that seem to be optimized for specific tasks; same with your computer. Arguably an LLM using tool calls and the like is doing something similar.

I think it matters because 1. I'm petty and like being right, and 2. I think it's worth thinking clearly about these things.

For quite some time, frontier models have generally been Mixture of Experts models (or, possibly, something even more proprietary and advanced - I'm not an insider). I think this fits with your intuition.

Thanks for flagging that link! Yes, I do think that fits with my intuition.

I implore you not to sane-wash Gary Marcus. Modern models like Astra are more than capable of feats that he swore up and down couldn't be done by LLMs, even if they could be run as bare as possible, without tool calls.

The most symbolic component in OpenAI's release is the Lean checker, and its job is to grade the LLM's work.

I implore you not to sane-wash Gary Marcus.

Gary Marcus could be literally insane and it would still not be right to mock him for saying something that is true.

The most symbolic component in OpenAI's release is the Lean checker, and its job is to grade the LLM's work.

Do you know for sure that no tools were used? My understanding is that the reasoning traces released for the most recent models were summarized; however, we know that the agents that solved Napier-Stokes had tool access. I'm not sure I'd assume something different was done here, but maybe they said so somewhere.

I'm not sure I'd assume something different was done here

I'd be fascinated to be corrected if the facts disagree, but just based on priors it would honestly be irresponsible to do something different here. LLM output is indispensable for problems that aren't well-posed or don't have a straightforward algorithm to find a solution, but inefficient and risky for problems that are and do.

When I was a first-year grad student, a TA gently pointed out at the bottom of half a page of my handwritten symbolic calculus simplifications that my doing that much by hand was (although correct in this case!) both risky and a waste of time, and that we all had Mathematica/Maple/etc licenses we were allowed and encouraged to use. I'd guess an LLM might prefer SymPy in the same circumstances, but the general principle is surprisingly just as true.

I'd be a little less surprised if the available tools other than Lean weren't at all useful for most of these problems, but surely they were useful for many of them. There are some really good open source tools out there these days (just looking around: SymPy has had a differential geometry module for a decade!) and OpenAI's models surely know about effectively all of the best ones.

Yeah I agree, I assume you'd want to give them the best tools, if they might be relevant at all.

Stock market can still be a bubble, even if the underlying tech is amazing. Research has cost an absurd amount of money, and it is unclear whether any of these companies will recoup their investments. Open-weight models are increasingly powerful and there is a lot of competition in the field. No one has a monopoly at the moment, and it doesn't look like one is forming either. Instead, everyone is trying to release the best possible models at the lowest possible cost.

Dang it, I spoke too soon! :) But fair enough; I was being a bit snarky. I do think a lot of the "AI is a bubble!" folks were arguing that AI would never be valuable, not that it's incredibly valuable but hard to capitalize on. Hopefully the former argument, at least, has been put to bed.

My theory is that a lot of the anti-AI crowd has simply been motivated by fear and ideology. They latch onto any argument that can be used to stop investment into these things, either because they distrust the tech bros, or because they view continued improvement as an existential threat to their livelihoods.

Since neither of those factors have changed and the existential threat has only worsened, I think the anti-AI crowd will only become louder in the coming months. If they can no longer credibly point to the abilities of the robots, they will just find a different angle. Water consumption, energy supply, or profitability are all arguments that are still being spread around.

You don't stop these people by being right. You stop them by convincing them they will benefit from the technology, or at least that it won't harm them or a cause they care about.

If only it were so simple!

Alas almost all technological hype cycles are for technology that works and is transformative. Trains, canals, optic fiber, you name it.

There is most definitely a world where the massive investments in AI don't meet enough demand quickly enough and Anthropic goes the way of Cisco.

Be honest. How much of the above did you actually write? I’m reserving judgment. I’m just curious.

95% of it? I typed it out during my lunch break (I had lost my appetite), and then the only extent of LLM involvement was to dig up and then wrap in Markdown the relevant links.

Which are all real, anyway. I just didn't have the time to bother.

Hmm. Let me check closely:

Several throwaway statements had a sentence or two of context added. For example, "The binding constraint has shifted to elicitation and persistence".

If you want to be technical, it inserted the full title of Aaronson's post, or surfaced it in the first place. Initially, I just happened to come across it secondhand on Twitter and had pasted the quote in bare. Uh, now I see that it claims that I had strong opinions on the algorithmic floor of integer multiplication. I did not, beyond being vaguely aware that there was a better option than a naive n^2 based off an article I think I read on Quanta. I couldn't have told you off the top of my head that the previous SOTA was O(n log n).

Uh, now I see that it claims that I had strong opinions on the algorithmic floor of integer multiplication. I did not, beyond being vaguely aware that there was a better option than a naive n^2 based off an article I think I read on Quanta. I couldn't have told you off the top of my head that the previous SOTA was O(n log n).

So the LLM inserted something dumb that you have no personal knowledge of? I'd be embarrassed, myself, but you do you I guess.

Oh really? What's so dumb about it? When I checked X earlier, there were plenty of mathematicians expressing surprise at getting any lower than O (n log n). I would have been surprised too, if I had remembered the exact threshold.

In my very scholarly and medical opinion, I think you're suffering from Human to AI writing transmorigification. I thought your post was also at the very least highly AI modified but it may well be that AI has addled your brain so much that now you just naturally write that way.

A curious specimen indeed.

Count, I always give your opinions as much weight as they're due. Just don't ask me how much right now.

Uh, now I see that it claims that I had strong opinions on the algorithmic floor of integer multiplication. I did not, beyond being vaguely aware that there was a better option than a naive n^2 based off an article I think I read on Quanta. I couldn't have told you off the top of my head that the previous SOTA was O(n log n).

Aw, dang, I was actually pretty impressed you had such specific CS knowledge!

Hey man, I read HN religiously ever day, and I once solved a LC medium in 2021.

Jokes aside, I do like maths and CS. The former all the more so when I'm not forced to for the sake of exams. And I try to learn what I can as the mood takes me. I even tried to learn Lambda calculus once, even though it provided literally no practical utility.

Someone on hacker news had for the last 20 years a specific problem in graph theory (Barnette's conjecture) and tried to understand the proof (and tried to translate it for non-math people) and was astonished that it used multiple esoteric "symbolic tricks borrowed from theoretical physics". He used Anthropics Fable to understand the OpenAI proof and first thought Fable was hallucinating techno-babble:

https://news.ycombinator.com/item?id=49988654

Forgive me for posting LLM output but I find this darkly hilarious:

"it's a matrix-tree cancellation wearing Kasteleyn's planar signs, run as a Witten index over Penrose-lineage states, evaluated as a fugacity-zero loop gas in an infinitesimal magnetic field — and the reason it reads like physics is that every one of those tools was built for partition functions"

I thought this was pure slop when I read it but there are some clear analogues in these other areas of physics, really neat computational tricks, and a very interesting paper by Penrose calculating Tait colorings I never knew about previously

So at least this proof (and at least to some guy on hn) this proof were not from the LLM digesting prior art papers on this problem or general graph theory, but from "the entire corpus of neat computational tricks that physicists derived to make their equations spit out something other than zero or infinity". Which makes it very unlikely a human could have thought of it, as few graph theorists are familiar with infinitesimal magnetic fields (and vice versa).

This eliminates valuable intellectual space for random people in non-math fields to reinvent Calculus like Dr. Tai did in 1994.

(The back and forth, here, was entirely me saying "you can do it!", "keep at it, I believe in you", and "you've got this, finish it!")

Can you expand a bit on how this conversation went for you? You put the paper in, turned Pro mode on, and then gave the verbal equivalent of a 'continue' when the model turn came to an end? As a researcher, I'm curious to know how others are working.

Here, yes, just that. Used higher diction than my comment, but you could put it in a loop.

For my more usual interests, it's more involved; oftentimes the model will go off on a tangent or latch onto something I don't care about, and I need to tell it to stop going in a certain direction. There's also specifically generating artifacts that're useful for later sessions to build on (OAI has a library tool that does something similar now, but I don't love the implementation.) But for this math paper extension, it was just the dumbest elicitation repeatedly.

Do you think it's some kind of token-checkpoint, like it goes so far and not further because it has a built in "Don't expend too many tokens without more prompting" feature? And then you say, "Keep going" and so it does.

Thanks, that's interesting.

For my more usual interests, it's more involved; oftentimes the model will go off on a tangent or latch onto something I don't care about, and I need to tell it to stop going in a certain direction.

This is more what I'm used to. I wondered if I was crippling it by stifling it, and would get better results from direct elicitation.

I guess worth trying both depending on the problem :)

I am pretty torn about it; originally I had a fairly sophisticated research ontology. But I've gradually simplified it to just files that can be discovered in a hierarchy of progressive disclosure, and trusting the agent is smart enough to find what it needs when it needs it.

I get that. I have moments of thinking, "is the AI a useful tool implementing my ideas, or am I increasingly a useful tool allowing the AI to work around the bullshit that human-designed systems put in its way?".

It's why I ask the kind of questions I asked you: to see if I'm slowing things down by insisting on participation to justify my existence, rather than just solving the problem. I think still largely no, but it's good to be aware.

I'd argue math matters -- there's a reason that people like me who need to use our knuckles for times tables still sometimes know Dijkstra -- but that's probably not going to help math PhDs choosing between a barrel of bourbon or a different type.

In the short term, this might boost math as a job, since now there's literally hundreds of lifetimes of proofs needing deep examination and translation being dumped on a weekday night... but even in the best-case, that's not what people signed up for. And worst case, the machines might be better at it than the humans, and if they aren't yet, they probably will be soon.

In theory, math PhDs are well-educated extremely smart people who should adapt best and most easily to a career change, but again, now what people signed up for, the PhD isn't cheap, and a career change to a different field that might just become the next benchmark isn't encouraging. Worst-case, this seems up there with 'moving to a new culture' or mega-dosing on THC that pegs the schizophrenia risk.

Notably, ML related results are nearly entirely absent from this batch of proofs; maybe a half dozen touch on it, distantly.

Yeah, and both anti-distillation and anti-LLM-training-assistance techniques are already known and disclosed. And given the 3SUM problem, if I'm understanding it correctly, I would be very surprised if there's zero here.

Related really fun thought: previous LLMs have been very prone to 'leaking' out to the public net as they do more advanced work or get stuck on genuinely impossible problems. Here, either OAI has really stepped up their sandboxing game, the leaked outcomes have become internalized universally without public fanfare, my search-fu has been flopping... or a lot of this stuff doesn't count as 'advanced work' or 'sticking'.

Also in my experience with Math PHDs a lot of them aren't really wired for the general workforce even if they've got unbelievable processing power. My high school ex's dad is a top tier Astrophysicist and has won a major prize for Astronomy but I wouldn't really trust him to do any job that requires particular lateral thinking and he's relatively sociable and flexible by the standards of his field. There is a certain filtering, especially these days, where only the most stubborn and pure-mathy are the ones who decide to pursue Academia (especially if they don't have good diversity points) in the first place and the ones with more lateral abilities drop off the curve when they realize how insanely difficult a slog it is to get consistent paychecks in Pure Math.

the ones with more lateral abilities drop off the curve when they realize how insanely difficult a slog it is to get consistent paychecks in Pure Math.

Even in applied math there's the same effect. The people with the most focused talent and interest stay in academia, in a tournament where only a fraction of entrants can win tenure. (How many grad students does the average tenured professor advise through a PhD? In a field that isn't expanding like mad, the reciprocal of that is roughly your odds of getting tenure.) Everybody else goes into industry or other research labs, trading respect for some combination of extra salary and lower stress.

ex's dad is a top tier Astrophysicist and has won a major prize for Astrology

This seems entirely possible (Isaac Newton was certainly very much into alchemy, for example), but it seems more likely you meant astronomy.

And Brian Josephson took up research into psychic phenomena after he won the Nobel for his work on superconductivity. The PR required careful management when he was the only living Nobel laureate in the Cavendish.

I feel like people being shocked that the sort of person with the single-minded drive, curiosity and intellectual firepower to get a Nobel ends up wandering into other esoteric stuff after their big win. Especially when the big win means they've got tenure and budget to just do whatever they feel like exploring

Always forget which is which, oops. Fixed.

I always assumed pure math was extremely easy to get academia jobs. Even in engineering or b-school math profs tended to be foreign born with awful English schools that resulted in making any college level math class a teach yourself affair.

Being that math departments seemed stuffed with foreigners who can’t speak English I assumed that there were zero domestic demand to be math professors.

Never actually tried myself but I've got a few white male friends with Maths PHDs who've said vast majority of job openings are looking for people with more diversity points (and foreign born counts towards that) and the compensation is low enough it only really suits visa-desperates or ultra autists.

The USA's Largest Employer of Mathematicians is not really academic, and almost certainly doesn't hire foreign nationals. I genuinely have no clue what their demographics look like, though.

Mathematicians are big mad. View the relevant subreddit on our progenitor site.

I did. Most of the top comments do not have an angry tone. It's more like, a mix of resignation and excitement. With some mad mixed into it.

Anyway and more importantly, it might be time to start thinking harder about the key question of our time: is it over for humancels?

Not any time soon, I think, but I tell you this: social positions in which being human is an inherent part of the value proposition are starting to seem like much more appealing things to devote one's time to than STEMing it up.

I'm thinking positions like prostitute, politician, athlete, stage actor, bartender, socialite, rock star, cult leader, man who makes money by befriending rich people, or man who makes money for taking the fall when an organization gets in trouble.

If y'all got any other ideas, share 'em. Cause it's time to strategize. Unless you're rich already, in which case: all good. Unless Skynet scenario.

I hope the AI-boosted productivity gains soon outstrip the job losses and we get the utopian diseases-cured Star Trek scenario instead of the Great Depression or Skynet ones.

From "The Power of Intelligence" by Eliezer Yudkowsky:

I have observed that someone’s flinch-reaction to “intelligence”—the thought that crosses their mind in the first half-second after they hear the word “intelligence”—often determines their flinch-reaction to the notion of an intelligence explosion. Often they look up the keyword “intelligence” and retrieve the concept booksmarts—a mental image of the Grand Master chess player who can’t get a date, or a college professor who can’t survive outside academia.

“It takes more than intelligence to succeed professionally,” people say, as if charisma resided in the kidneys, rather than the brain. “Intelligence is no match for a gun,” they say, as if guns had grown on trees. “Where will an Artificial Intelligence get money?” they ask, as if the first Homo sapiens had found dollar bills fluttering down from the sky, and used them at convenience stores already in the forest. The human species was not born into a market economy. Bees won’t sell you honey if you offer them an electronic funds transfer. The human species imagined money into existence, and it exists—for us, not mice or wasps—because we go on believing in it.

Which is to say, no, soft social skills are not going to become some last redoubt of human economic value. AI is perfectly capable of sending e-mails, talking to you over the phone, doing sales, writing a story, drawing a painting, acting in movies, composing a song, working as a therapist, dating you, becoming a pop idol†, etc.

AI can do anything a remote worker can do. And it's only a matter of time until drones and robots become good enough that human bodies are unnecessary as well.

We have to deal with the fact that the human era of producing value is coming to an end, probably in the next few years, certainly in the next couple of decades.

If our AIs are not aligned (they aren't), we all die. If our institutions are not aligned (they aren't), maybe a few elites survive but everyone else lives or dies according to their whims.

†Hatsune Miku was a vocaloid, not an LLM, but she is a proof of concept that you don't need to be a physical person to achieve stardom.

It is worth remembering that value is assigned by humans. A mass-produced ceramic cup and a hand-made one may look the same and be made from similar materials. But the one produced by hand will sell for more in the right situation, because some humans believe that the work put into it increases the value. They will imagine all sorts of ways that the hand-made cup is better, or maybe they simply want to preserve the skills required to make it. Regardless, the increased value is completely intagible and only exists within the mind of the buyer.

It is real regardless.

We have to deal with the fact that the human era of producing value is coming to an end,

Value is really an agglomeration of scarcity and social elements. Even today, it's possible to see a trend of concrete material wealth losing "value" to a human element. People often are willing to pay extra for functionally equivalent handmade products, farmers markets, and original art pieces. Food in terms of calories is valueless in a way our ancestors might not be able to comprehend, but there remains expensive food.

It seems quite possible that post-material-scarcity doesn't necessarily imply that human physical and emotional labor is zero-value.

If y'all got any other ideas, share 'em. Cause it's time to strategize. Unless you're rich already, in which case: all good. Unless Skynet scenario.

Or a bunch of scenarios short of that where the money simply becomes taken.

Like others have alluded to, these are already starting to be replaced. "Influencers", a job which doesn't even make sense to me to begin with, are now entirely generatable. The other day, I was on a drummer Facebook group, and someone posted about "see this photo featuring this cute twenty-something drummer chick? She doesn't actually exist."

If y'all got any other ideas, share 'em. Cause it's time to strategize. Unless you're rich already, in which case: all good. Unless Skynet scenario.

I'm lucky enough to have had a career in STEM, so I'm rich already, but I'm not confident that's going to "matter" in the long run. Where the long run might be, uh, a few years. In an economy where all real work is done by AI, are we even going to honour the increasingly-fake number that is the size of your bank account? In an AI world, does it make sense for me to be able to commandeer 1000x the resources of my never-working friends?

You might say that this would cause upheaval, but there are an awful lot more poor people who'd be in favour of this upheaval than rich people who'd be against it.

In an AI world, does it make sense for me to be able to commandeer 1000x the resources of my never-working friends?

No, and it's magical thinking based on a heuristic to think otherwise. People really need to start seeing past this. Your stuff isn't your stuff. You're just renting it. Pray the deal isn't altered.

In an economy where all real work is done by AI, are we even going to honour the increasingly-fake number that is the size of your bank account?

The question of who owns the miraculous bounty of industry has occupied organisms long before man was even around, few things are as lindy. It didn't go away when we invented machines, it's not gonna go away this time.

Even in the scenario where productivity gains are extremely high, somebody or something will still be fighting over who gets to turn which planets into computronium.

More plausibly, the secular run of the educated bourgeoisie is coming to an end, and a new aristocracy is about to rise. People always underestimate how much time these things take though.

You might say that this would cause upheaval, but there are an awful lot more poor people who'd be in favour of this upheaval than rich people who'd be against it.

Middle income families likely disagree, it is a dice roll not in their flavor, they can still go with their snp500 index fund and command the poor people, then you got the historical interesting norm where most successful revolutions were started by nobles or individual with higher wealth, they certainly aren't lossing in the current capitalism system

where most successful revolutions were started by nobles or individual with higher wealth, they certainly aren't lossing in the current capitalism system

A lot of the big successful revolutions were started off by the downwardly mobile and/or somewhat relatively-unsuccessful children of nobles, merchants and landlords. They had the education and the resources to dedicate themselves beyond subsistence, but also the motivation to take shots at the established order.

Suddenly the world has a lot of those.

Education I agree, but I disagree on resources. If they have the required resources to take a shot, isn't it better to just invest in a diversified index fund like SPY or VOO, compare to shooting the shot and gamble?

Well most historical communist revolutions predate access to index funds, and 'My rich parents/family paid for my education but I'm one of many children and I'm going to have to wait for 40-50 years old to get my inheritance isn't uncommon'

I agree with you on this, my comment on index fund is focus on modern time, where normal rich people rarely have more than 3-4 children, and index fund is a thing.

My disagreement is that given the 2026 circumstances, it is a lot harder to convince these people to break the current system

I think there's plenty of people in the spot where their parents are solid upper-middle/lower-upper class in 2026 and they're despairing of getting access to a meaningful inheritance in any time period that'll let them meaningfully change their own life circumstances.

If you're the only child of a Doctor or two (in their 30s at the time) in the mid 90s, most of their say 5-10M net worth is tied up between retirement accounts and their PPOR and they're from a culture of self-reliance then you're likely going to have had access to a great education but you're not going to be able to meaningfully touch an inheritance until your 50s or 60s and you're dealing with the employment market of today...

And I'm sure the above describes a lot of Mottizens and the current dissatisifed, especially those who took that great education and turned it into something other than skills or connections that the market values.

More comments

stage actor, bartender, socialite, rock star, cult leader,

I think all of these can be LLMified as well. People already talk to LLMs like they are their therapists, extending this to bartending and cult leadership is trivial. Converting K-Pop idols into digital format is also plausible. Imagine being a penpal with your favorite BTY member! Who cares if he's a machine if he replies to your messages and likes your memes!

Man, I guess I won't see any more job listings for "OnlyFans message responder" on Indeed anymore...

"It's not just a masturbation video, it's an exploration of female sensuality."

This is what I’ve been thinking. Except most places today seem less human than ever, and most people (maybe just around me, in STEM, and I’m worse than all of them) are bad at acting “human”.

So I think there’s still exclusive opportunity for nerds, until society is fixed so people can behave naturally, or reaches the point of no return. Whether that be improving the AI, or combining their nerd thinking with tacit human thinking that AI hasn’t learned to do what it can’t.

I'm thinking positions like (...) politician

I'm rooting for the Deus Ex ending where we replace these too.

Pros: No more politicians.

Cons: You've been mentally merged with all other people and the machine intelligence.

STEM is over, IMO; even more broadly, abstract thought will soon be economically valueless, including the humanities (which can have meaningful rigor to them!)

I would say that, to the extent labor income exists in the future, it's the physical aspect that matters. Prostitute, yes; OnlyFans star, no, you'll be competing with millions of AI bots who will outcompete you on every dimension.

Hoping I'll have enough to retire by the time that transition is complete.

you'll be competing with millions of AI bots who will outcompete you on every dimension.

Even in a theoretical AI distopia/ Star Trek tier utopia people still picked those grapes up, they made the wine. The starfleet cadets drew up new plans for new and exciting shuttlecraft, they had AI too. The human element will not go away unless we are all wiped out, driven by pure preference for the human connection alone.

The mechanical turk jobs will all disappear, but I shed no tears for those kinds of sweatshops.

Star Trek is fiction.

Reasoning from fiction about the real world is a fallacy (imaginary map vs real territory?), and science fiction says more about the time it was created as the depicted future. Old scifi imagined any household chores done by robots, but they couldn't imagine that the washing machine freed women from being housewives and the coming waves of feminism. In 60s StarTrek they couldn't get female commanding officers as women in the test viewership hated the bossy actress, in 90s Star Trek there were no gay people (not to mention transgender), and Picard was bald. In a real future with Star Trek like tech everyone will have fantastic hair and some otherkin/wolfkin will use transporter tech to degenerate themselves to furry-hybrids. And the holodeck itself is a can of worms...

Star Trek tier utopia people still picked those grapes up, they made the wine. The starfleet cadets drew up new plans for new and exciting shuttlecraft, they had AI too.

Yeah, but Star Trek is fiction, and the things you mention are why I believe that the only way it passes for progressive was because of the time it was made. People in ST are shown to posses a fanatical devotion to preserving the traditional human way of life. It's all reactionary as hell.

a fanatical devotion to preserving the traditional human way of life

Which, in my opinion, is the only way humanity could ever survive the age of artificial general intelligence.

I think it will take a while, particularly for large bureaucratic entities with entrenched market positions. Even in software engineering, a rather ridiculous amount of time is spent on coordination problems, figuring out what customers want, compliance and legal, flip-flopping management... If that overhead was previously say, 50%, even if LLMs can cut the dev time to zero, the maximum speedup is still only 2x. It's effectively an organizational version of Amdahl's law.

Organizational frictions are real. A year ago, I'd guess I spent 60% of my time writing code, the rest spent doing organizational bullshit. Now it's something like 95% organizational bullshit, and 5% prompting and reviewing LLM-generated code.

This seems temporary to me, though. Once everyone finishes shifting to LLM coding, organizations (mega corps or startups) that minimize the organizational aspects will outcompete those that don't.

The organisational "bullshit" is what maintains a causal connection between what the end user wants and what your code does. The big idea of the Torvalds/Raymond side of the open source movement is that when hackers work for themselves or each other they can cut 90% of it out. That gets you Linux. Cutting 100% of it out gets you GNU Hurd, so be careful.

The big idea of the Torvalds/Raymond side of the open source movement is that when hackers work for themselves or each other they can cut 90% of it out. That gets you Linux.

And we all know how good Linux is on desktop or at being user friendly...

Unironically Ubuntu desktop is very good these days and very user friendly. More user friendly than Windows actually, which today makes me want to tear my hair out with the ads it forces onto you and how it won't respect your registry choices as well as three dozen other serious problems.

That's a bit optimistic about the organizational stuff in a large org. E.g. make a Gantt chart for reporting to higher ups; convert your team's tickets into a new particular spreadsheet structure; write a design proposal for a system that's already done so it can be used for promo. They aren't entirely valueless because they give organizational legibility into otherwise opaque processes for decision makers, but they're not the pleasant part of software engineering.

Agreed. I will correct my comment to "The minority of the organisational bullshit that isn't pure waste is what maintains a causal connection between what the end user wants and what your code does."

Good software management (actually good management anywhere) is partly about protecting your team from the bits of organisational bullshit that don't align them - or protecting the whole organisation by keeping them out altogether if you are the founder. And it is hard because throwing the baby out with the bathwater produces a misaligned organisation.

Physical aspect or legal aspect, I figure. By legal aspect I mean that for a while to come laws will probably mandate that certain social positions must be filled by humans. No AI politicians yet, although many might soon become meat proxies of AI instead of being meat proxies of teams of human staffers.

Maybe there are other aspects as well.

In STEM, it's probably mostly over for pure programmers. There's still going to be room for being a person who understands what customers want, coordinates the various aspects of the business, tracks down humans with domain knowledge to pump their knowledge out of them, and prompts the AI. But AI will probably cut more and more into that too.