site banner

Culture War Roundup for the week of April 13, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

3
Jump in the discussion.

No email address required.

Another indicator that AI is a bubble. Anthropic just released Claude Opus 4.7, and users are reporting significantly higher token burn rates (and therefore costs) for what appears to be a minor improvement over Opus 4.6. Discussion on Orange Reddit is here: https://news.ycombinator.com/item?id=47816960 and a tracker of the increased token burn rate is here: https://tokens.billchambers.me/leaderboard

The token tracker is based on user reporting, but has been fluctuating between 37% and 45%.

Even if AGI is actually possible with LLMs (or at all, but I'm not trying to start a discussion on metaphysics here), it looks like the capital needed to achieve it is drying up before it can be reached. Anthropic's move here (combined with them handicapping Opus 4.6 a few weeks ago) seems to clearly be an attempt to achieve profitability. The free/subsidized rate train for end users has pulled into the station, and now you have to pay more for the same (or worse) capabilities you were enjoying before.

I normally don't care much for the median Hacker News commenter (if me calling it Orange Reddit didn't already give that away), but I do find them to be a useful barometer for general sentiment in the tech industry. And a few months ago I would have said roughly 60% of HN users were AI believers/enthusiasts, 20% neutral or unsure, and 20% anti/negative. Anthropic's antics over the last few months (and Sam Altman's antics for his entire life) seem to have soured their views significantly, and I see this as a big sign of a sea change in sentiment about AI in the tech industry.

At least for me personally, I just hope this leads to less retarded mandates from my higher-ups about using AI X times a month etc. (we're literally tracked on usage and it can affect our raises/bonuses).

For everyone here, nut perhaps especially the AGI believers, have your feelings changed at all over the last few months?

Even if AGI is actually possible with LLMs

I'm pretty convinced it isn't, based on a thought experiment I read about.

The argument goes basically like this:

Suppose you take the latest and greatest LLM and use it to generate a huge corpus of text and use that text to train a new LLM. And then repeat the process a number of times. Intuitively, it seems unlikely that the result will be any better than what you started with. And apparently both experiments and mathematics indicates that what happens is "model collapse," i.e. with each iteration the new model performs worse. Because you always lose a little with each iteration. Assuming that's all true, it follows that LLMs must be missing some essential attribute possessed by human brains. Because we apparently picked ourselves up by our bootstraps and created from scratch all the text which is used to create LLMs.

Anyway, it's just an argument I read and found to be persuasive. Feel free to correct me.

Another indicator that AI is a bubble

To me it's pretty obvious that AI is wildly over-hyped. But even so, the progress which has been made in the field is nothing short of astounding.

it looks like the capital needed to achieve it is drying up before it can be reached.

If nothing else, it's seems virtually certain to me that governments have realized the strategic implications of AI. Even without any private investment at all, the United States, China, and various other countries can throw quite a lot of resources at the problem.

For everyone here, nut perhaps especially the AGI believers, have your feelings changed at all over the last few months?

Not really, I'm still pretty confident that (1) within the next 10 years or so, we (humanity) will get to AGI; and (2) regardless, there will be huge changes to the world economy.

And apparently both experiments and mathematics indicates that what happens is "model collapse," i.e. with each iteration the new model performs worse.

Model collapse is not really a major concern. The original researchers in that paper trained small models on only AI outputs (of the previous model). Them being small models, they made mistakes and the mistakes compounded over time. It's more like a Chinese whispers experiment.

Big companies make great use of synthetic data and autonomous training, in addition to human originated data. For example, consider Deepseek R1-Zero, which was just trained on reinforcement learning, verified signals and not human reasoning patterns. It was kind of weird and switched languages a lot but it did work and got smarter over the course of training. In fact, all modern models are trained in this way. When Claude occasionally slips into Chinese for a single word it's not because any human ever does that in the training corpus, it's because during the training process they have them autonomously bootstrap and get smarter over time and that's just how it goes. AIs are omnilingual by nature it seems.

Model collapse is not really a major concern.

If you say so, I have no reason to doubt you. But what does that say about the thought experiment I proposed? Are you saying that potentially the 1000th model could be significantly better than the first?

Yes, I think so, provided you were doing the training in a sophisticated way rather than solely training on the outputs of previous models without grading for quality or accuracy. You could get AIs to review the data for example for any errors or issues or have them work out a testing suite to check if the data is right. Data quality is very important, that and the right RL techniques are basically the two key things you need most to get right.

Microsoft Phi trains just on synthetic data and is very cost-efficient, that was its primary goal, making a good very small AI that can run on most PCs. But they curated the data a fair bit to make sure it was good.

In principle I think you could do the same for big first rate AIs too. It's just that it wouldn't be efficient to leave out human data and human curation (it's there, why not use it, the competition will) and you want something humans enjoy working with and not a schizo-sounding model. It'd be like o3 at its most alien but more so:

https://arxiv.org/html/2510.27338v1

they soared parted illusions overshadow marinade illusions overshadow marinade illusions overshadow marinade illusions

Number of relevant organic products depends on whether both of!mena get.demoteudes someone and gem jer eats SAND the protonation-bids, leading possibly to three product calculation

Like wtf does that mean? Who knows? This is an artifact from inhuman RL processes. The inhuman RL processes work, that's why they're used.

I'm also convinced that LLM's aren't the path to AGI. They were grossly overrated from the very beginning. If you want real AI, you have to dump the snake oil. It's already empirically well adduced (1, 2) that LLM's can't get you there.

There's a 'lot' of things you need before you've actually got an intelligent machine that can think. It begins with constructing mental models. From there, you navigate those models in the imagination or perceptual space laid out to work out answers to questions and work out alternatives. Once you build those spaces (and in this case I mean building creative and entirely novel ones) and navigate them to accelerate anticipatory learning. Cats for example actually "learn" how to hunt by doing this. Once you've learned how to model spaces you can them move to modeling "systems;" and that's when you get to the point where it becomes possible to give AGI a theory of 'other' minds. And a "mind" in purely mechanistic and computational terms is simply another causal system; just like "spaces" and "systems" are particular causal systems.

Notice that's exactly what an LLM 'doesn't' do. If you take a look at Waymo's World Model for instance, this is exactly more along the lines for the correct pathway of approach that you need. When you're continuously inventing new models of imaginary environments, you begin to build the skillsets that slowly become applicable to the real world. When it can do that, that's more along the lines of where scaling becomes effective. Nothing like Sam Altman's idea of where it's relevant. When you've got to that stage, AI can then begin to model it's own causal system to think about it's own thinking such that it's capable of asking itself when it's wrong; or how to stack a particular sequence of events to achieve a desired end result.

Incidentally this is the exact pathway natural selection determined for human beings and I'm thoroughly convinced it's the 'only' way to get AGI. This is what human beings fundamentally are on the naturalist paradigm: models and model builders that also navigate and move about in those models. There's zero evidence that I've seen to indicate that the money flowing into AI at the present moment is traveling down that research pathway and make no mistake, eventually the supply of it is going to run out. But make no mistake. A 'lot' of rich people are stupid, so it's doubtful it'll ever go to the real thing. They'll throw it all at the next bullshit snake oil, get their fiscal bailout and blame it on immigrants on something. Presently I'm not left feeling very optimistic about the current state of the industry. It's incredibly destructive to the environment, it dumbs down the human intelligence, and it hasn't even been proven to even work. Why is the “world” so excited?

What made me even more dubious about the entire grand project than I had already been was the news that now they were generating their own data to train models on. We've scraped every single bit of text produced by humans in all of history to date (ahem ahem I take my leave to doubt that, what you mean is 'we've scraped all the available English language text online') and now we need even more to feed the gaping maw of Behemoth, so now we have to invent our own synthetic text generated by AI.

Do they not remember "garbage in, garbage out" or, indeed, Flanderization? Generating your own synthetic data off sythetic data and using that to create more synthetic data and synthetic demographics is getting further and further away from reality, then some poor fool uses the conclusions your AI served up so prettily to make real world decisions and it turns out that in fact 15-24 year old mixed race lower middle class exurban teenagers with sports scholarships do NOT want to wear pink clamdiggers topped off with stovepipe hats. That's your entire chain of stores' summer wear stock now useless even on sale.

While I share your skepticsm, this Dineen makes intuitive sense. But given that all the labs seem to be doing this and are not super concerned about it, I assume it serves it's purpose.

Mythos is a big boy, I imagine there's lot of synthetic data in there and it seems to be working

Intuitively, it seems unlikely that the result will be any better than what you started with. And apparently both experiments and mathematics indicates that what happens is "model collapse," i.e. with each iteration the new model performs worse.

Yes, this follows from data processing inequality.

Assuming that's all true, it follows that LLMs must be missing some essential attribute possessed by human brains. Because we apparently picked ourselves up by our bootstraps and created from scratch all the text which is used to create LLMs.

No. It applies just as well to humans. And humans did not build a civilization by thinking really hard at a corpus of word sequences. Oh, we tried this too, to an extent, and got wonders like Sophistry, Rabbinical Judaism, Medieval Scholasticism, Marxism and Rationalism. But we mostly progressed by receiving environmental feedback, filtering the generated data and preferentially training on validated fraction. Similar logic can be applied to LLMs (or any ML artifacts). This is why the basic trick of the current paradigm is RLVR (reinforcement learning with verifiable rewards). You finetune a model on successful trajectories, then you give it tasks and update towards policy that has generated correct conclusions. The primary source of updates is the model itself, steered by an external verifier. In principle they can do this fully autonomously, by building an ontology of possible tasks that can be algorithmically verified, coding these verifiers, and generating (eg relying on web search) queries against these tasks.

Even under very rudimentary realistic assumptions, generated data improves model performance.

Marxism and Rationalism

I don't understand how these fit into the category like the religious examples

Sophistry is not religious either.

I don't want to make a definition for this category because it's very loose, but basically it's "attempts at recursively improving your understanding via introspective self-play starting from a given set of verbal premises, without any significant role for procedures of updating on empirical, physical evidence".

One can see how this might well work in fields which really don't need an empirical physical component, such as math. Physics can inspire new subdomains of math, but strictly speaking we don't need this. An AI could train on its own data (+ easy verifiers and just corrected majority voting), entirely autonomously, to become an ever stronger mathematician.

No. It applies just as well to humans. And humans did not build a civilization by thinking really hard at a corpus of word sequences. Oh, we tried this too, to an extent, and got wonders like Sophistry, Rabbinical Judaism, Medieval Scholasticism, Marxism and Rationalism. But we mostly progressed by receiving environmental feedback, filtering the generated data and preferentially training on validated fraction. Similar logic can be applied to LLMs (or any ML artifacts). This is why the basic trick of the current paradigm is RLVR (reinforcement learning with verifiable rewards). You finetune a model on successful trajectories, then you give it tasks and update towards policy that has generated correct conclusions. The primary source of updates is the model itself, steered by an external verifier. In principle they can do this fully autonomously, by building an ontology of possible tasks that can be algorithmically verified, coding these verifiers, and generating (eg relying on web search) queries against these tasks.

Sadly, I do not understand this. Would you mind giving me a concrete example of the RLVR process you refer to?

Question: What is 2 + 2

Model: Hmm, that’s 2 and then another 2, so 22.

AUTOMATIC VERIFIER: WRONG

——

Model: Hmm, that’s the sum of 2 and 2, so 4

AUTOMATIC VERIFIER: CORRECT.

The model is tweaked slightly to make the second output more likely, and that output is potentially added to the training set. Repeat for arbitrarily complex mathematics and other problems as long as the solution can be verified, even if it isn’t known in advance. In this way you can generate potentially infinite amounts of data, albeit limited to certain domains. However, problem solving ability has so far extended quite well to other domains even when trained in this manner.

Question: What is 2 + 2

Model: Hmm, that’s 2 and then another 2, so 22.

AUTOMATIC VERIFIER: WRONG

——

Model: Hmm, that’s the sum of 2 and 2, so 4

AUTOMATIC VERIFIER: CORRECT.

The model is tweaked slightly to make the second output more likely, and that output is potentially added to the training set. Repeat for arbitrarily complex mathematics and other problems as long as the solution can be verified, even if it isn’t known in advance. In this way you can generate potentially infinite amounts of data, albeit limited to certain domains. However, problem solving ability has so far extended quite well to other domains even when trained in this manner.

Generally speaking, how does this "automatic verifier" work? Obviously I am not an expert but it seems like this automatic verifier would require human level intelligence.

In this toy case it's just literally a calculator (a snippet of python code). The problem is 2+2, the calculator just does 2+2 and checks if the answer is the same as the LLM output. (The LLM is trained to format the final answer in a particular manner and wrap it with special tokens, so the verifier doesn't have to be able to interpret natural language.)

You can get surprisingly far with this. If it's a calculus question, you can use an automatic differentiator to check it. Likewise for factorisation questions, metric conversion questions, algebraic manipulation of formulae, etc. you put a little work into programming the automatic verifier and you can get an infinite number of problems.

If you're a big company, you might have human domain experts doing some of this work too. If you're a smaller company you have a big LLM do verification for the smaller ones.

Then you have leetcode and programming problems, and again you can verify these automatically. Does the program compile? Is the program output what was requested? Is it faster than the previous solution?

Like I said, this only works for maths, programming, and other domains where you can verify the answer with a computer relatively cheaply, but contra the model of multiple intelligence factors, heavy training on maths and programming seems to improve general intelligence and reasoning quite well.

Like I said, this only works for maths, programming, and other domains where you can verify the answer with a computer relatively cheaply,

This is what the armies of Kenyans are for. I'm actually surprised progressive libs don't use "muh mechanical Turk sweatshop" as an anti AI talking point.

Also I think in some ways what they use the thumbs up/down for? I saw people saying the sycophantic behavior of the 4o era was people love being glazed and thumbs that up a LOT so it crept in.

In this toy case it's just literally a calculator (a snippet of python code). The problem is 2+2, the calculator just does 2+2 and checks if the answer is the same as the LLM output. (The LLM is trained to format the final answer in a particular manner and wrap it with special tokens, so the verifier doesn't have to be able to interpret natural language.)

You can get surprisingly far with this. If it's a calculus question, you can use an automatic differentiator to check it. Likewise for factorisation questions, metric conversion questions, algebraic manipulation of formulae, etc. you put a little work into programming the automatic verifier and you can get an infinite number of problems.

If you're a big company, you might have human domain experts doing some of this work too. If you're a smaller company you have a big LLM do verification for the smaller ones.

Then you have leetcode and programming problems, and again you can verify these automatically. Does the program compile? Is the program output what was requested? Is it faster than the previous solution?

Like I said, this only works for maths, programming, and other domains where you can verify the answer with a computer relatively cheaply, but contra the model of multiple intelligence factors, heavy training on maths and programming seems to improve general intelligence and reasoning quite well.

Thank you for the explanation. My instinct is that even with this type of training, LLMs will still be missing something essential, but I will give it some thought.

My instinct is that even with this type of training, LLMs will still be missing something essential

Your instinct is probably correct IMO. This form of synthetic data generation is just another tool in the box, it's not the key to everything.

I will say that we've got far further than I ever expected us to get using these methods. I'm instinctively a Gary Marcus-style fan of embodiment and unsupervised learning, it seemed clear to me pre-LLM that models wouldn't be able to be anything resembling intelligent without a body and the ability to interact with the real world and 'test' their understanding in real time. When LLMs came in, I felt I had to admit that I'd been wrong. It seems clear to me that we have managed to get to something I would call 'intelligence' (even if it's spiky and fails in some cases where humans would not fail) through these means. So I no longer trust my instincts as much.

This kind of semi-supervised exploration seems like a good compromise for now. I am also very interested in LLMs that can combine next-token video generation and text generation, because video generation requires understanding a bunch of stuff about the real world in order to produce consistent results, but that's a way off.

We formulated our understandings of the world and our interactions with it into techniques and theories, and when we build stuff we do so by employing those techniques and theories from a standpoint of engineering and design. LLMs are merely next word generators. They can recall many of the things in their databases and expurgate them to us, but their outputs aren't the products of strategically employed techniques and theories. This is inherently limiting for the complexity of the outputs they can give us.

I don't understand this claim. Who "we"? Most people learn almost everything they know about economically valuable complex domains from textbooks, manuals, teacher's answers and such second-hand information, and then polish it with on-site instructions and increasingly long-range, open-ended training. They don't build much in the way of their own "techniques and theories" and there's not a world of difference from what LLMs now do. Maybe you're overestimating how much they depend on pretraining at this point. Well, it's believed that >50% of compute in some of the last-generation models goes towards RL, not pretraining on human data.

And as I've said in the opening post: we have literally just seen an LLM employ a technique no human mathematician had thought of using in this specific context, to solve a problem that had remained unsolved since 1968 – over half a century! It wasn't some Riemann hypothesis tier challenge, but it wasn't exactly obscure either, smart professional mathematicians had been working on it for years before GPT 5.4 Pro came and did this. Moreover, GPT does this reliably. In the comments you can see Terence Tao, arguably the guy with the greatest knowledge of "techniques and theories" of math on the planet Earth, an expert of such level that he actively avoids getting roped into solving other people's frontier research level problems, seriously engage with GPT's work:

Thanks! So there does seem to be something special about the original von Mangoldt process - the associated invariant measure ν is extremely smooth (in the Archimedean sense), being asymptotic to 1/nlogn , while all the variants of this measure pick up arithmetic factors such as 1∏pvp(n)!

  • A little surprising to me that removing individual primes instead of prime powers makes it less likely to have prime multiplicity, but I'll chalk it up to one of the numerous probability paradoxes that arise when one tries to compare various weighted expectations. But these factors mean that one cannot immediately solve #1196 by using these processes instead of the von Mangoldt one, as the invariant measure is no longer asymptotic to 1/nlogn
  • So in some sense the AI was "lucky" in finding the one approach that actually worked; it would be interesting to publish the traces to see if there was a lot of brute force involved in trying nearby approaches which didn't quite work.

……

Arb Research has kindly shared with me ten separate runs of GPT 5.4 Pro on this problem #1196 (with a request not to use internet search). From a quick reading, it appears that 8 of them claimed successes, with the other 2 rating the claim as plausible. Interestingly, several of the successful runs actually obtained the sharper formula ∑n≤Aν(n)≤1 that was also derived here, with ν essentially the Mellin transform of 1/ζ(s)

  • Almost all of the runs latched on to the approach of constructing a random chain with a good hitting probability (many runs referred to this as the "Lubell method", after the Lubell of the LYM inequality).

Another notable fact is that none of the runs highlighted the von Mangoldt process that was a prominent feature of the original run (and none of them mention flow networks either). Runs 4 and 7 have an interesting alternate construction of the upward divisibility chain in terms of exponential clocks in the prime factorization indices that actually looks rather tractable to work with; I will need to study this construction further when I have more time.

Basically it seems that for this particular type of problem there are several natural ways to proceed that make the problem actually quite tractable; the literature had managed to focus on a somewhat suboptimal approach in which the opening move was to transfer the problem to a continuous setting, but the AI runs consistently stayed in the discrete world and managed to utilize various existing tools from discrete mathematics (mostly centering around methods relating to the LYM inequality) to reach a solution.

So I don't know. Where's this inherent limit on complexity that you're talking about? What in our culture is truly irreducibly complex, if not math that can surprise Terence Tao?

This is getting a bit comical, don't you think?

I must differ here as I do not see evidence (in domains I'm able to judge) of AI employing techniques and theory in its tasks. Ask it to mimic Stephen King and then compare the output to actual Stephen King. You'll understand what I mean.

I cannot speak to math here as I lack competency in that. But from what I hear from coders, its similar in that domain as well: AI can expurgate volumes of legible code, but it cannot utilize structure.

Humans have techniques and theories which inform their decisions high and low as they layer things together using judgement, intuition, etc., while AIs appear to generate text using probabilistic hacks. AI appears to be able to recreate low-complexity patterns from its dataset. I disagree that these processes are related except at a very basic level.

We have a good idea of how to train AI to solve mathematical problems, of virtually unbounded complexity. In the course of this, AI clearly learns "techniques" as shown here, if not "theories". I don't think King's prowess is theory-driven either, but in any case we don't have a good idea of how to train AI to be a good prose writer. We have some ideas, but are unlikely to act on them. There's not much money to be made in it, and plenty of highly motivated enmity – AI is already widely hated. and yes, autoregressive generation for the prompt "write like King" is not like King actually writing a novel. We have such tricks though.

My point is, it's not a general principle that AI will only rehash human techniques in some uninspired "probabilistic" way. If there is a hill to climb, such that "good" and "bad" outputs with regard to the problem statement can be distinguished, AI can bumble its way up the hill and also find new tricks. We've seen this before LLMs, with AlphaGo and move 37, we're starting to see it with LLMs.

while AIs appear to generate text using probabilistic hacks.

Human mind runs entirely on probabilistic mush. Neural networks were invented as approximation of our own approximate learning. But probabilistic decision processes can have clear enough decision boundaries that they become able to operate with "abstractions", "symbols" or "theories". They also remain able to fail. For example, you are failing to update on evidence, because you haven't been trained to take input like "Terry Tao is surprised" seriously and think it's infinitely less interesting than your preconceived notions, basically some dweeb noise. Unlike an LLM, you can update at lifetime, so maybe you'll reread the above post and see how it contradicts your position.

This is getting a bit comical, don't you think?

Seen on X:

"As the Earth is being disassembled:

"Guys, stop over-reacting! The concept of a Dyson Sphere was already in the training data!"

Heh. See, the AI making that Dyson Sphere doesn't have general intelligence, I bet it can't get the Wordle 6 days in a row like me.

Because we apparently picked ourselves up by our bootstraps and created from scratch all the text which is used to create LLMs.

Or it holds for human brains, but we train on something higher than our text, and so LLMs are upper bounded by us if they train on our text, which is like our cognitive discard.

Suppose you take the latest and greatest LLM and use it to generate a huge corpus of text and use that text to train a new LLM. And then repeat the process a number of times. Intuitively, it seems unlikely that the result will be any better than what you started with. And apparently both experiments and mathematics indicates that what happens is "model collapse," i.e. with each iteration the new model performs worse. Because you always lose a little with each iteration. Assuming that's all true, it follows that LLMs must be missing some essential attribute possessed by human brains. Because we apparently picked ourselves up by our bootstraps and created from scratch all the text which is used to create LLMs.

Where did you hear that anyone is proposing to reach AGI via LLMs by training LLMs on their own generated output? That's clearly dumb and not what people propose. The model has to interact with something real, it has to "touch grass", for it to work. That's the external information. For example a coding LLM can get an informative learning signal by running its generated code through the compiler and running tests and seeing if the resulting program compiles, passes the tests, uses less RAM or is faster, etc. I'm not saying that leads to AGI, but there are clearly ways to obtain information from the outside world, and it's not just about sewing a pipe from the LLM's ass back into its mouth.

Where did you hear that anyone is proposing to reach AGI via LLMs by training LLMs on their own generated output? That's clearly dumb and not what people propose.

I presented the idea as a thought experiment, not as an actual proposal.

No, you presented it as a conceptual proof that LLMs will never get better. All it takes is one innovation that addresses your concern about recycled data to make it invalid. All arguments about intelligence are necessarily a bit wishy-washy, mind you, so I'm not saying your thought experiment is useless.

I think if you really want to argue that LLMs have an inherent cap on their capability, you should address their actual algorithm rather than how they're trained. However much we rejigger them with CoT thinking and non-text data sources, they're fundamentally not designed for anything more than next-token prediction. It should be a source of constant surprise that they do so well on such a wide variety of non-creative-writing tasks (look at early SSC posts about GPT3's output to see this surprise evolve in real time). You could argue that if LLMs end up hitting a soft or hard limit, that's really just the "surprise" petering out, that we really can't just take a glorified text completer and keep pumping neurons into it until it's a genius.

I don't personally believe this will happen, but hey, I don't think anyone really knows for sure.

No, you presented it as a conceptual proof that LLMs will never get better.

Umm, no. In fact I totally think that LLMs will get better.

I presented it as a thought experiment to show that LLMs seem to be missing some essential attribute possessed by human brains.

ll it takes is one innovation that addresses your concern about recycled data to make it invalid.

Yes, of course. Well, perhaps more than one innovation. But yes, if LLMs are missing something important; and we create LLMs 2.0 which include that important thing (or those important things), then yeah, we'll have AGI.

Because we apparently picked ourselves up by our bootstraps and created from scratch all the text which is used to create LLMs.

This is clearly proof that the Saurian Overlords of Agatha in the Hollow Earth taught us language.

More seriously, isn't there a lot of research going into using synthetic data safely? I thought that the current consensus was that you can avoid model collapse with synthetic data if it's properly labeled as such.

More seriously, isn't there a lot of research going into using synthetic data safely? I thought that the current consensus was that you can avoid model collapse with synthetic data if it's properly labeled as such.

I have no idea, but intuitively it seems to me that training with synthetic data is something that can't possibly work. To be sure, I am neither a mathematician nor a computer scientist. But as I understand things, the basic operation of an LLM is to predict the most likely words to follow a string of words. Which is done by training a neural network on lots of text. It's difficult for me to see how an LLM could get better at predicting words if it's trained on its own output. It seems to me that if you created an LLM using a corpus of synthetic data, the best you could realistically hope to do would be to reverse-engineer the original LLM which had created the synthetic data in the first place.

Anyone who is a subject matter expert, feel free to correct me.

I have no idea, but intuitively it seems to me that training with synthetic data is something that can't possibly work.

I agree with you intuitively. However large amounts of Serious People are spending large amounts of Serious Economic Resources to do literally just that. So they clearly see something there.

Mythos seems to be yet another OOM of compute and training and we can be sure a good % of that was synthetic at this point as they already ate the entire corpus of human writing a while ago

I agree with you intuitively. However large amounts of Serious People are spending large amounts of Serious Economic Resources to do literally just that. So they clearly see something there.

Part of me wonders whether this might be some kind of mass delusion and/or grift. It wouldn't be the first time something like that has happened. That being said, I am not a subject matter expert and haven't studied these issues carefully so I couldn't really say.