site banner

Culture War Roundup for the week of August 31, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

3
Jump in the discussion.

No email address required.

Dead forum?

Today in AI news, OpenAI has reportedly released GPT-6 Astra. Rumors on X suggests that it scores 99% on Arc-AGI-3, meaning that the benchmark is essentially dead now. For comparison, Fable only scored 30%. Does this mean AGI is here? Many are saying that it is. Color me skeptical, but its getting harder and harder to find concrete tests that AI cannot pass. The god of the gaps gets smaller.

Many intelligenct people have important things to say about this. But what do the 84 year old communists think? Bernie Sanders today has introduced legislation calling for a ban on superintelligence with a 20 year prison term for attempted superintelligence. He also wants (shocked Pikachu) a government regulatory agency. Although he mumbles something about international cooperation, its clear that China will not be subject to these rules so it solves nothing.

I will give him credit for at least recognizing the importance of the issue.

We could be entering a world of unimaginable change. And, given the political systems in place in the US and China, the midwives of that changes will be those on the cusp of senility.

AI is culture war in a way (although contempt for it often seems bipartisan) but for now, we don’t really know how any of this plays out. You can write science fiction like Scott and dress it up as prediction, but that feels insincere. Interesting things are going to happen, yes. Beyond that, we know little.

I don't know what any of these eval scores mean. I no longer have the ability to intuitively "feel" how much more intelligent each iterative model release is. I don't know how to code. I don't know how to do graduate-level math. These models have been smarter than me for quite some time.

As someone who uses these models to code every day, I also don't intuitively feel the jump from model to model. I suspect the people who pop out of the woodwork to say there's a huge difference are mostly just disingenuous engagement farmers at this point, with the occasional normal user who just happened to have a specific niche use-case that randomly got big improvements from one model.

Though I will say that I can still feel the difference in the long term. The models today are a step up in terms of coding compared to the models from a year ago, though many issues around context length and basic computer usage still persist.

I'm surprised people aren't humble bragging more about how their work is so fucking heady and sophisticated that they have no choice but to pay a premium for Fable 5.1 all of the time and even being forced to use 5.0 would delay completion of the Dyson sphere by years.

The ARC-AGI-3 comparison isn't apples-to-apples; they have a custom harness to call the API more intelligently. Using the old one, IIRC it was more like 60% or something. Still a meaningful step up from Fable.

More broadly, keep in mind: the speed things are moving this week is the slowest they will ever move. Wondering if we'll get a new Millenium solution by EOY.

custom harness

BS like this is way, way too common in benchmarks.

Dead forum?

I guess there's just no-one left here who cares enough about the issues that energized discussion in Reddit days.

In the announcement blog post, footnotes 9 and 10 concern mathematical discoveries regarding prime gaps. In 9 they link to proof to there being an infinitely many primes that differ by at most 186, in 10 that there are infinitely many primes p_n that are greater than the previous prime p_(n-1) by at least ln(n)ln(ln(n))^2 ln(ln(ln(ln(n))))/ln(ln(ln(n)))^2.

Not linked in the announcement, but also products of this model are: disproof of Erdos Problem 1 (according to Bloom's numbering), disproof of EP74, proof of EP126, proof of EP548, proof of EP571.

As mathematician's are usually asked to pre-review OpenAi's proofs (but not those published by Alpöge of Anthropic), Terence Tao must have known about this prior to thus publication. This would explain today's doomposting of his on mathstodon about how open problem are needed by mathematician to discover new techniques and proposing certain problems be offlimits to AI investigation. In my mind this summoned a thought of ChatGPT treating a prompt to consider prime gaps as it does a prompt about explossive manufacture.

I just want to know how many credits it costs me with Github Copilot.

It's not AGI, and we'll never have AGI until we have an AI that can do everything better than humans can, at which point it will be ASI. Anything less than that would be prone to being made fun of on social media as "if it's AGI then why can't it do [thing] that is totally easy for humans???" E.g. the best AIs can only barely play Pokemon games that are meant for toddlers, finishing them in 10x-100x the usual playtime that a human would require.

Scoring well on one benchmark might be interesting but it's unlikely to be a step-change in overall capabilities. AIs still have big issues with context limits and real-world interaction, even just in terms of computer use. Progress in those areas remains relatively anemic, and is probably a requirement for lots of other capabilities.

Expect regulations to be nonexistent until it does something to shock the public out of thinking AI is useless. The median voter hates AI but thinks it's just a slop machine, and how could a slop machine ever do anything harmful if it's so stupid? Eventually an AI will do something that can sort-of somehow plausibly be linked to something like killing children, and thereafter the regulations will be swift, extensive, and unreasonable.

Eventually an AI will do something that can sort-of somehow plausibly be linked to something like killing children

This has already happened a bunch of times. I think Tumbler Ridge was the first mass murder where those conditions were true; there were a few scattered suicides before that, but nobody really cared about those guys.

(One imagines much jealousy by state organizations responsible for fighting this sort of thing; they at least have to make house calls to convince some retard from $group to do it. At least the AI hasn't gone full Eagle Eye... that we know of, at least.)

E.g. the best AIs can only barely play Pokemon games that are meant for toddlers, finishing them in 10x-100x the usual playtime that a human would require.

Are there new benchmarks on this? Best I can find in a quick search is a year old (i.e. 7 dog years, 20 LLM years).

Astra 6 beats FireRed in 18 hours, comparable to a (non-speedrunner) human player.

I want to know which starter the different models pick.

Thanks for the link. That seems like a fairly significant increase in capabilities, though

  1. I wish they streamed it and had vods, and
  2. I wish they were doing it on the old Red/Blue for a more direct comparison.

There is a stream going on right now! It's three hours in, and Astra is fighting bug catcher Cale in route 24. Come be a part of history!

There's pretty recent data, and the most recent runs are close to human parity, but they're also least well-documented.

((I'd also argue that Pokemon is intended more for the 10-16 range, but the general problem remains for them, too.))

I had trouble with Pokemon games when I was 6. Mt. Moon was an almost insurmountable challenge.

Yeah, Pokemon Blue was my first proper video game at age 5 (though I might have played Freddi Fish before it, I don't remember the exact time line), and I needed help from my cousins for two sections of the game.

((I'd also argue that Pokemon is intended more for the 10-16 range, but the general problem remains for them, too.))

Red is 11. It makes no logical sense for a kid that age to leave home, defeat every adult he encounters, destroy a criminal organization, and become world champion. Only reason to do that is because your target audience is 11-year-olds.

Does this mean AGI is here?

AGI is here when we can't make ARC-AGI 4 (that a median human can pass but the frontier can't).

Bernie Sanders today has introduced legislation calling for a ban on superintelligence with a 20 year prison term for attempted superintelligence

He should've just cited AI 2040 instead of that vague proposal. The rationalists are genuinely smart and dedicating their lives to this, and I really believe they are the ones who are best equipped to handle regulation, even if their plans sound infeasible and sometimes it seems they've given up (and the other times they're naively optimistic up to delirium). Imagine what our pre-ASI utopia would've been if only politicians listened to autists.


I'm interested to see if it can write like a normal human and be interesting. If not, maybe it's for the best.

AGI is here when we can't make ARC-AGI 4 (that a median human can pass but the frontier can't).

That's a good definition, but there's still the additional hurdle of being able to learn after pre-training. LLM's still do pretty poorly when they have missing or poisoned data in their pre-training. Whereas humans (some of us at least) really can "generalize" with much less data.

I'm interested to see if it can write like a normal human and be interesting.

One thing that LLM's are really weak at is modeling what humans find interesting. Also probably for the best.

being able to learn after pre-training

True AGI must be able to learn continuously, or at least emulate it well enough.