Be advised: this thread is not for serious in-depth discussion of weighty topics (we have a link for that), this thread is not for anything Culture War related. This thread is for Fun. You got jokes? Share 'em. You got silly questions? Ask 'em.
- 202
- 1
What is this place?
This website is a place for people who want to move past shady thinking and test their ideas in a
court of people who don't all share the same biases. Our goal is to
optimize for light, not heat; this is a group effort, and all commentators are asked to do their part.
The weekly Culture War threads host the most
controversial topics and are the most visible aspect of The Motte. However, many other topics are
appropriate here. We encourage people to post anything related to science, politics, or philosophy;
if in doubt, post!
Check out The Vault for an archive of old quality posts.
You are encouraged to crosspost these elsewhere.
Why are you called The Motte?
A motte is a stone keep on a raised earthwork common in early medieval fortifications. More pertinently,
it's an element in a rhetorical move called a "Motte-and-Bailey",
originally identified by
philosopher Nicholas Shackel. It describes the tendency in discourse for people to move from a controversial
but high value claim to a defensible but less exciting one upon any resistance to the former. He likens
this to the medieval fortification, where a desirable land (the bailey) is abandoned when in danger for
the more easily defended motte. In Shackel's words, "The Motte represents the defensible but undesired
propositions to which one retreats when hard pressed."
On The Motte, always attempt to remain inside your defensible territory, even if you are not being pressed.
New post guidelines
If you're posting something that isn't related to the culture war, we encourage you to post a thread for it.
A submission statement is highly appreciated, but isn't necessary for text posts or links to largely-text posts
such as blogs or news articles; if we're unsure of the value of your post, we might remove it until you add a
submission statement. A submission statement is required for non-text sources (videos, podcasts, images).
Culture war posts go in the culture war thread; all links must either include a submission statement or
significant commentary. Bare links without those will be removed.
If in doubt, please post it!
Rules
- Courtesy
- Content
- Engagement
- When disagreeing with someone, state your objections explicitly.
- Proactively provide evidence in proportion to how partisan and inflammatory your claim might be.
- Accept temporary bans as a time-out, and don't attempt to rejoin the conversation until it's lifted.
- Don't attempt to build consensus or enforce ideological conformity.
- Write like everyone is reading and you want them to be included in the discussion.
- The Wildcard Rule
- The Metarule

Jump in the discussion.
No email address required.
Notes -
So I know a lot of posters here are doing some really advanced stuff with AI, but I have a quick little anecdote from my own luddite life that might help explain why most normies don't really see AI as something particularly useful.
I just asked Google's free AI what the current value of 0.035 ounces of gold is and it completely face-planted, ignoring the decimal and the zero and giving me the value for 35 ounces. This isn't a case of overactive guardrails or confusing syntax or hallucinating plausible details into a news story; it's the most basic of transcription errors, the kind that dumb, pre-LLM systems have solved for decades. Old google never would have made such an error.
This is the technology that's going to completely upend the world order? It makes sense for normies to be skeptical. Regardless of what the frontier labs have cooking, the AI most people actually interact with is worse than useless half the time.
Anybody else have one of these "the fuck are we even doing here?" moments with AI recently?
Which model was that? Instead of saying "well yeah use a better model", I thought I'd try to replicate it. So I just went to google.com in incognito and dropped in
what is the current value of 0.035 ounces of gold, and it gave a good, accurate answer from it's AI, with a correct calculation and a link to the page it got the current price from. That's the regular Google search interface and no login, so it's got to be the most heavily cost-optimized model they operate.I don't generally make a habit of trying to seek out the dumbest-possible AIs, and I know they're non-deterministic, but I haven't ever seen one make that kind of totally braindead error before.
The fact that AI may only generate these kinds of errors a small percentage of the time is only a reasonable argument when the comparison is with something that makes similar errors a greater percentage of the time. Writing a computer program that queries a database for commodity prices and gives the current values of a given quantity is trivial to write, and can be done relatively easily on a home computer from the 1980s. Despite the relatively low tech, it won't make mistakes very often, and when it does they will be in the form of the program crashing entirely, not misplacing a decimal point or giving a plausible but incorrect response. I have been using Excel for decades, and this has not been a problem. Even if the LLM output is only wrong one time in every thousand attempts, that's still immeasurably worse than our simple computer program. And then you add in the fact that you're making this calculation through a process that requires a billion-dollar offsite computer that cost untold billions to train to get an answer that is less reliable than from a pocket calculator given away for free at a bank, and people are going to look askance.
"But," you may say, "what you describe requires a separate program. The beauty of the LLM is that it can answer general purpose questions." Which is true, but if the error rate is too high, it still isn't particularly useful. In the commodity price example, if I actually need to know the price of 0.035 ounces of gold for business purposes, it's worth it for me to have a computer program designed for the purpose that will always give an accurate answer. If I don't, then I'm not paying for a technological solution, period.
More options
Context Copy link
More options
Context Copy link
"I've heard all about these Chinese cars totally upending the global auto market and I'm told they're absolutely mogging the established auto players in basically every non Western country, so to investigate I test drove a 2003 BYD Flyer and it sucked balls, so clearly this is all fake news"
What's funny is I actually agree with your thesis. I think that the various free/quantized models that normies largely interact with (most notably, google search and copilot) have totally given everyone a terrible impression.
But this style of the motte investigative reporting is just ... Lmao
The difference is that BYD isn't pushing their 25-year-old cars for marketing purposes. For example, most of the time that a new food product comes out, you have to buy it if you want to try it. Sometimes, the manufacturer will hire a company to distribute free samples in grocery stores. Imagine if, in order to save money, the samples all came from batches that QC rejected for being not up to par. If the product failed to sell, and market research suggested that people weren't impressed with the samples, and the company's response was "You can't judge the product based on the free samples; you have to get the real version", nobody would have any sympathy for them. It would end up in marketing textbooks as an example of what not to do, and people like Rory Sutherland would repeat the story every time they gave a talk.
I agree that the free AI samples are trash, I just think once we're all operating at the meta-level of understanding this, repeatedly pointing this out is a pointless exercise. Like congratulations, you have successfully noticed that something that is obviously low quality is low quality. Are the AI labs dumb for letting this pollute their brand/the overall "brand" of AI? Yes.
It's like mild "boo outgroup" posting but the outgroup is free shitty AI models
If people repeatedly went to McDonald's to make effort posts about how bigmacs are trash compared to their local favorite smashburger place, I'd have the same opinion.
It's also mildly funny to me that ~4 years ago AI was primarily used to make me laugh in /r/subreddit_simulator with shit tier Markov chains/GPT-2 and now we're mocking an AI model so cheap it can be served for free missed a decimal but otherwise was able to instantly answer a question with live information. In your food examples, yeah they're failing QC, but the food failing QC is significantly better than the failed QC food 4 years ago, the QC bar is a few orders of magnitude higher, and overall quality is skyrocketing year over year.
The companies are using the free samples as a sales come-on though -- like if the smashburger place were giving out shitty frozen burgers on the street as an ad campaign to get people to come pay for their good burgers, it would be weird, right?
Most normies are experiencing AI through the lens of Google's shitty free one -- which seems particularly shitty if you ask it about normie things. It seems like this should be anti-convincing in terms of driving adoption of the paid product.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
You know those posts and news articles where the author finds some Democrat lunatic or Republican influencer saying something stupid, to paint the entire party that way?
Unfortunately for you here, Google’s AI has about as much sense as a lunatic influencer. You’re better off asking a Markov chain.
I totally recognize that. But the primary point of my post was to highlight how vast the gulf is between the average Motte-poster's experience with AI and the rest of the world's experience. Part of the reason people like Scott Alexander keep spectacularly failing to connect with normal people on AI safety and related issues is that the vast majority of people's interactions with AI involve shitty free versions. With the exception of the absolutel worst and cheapest AI-assisred search and automated customer service lines, AI technology has penetrated into ordinary people's lives a lot less than you'd think just by reading here.
More options
Context Copy link
More options
Context Copy link
I get them every single time I use one for anything subjective or creative. I suspect LLMs are a dead end. The comprehension of prose is terrible, their ideas are terrible and wishy-washy, and the naked glazing disgusts me.
I suppose that biological neural networks were ultimately evolved in the same way; throwing them at a task (survival) over and over and over until they accomplished it, then iterating. It's why our brains arent truth-seeking machines, but survival machines, and even our sensory perceptions are a cobbled together mess of imaginary qualia and we see Shadow-Men and optical illusions.
Instead, these things are trained on convincing a human that the output pleases them; any accuracy or usefulness is ultimately incidental.
More options
Context Copy link
I'm amazed at how many people don't cough up even the $20 required for the full-fat experience. At least OpenAI has recently announced that they're making unlimited usage of a newer model, 5.6 Luna, available to the masses for free. It is still several steps down from Sol, which continues to genuinely impress me.
Google's "free AI" is an incredibly imprecise term. It can mean Gemini 3.5 (Flash? Lite?) when used through the Gemini app. It can be the search-specific Gemini 1.5, IIRC, used if you just "Google" something. As you can imagine, not even Google can afford to use the best of the best when serving so many people at such scale, with the majority being trivial queries.
I would pay a great deal more than $20 for an AI plan, but luckily, I don't have to. I've got tons of Claude Max, because I decided to accept payment in that form instead of figuring out intercontinental financial transfers after my participation in Unslop.
I remember when a Dutch junior of mine was consulting another senior for advice when working on a meta-analysis meant for a big name journal. I could have clapped like a seal when the latter mentioned Claude, in addition to ChatGPT, for help with the stats scutwork. Then I asked, and confirmed, that she only had a free plan. I was so dismayed that I immediately offered to share my Max, and walk her through things. It didn't hurt that she's a pretty girl, and it would have provided an opportunity to go out with her (that ended up happening anyway), but I am beyond annoyed when other doctors use free-tier LLMs for clinical work. They still function adequately, but anyone who can afford better should use better. $20 is really not an onerous ask for a significant amount of intelligence on tap.
I've got access to higher end stuff from work and I've found that Sonnet on the free plan is basically golden for most anything I want from it in my personal life. The main things I would get out of paying for a model aren't "better model" it would be things like better integration with my desktop files (since on the free plan you can't give Claude access to a folder) and longer uninterrupted sessions.
More options
Context Copy link
Why would anyone pay the $20/month? To be impressed? It wasn't long ago that someone asked what they're supposed to use AI for, and the response was a list of about a dozen completely frivolous things that can only be described as glorified low-grade entertainment. I'm sure some people find it useful for niche things. But for most people it's nothing more than a glorified search engine, and expecting people to spend money for it is like asking them to pay for a subscription to Google.
I am genuinely taken aback by this take. I know you're being sincere, but I can't understand where you're coming from. It's easier to say what LLMs can't do for me. And I very clearly benefit from the additional intelligence the paid models provide. They're not eating white collar work for no reason, they can do a surprising proportion of the work that a regular human can do on a computer.
More options
Context Copy link
Ignoring code (and writing/smut), non-exclusive, I've used Grok for :
Some of these might have been solvable with google (air conditioner motor is pretty much a process of elimination thing, although even there Grok was a lot better at finding a local seller than Google is). Some of them not: the plain language search for a crankshaft position sensor error was universally 'bring it to a shop', and I'm very bad at 3d modeling. I'm skeptical that they're all frivolous.
I've gotten a few more uses from Claude, though I'll admit it tends to be something I burn more code-facing or writing-facing.
More options
Context Copy link
If LLMs cost 20 $/mo for "low-grade entertainment", they're in the same league as Netflix (without advertisements, 20 $/mo) and YouTube (without advertisements, 16 $/mo paid monthly or 13 $/mo paid annually). It's my understanding that lots of people buy temporary subscriptions to streaming services, in order to test them out or to watch a specific platform-exclusive show.
More options
Context Copy link
More options
Context Copy link
In my ime, all cost optimized models are absolute worthless dogshit. A current gen cost optimized model is not as useful as early gen normal models (llama3 70b, gpt-4o).
My understanding is that current chatgpt free is 5.5 real with thinking=0 which is still a decent model. Luna (equivalent to nano, much worse than even mini), is a maaaasive step down. I think it's criminal to even let the masses have access to this piece of hot garbage, as it will provide negative information to a 100iq midwit.
This is true and I'm currently doing a project that has ABSURD unavoidable token read/write overhead (I'm processing ~300 cookbooks) and it takes 100 millions of tokens per book to ensure a reasonable level of accuracy.
So I'm basically stuck with subscription subsidized tokens, but Codex only serves the most recent models. I'd absolutely kill for a gpt-5 full size at low cost, but they don't offer the big old models at low cost (presumably bc expensive to run, but they're still much smaller than the base model of 5.5+) and Luna, while now gloriously cheap, is just so fucking stupid and myopic.
I could do this project with a 2 year old LLM most likely, it's just so deeply unergononic
More options
Context Copy link
Luna is incredibly useful if you have a task scoped right for it. Funnily enough, low effort seems to perform better than high/xhigh, because the latter keep trying to find a way to galaxy-brain their way around simple evaluation tasks.
More options
Context Copy link
More options
Context Copy link
Yeah, very recently decided to try the Claude Pro-tier AI for a month, and I'm very impressed. I did have access to Gemini 3.1 Pro through work, but it feels like a night-and-day difference.
We'll see if I still feel this way in a couple weeks, but for now I'm firmly in the "very, very useful" camp.
I'm actually a huge fan or 3.1 pro (high) personally. Opus and 5.x xhigh are definitely better at difficult problem solving but 3.1 pro can handle most things and is much more chill and less anal about many things.
Would be interesting to hear what specific tasks 3.1 pro underperforms on that opus can take care of
I've tried to use Gemini 3.1 pro to handle some basic coding tasks at work, and it does not do particularly well. I don't want to give exact specifics to avoid doxxing myself, but I can give you the shape of it. Skip to the most direct example if you don't feel like reading.
I have a java project. Currently it uses all-java libraries for image processing. Java kind of sucks for image processing because the integrated entry point always loads the entire raster into memory at once. If you have a limited heap, you're going to run into a world of pain trying to work with the standard library.
To get around that problem, I've looked at a native library that has java FFM bindings. Unfortunately, to use FFM, I had to upgrade to JDK 25. We had some test failures, and I figured that Gemini/antigravity could handle that kind of scutwork. It could not. It tried to make massive, architecture-level changes to two different subsystems in our codebase rather than just fix a classloader problem. Eventually I gave up and did it myself. I lost about a day to this.
After upgrading the JDK, I handed off the work to another developer to handle writing a small wrapper around the FFM library to unpack the native libs. She immediately tried to use Gemini, and lost four days to its confabulations. She came back to me repeatedly telling me it was impossible, and that we couldn't possibly do this on windows because Gemini gave her a trivially disprovable assertion. I eventually gave up and handed it to another developer who engaged his brain and had it done in a few hours, on all our supported platforms, with tests.
The most direct example: After that, I started converting one of our image processing routines to use the native lib. I figure that since the problem was easy and both the native library and the FFM bridge are both exquisitely documented, and they're both open source, this should be trivial for Gemini. Well, it turns out that both the native lib and the FFM bridge were both mostly written after January 2025, so neither 3.1 Pro and 3.6 flash consistently had the APIs inside their training window. It also didn't really have many examples to match on because this isn't a basic Python CRUD app.
You would not believe the absolute fever dream of a codebase it tried to cook up. It couldn't get function names right, and when it could, it couldn't consistently distinguish between the native library and the bridge. It frequently failed to even be consistently wrong. Eventually it finished, and it solved the problem by importing the FFM bridge but not actually using any of the underlying native calls.
I was a little disappointed.
RIP in pepeloni. But to be fair most agentic ai loves to do this unless you prompt it correctly.
More options
Context Copy link
More options
Context Copy link
To be honest, I don't know if I can compare them fairly. My access to 3.1 pro was clearly a budget option or something, because I would constantly be shunted back down to "Thinking" because of pro being in high demand.
I also didn't have a Claude Code-like option, but attempting to feed 3.1 pro my code worked for a single response, and then it claimed to be unable to remember it for the next question. Possibly, I could have kept giving the entire 5k lines of code for every question, but that seemed ridiculous.
In contrast, I pointed Claude Code at my project directory, typed "/init" and it very accurately determined what the purpose of the project was and all the specifics from just the code, test output, and poor documentation.
With that said, for smaller code questions 3.1 pro was certainly adequate.
Ohhh you got access to the gemini.com app not access to the api wrapper. Yeah the app sucks, idk how google screwed it up so much.
Also the gemini bundled with the $25 google workspace accounts is especially gimped, it's much worse than thr $20 plan that normal users get.
3.1 pro high is good in antigravity, aistudio, and api.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Not quite that extreme, but https://www.themotte.org/post/3791/smallscale-question-sunday-for-june-7/450717?context=8#context
More options
Context Copy link
Google's AI is embarrassingly bad.
I think LLM-tech is over-hyped, and I think this site's rationalist/EA DNA makes its commentariat hype it even above that over-credulous baseline.
Even still, it's probably unfair to judge the current state of the LLM frontier by whatever Google is doing. Their latest Gemini models are dogshit, and anecdotally represent a regression over previous releases. The fact that they can't even get a new pro model out the door reinforces my opinion.
My work is using it along with GPT for various tasks. Management loves Gemini for some reason. It can't even consistently summarize emails. It will hallucinate basic things like who said what, or even completely invert the plain statements of speakers. This doesn't do too much damage, because management tended to hallucinate a lot in the first place, so all we've really done is streamline their rejection of reality.
When I have tried to use it for coding tasks, it frequently shits the bed because it invents APIs that don't exist, refuses to use tool calls, and goes into schizophrenic reasoning loops. I use it because politics demand that I be a team player, but at this point I'm pretty sure I've already lost more time to it than I will ever save.
In contrast, the latest GPT models are... OK I guess. They still hallucinate once you go a few steps off the beaten path, but it's not omnipresent and endemic like it is with the Gemini lineage. I think it's insane that we've spent nearly a decade and approximately 50 Manhattan projects worth of money to accomplish that level of utility, but it's... OK.
Gemini is so bad. It's actually kind of mind-blowing how hard google fell off this year.
I haven't had a chance to catch up on the podcasts about it, but sounds like they did a massive AI shakeup last week
More options
Context Copy link
Poor old Scott Adams. This would make a great Dilbert punchline.
If Scott Adams weren't dead, the current business zeitgeist around LLMs would kill him.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link