This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Dan Ariely Is In The News Again
[previous discussion here, hat-tip to TraceWoodgrains for bringing it to my twitter feed]
Data Colada Reports:
For those who missed the first one, at least one of the studies in this data is fake. There's an orgy of suspicious problems with it: incredibly implausible reported numbers, ridiculously strong effect size, negative correlations for things that should have been strongly tied together, sample 'twins', inconsistent rounding where only three out of 360 survey values were not rounded to the nearest 5%, miscalculated statistics, means changing from a preprint without derived statistics changing with them. About the only thing missing is the text format mistake showing exactly what was copied and pasted in the Honesty trials; we should be glad there's not a spurious western blot that someone made its way in. DataColada doesn't even cover all of it: there's numbers in the underlying spreadsheet that don't match in the basic math sense.
There's also a bit of a buried lead, though. DataColada needed the raw data to do this analysis. That doesn't mean anyone needed that data to make a serious criticism. The errors here are not subtle. You can tell there's something fucky wucky with this paper just by looking at its charts, and there's only one for this study.
The experimental protocol described people receiving three papers with a hundred errors each, and were told that they would receive 10 cents for every correction, but would be penalized a dollar for every day they were late.
(errors and delay calculated from the underlying source data, but nothing within a round-to-10% value works, either, so this fails the MK1 Eyeball Test. These study's test involve machine-generated papers intended to be meaningless, so a best-case scenario of 70% errors found, while suspiciously even, is at least plausible even for MIT's best and brightest.)
The chart isn't, to be clear, close to a possible combination. The chart doesn't even have a negative value reported, period. And if you start thinking about what a negative value for earnings means -- any further work would have no value, and soon no plausible expected value -- and quickly many the numbers don't really make sense. At best, the trial used a different protocol than its own minimal results show, in a form clearly visible by just reading the paper. A flat rate for participation isn't impossible, but if it was included in the measurement, it should have been in the methods section. It wasn't.
But now we actually have e-mail conversations showing that wasn't the case, either: "There was no show up fee and we did not say anything about the fact that the payment rule could mean that subjects will lose money (and as far as I remember we never had to deal with this problem)". That's for a trial where, by the published protocol and data, a full 19 out of 60 participants had negative earnings, and even assuming a flat $10 participation fee needed to make the chart work, 8 participants would still have had negative earnings.
A fellow academic, Kyle Hyndman, received that e-mail in 2014.
Hyndman, to his credit, did forward the data and conversations to DataColada in 2023. And he was working on a replication attempt most of the time, albeit probably as a low-priority for a decade. He did, after his replication effort, publish in a footnote that : "In October 2024, at the request of the editors, we shared with Dan Ariely an analysis of the contents from the file purportedly for their Study 2 and asked for permission to include a summary of it in the paper. Dan Ariely denied our request, arguing, among other things, that the files we received may not be the actual data."
That does seem to be the defense, for another way this study rhymes with past scandals. Ariely can't remember the actual study protocol, can't be sure this data was what actually got published, and doesn't want a summary of the data that does exist being republished. It is quite possible no records of the experimental protocol exist to prove or disprove anything. I've got a search tool running through old MIT classifieds, and that's more a hope than a process.
There's a plausible, if disturbing, option that had the data fabrication been revealed before the replication failure and the separate scandals in other papers, it would have blown over. There's a plausible, if even more disturbing, option that even with those other data points, it will just blow over again. Maybe a retraction -- and to be fair, there's at least been a request for one -- probably Ariely doesn't get another tv show, probably not a slapped wrist, almost certainly no serious investigation by his school.
But that still leaves a lot of questions. Ariely has over a hundred other papers he's authored or coauthored listed on his own CV. It was plausible for the Honesty trial that no one else should have looked at Ariely's data, or looked at his conclusions with a skeptical eye. It remains plausible that, for two papers, his coauthors and peer reviewers didn't look at the paper with a skeptical eye. Two isn't feeling like a very real number here.
Trace mentioned this story with the opening line "I feel for the coauthor [Wertenbroch] here – he seems to have acted generally honorably and been caught off guard by having a liar for a collaborator." Before I read the paper, my gutcheck was "On one hand, it's nice to see Wertenbroch moving on it. After Gino, I'd seriously worried about the possibility Ariely had just been able to thrive so well because no one he was working with wanted honesty, either."
And then I saw that chart, counted on my fingers, and blinked.
Wertenbroch is, notably, CC'd on that 2014 e-mail where Ariely said that his study did not have any students with a negative balance, and had no participation fee to explain that 10 dollar offset. He could, plausibly, have not reviewed the data or the paper in a decade, and not realized the discrepancy even then. Leaves a bit of a question about what he does do, though. He could, plausibly, have not touched or looked at a single part of the study methodology or data in this entire paper.
In the Honesty trials from the previous scandal, there was an absolute mess where Ariely was responsible for Study 3, and there were serious questions about what, if any, exposure to the fraud the other authors might face. Gino seems to have had minimal responsibility of exposure to Ariely's original data... and separately produced some falsified data in a separate study in the same paper, and multiple other studies in other papers.
DataColada ends today's post with "In our next post, we will share analyses of the Study 1 data file [...]. That experiment is quite different. Our analyses are quite different. But our conclusions are quite similar." I'm working on writing up a post on the aftermath -- or lack thereof -- on the Hindawi scandal.
Eat at Arby's.
EDIT: while DataColada does not spell out the chart discrepancy, it is in their underlying R code. So the description here is more a dumbass noticing the same thing, not me finding something they missed.
Part 2 is funnier and more damning, somehow, despite there at least being no question of the study actually having happened.
In his defense, Wertenbroch does seem to have genuinely and enthusiastically provided information that could be plumbed to show the manipulation. Still not sure what he did on the study, but I guess that's a good thing for him right now.
More options
Context Copy link
I'm starting to think that Futurama was rather prescient in calling the turn of the century from 20th to 21st as the Stupid Ages. More and more of the facade of credibility of organizations, institutions, people, ideas, etc. that people used to believe were not stupid - justifying using them to guide their decisions to be less stupid - have revealed themselves to be deeply stupid. I hope that LLMs will become so cheap and reliable that all the previously hidden stupidity will be aired out for all to see soon enough. I just dread what unexpected stupidity will be found - sociology, [x] studies, psychology are the obvious no-brainer candidates which have already largely gone through the process, but what if it turns out even very "hard" or concrete fields like electrical engineering or chemistry are built on a bunch of stupidity that has just enough facade to be convincing to the in-field expert and certainly more than enough to the layman?
Most institutions become incoherent because of Goodhart’s Law, so as long as we disprove that one first, we should be fine.
More seriously, you don’t hear about all the industrial processes and social tricks that work on a day to day basis. Selection bias tells you about the factory which fails and not the ones which are chugging along according to the boring sort of science.
More importantly, reversed stupidity is still not intelligence. Disregard the institutions insofar as they have been Goodharted. Doubting the boring ones with no social status attached is how you become a crackpot.
More options
Context Copy link
Longtime readers will know what my take on this is.
The institutions in question have not been held accountable for outcomes in decades, don't hold their internal staff accountable, and oftentimes don't measure outcomes at all. The Institutions generally in charge of holding people accountable and punishing the misdeeds are, themselves, usually compromised as well. I suspect everyone prefers an arrangement where the institutions shamble along based on their prior reputation, everyone getting paid to maintain a facade, than risk collapsing it by exposing the extent of the rot. To the extent there are punishments and consequences they're rarely sufficient and proportional enough to truly dissuade the behavior in question.
And finally the institutions that manage to maintain integrity eventually get outnumbered and even if they are able to hold some people to account, can't possibly keep up with the volume of the problem. If that becomes widely known, even average people will be inclined to cheat/defect since not doing so makes you a bit of a sucker and there's rarely true consequences for doing so.
Re-insert skin in the game.
The good news is that in such hard fields, nature itself tends to have a correction mechanism for this, where the feedback loop for doing something wrong is tight and the punishments can be harsh.
Synthesize the wrong chemical, the reaction won't work, and you possibly kill yourself. Design the plane wrong it crashes, eventually. Reducing your safety margins too much will eventually come back to bite you, and possibly remove you from the system, and provide a warning to the next guy.
I'm sure there's stupidity to be found there, but when the real world REQUIRES you get things right in order to actually succeed, the accountability is to some extent baked into the cake, even if various institutions will try to put a few layers of obsfuscation between themselves and the outcomes.
Not sure if Physics still counts as "hard or concrete" field, but a lot of attention in the last few decades was spent on (mis-)adventures where the real world won't care one way or another. Of course, the fact that the world doesn't care also means that most people tuned out a long time ago, so at worst some people will have an egg on their face in a forest where nobody is around to notice.
More options
Context Copy link
That's exactly why if all that turned out to be a facade, it would be absolutely wild and crazy and turn my world upside down. I'd guess that the odds are low, but also, I'm quite sure that the true believers in those already-discredited fields would also judge the odds as being low for whatever field they're a true believer to.
Yeah solid point.
If it turned out that foundational papers in Engineering and studies in chemistry were false and everyone else had just been winging it and lucking out (or maintaining a massive coverup) then I would have to re-examine a lot of my core assumptions.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
No matter how stupid people are today, I'm sure they were dumber pre-1900, especially before we normalized hand-washing. Although if you want to call the entire history of humanity "the stupid ages" I'd agree to that.
I'd be surprised if these didn't have subtle errors that will be corrected later, but they (even the slightly incorrect old models) have already been replicated and extrapolated into new discoveries, unlike much of psychology, so I wouldn't call them "stupid" and certainly not useless.
Pre scientific people were often dumb in believing incorrect things, but they believed these things largely due to the limitations of their methods and observational equipment. Modern beliefs which are incorrect are intentionally so, lots of these beliefs are literally supported by changing the definitions of things when statements about them have been falsified.
Aristotle believed that women naturally had fewer teeth than men because he was unable to check the chemistry of nutrient flows in pregnancy that led to tooth loss. He believed that objects in motion gradually come to a stop because he had no way of measuring air resistance. He believed women’s mental illnesses were caused by their wombs moving out of place because he had no way to check hormone levels. People believing falsified things about gender and sex being meaningfully different can check those things, they know they’re wrong, because they have checked this. They changed the relevant definitions instead of updating their beliefs.
Pre scientific people were intentionally mislead by smarter ones too: mystics, false religions, etc. And science is sometimes like mystic religion, with incorrect beliefs perpetuated by dogma, but I think it’s more accurate than those of the past: exposed papers do get retracted, my albeit basic understanding of mysticism is that there isn’t really an analogue “oh our beliefs are wrong now that outsiders have actually questioned them”.
We have mystics and false religions even today. Woo is still a thing.
More options
Context Copy link
The difference in my view is that mystics, false religions, etc. didn't provide some structured, well-established method at arriving at facts and the truth the same way science does. Generally, mystics, to whatever extent they believe that their mysticism gives them access to the truth, doesn't purposefully corrupt their mysticism in an effort to "discover" that what is true is everything they wanted to be true all along (arguably the mysticism itself accomplishes that, which is one part of what makes it different from science). Right now, the people purporting to do science are doing exactly that, despite having all the tools and knowledge and information required to understand that they are indeed corrupting science and also how not to corrupt science. That's what makes their stupidity extra stupid.
Like how someone with training in arithmetic and calculators who declares the calculator as a tool of White Supremacy and therefore 12 x 12 = 1212, not 144 like what the calculator said, has a sort of stupidity that goes beyond the stupidity of someone who just never had access to mathematical education and declares 12 x 12 = 1221. 1212 is slightly more accurate because it's even like 144, not odd, and also closer to 144 than 1221 is, but the process by which the person arrived at that was arguably more stupid.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I think it could make sense to call the ages the Stupid Ages if the cause of us being stupid was the very thing that was for the purpose of addressing/avoiding/escaping stupidity, which is sort of what's happening with academia, which are ostensibly for the purpose of something like addressing stupidity, even if, overall, we're less stupid than any other age. There's something to hubris that is extra stupid than just being plain old stupid. Especially since hubris is, itself, a concept that is taught in academia as something bad and to avoid.
As flawed as academia is, I still think we'd be stupider without it. For example, imagine if we didn't have the psychological fallacies (although even today people fall for them way too frequently, imagine if you couldn't even point them out because they weren't established, and AFAIK these haven't been debunked).
I have seen exactly 0 people convinced anytime anywhere by an argument via someone else saying they are committing a psychological fallacy. It's nice to have them labeled but I'd put the real world impact as pretty de minimis
Even that might be overly charitable. If you start with someone who is arrogant and partisan and teach that person various fallacies, quite possibly he will become more arrogant, more confident in his partisan opinions, and less willing to consider contradictory information (because he is better at spotting arguable weaknesses in his opponents' arguments).
More options
Context Copy link
More options
Context Copy link
I think you're right. But:
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Your phone works, doesn't it? Everything from the 5G radio protocols through the processors and OLED screens here isn't trivial and relies on a lot of in-field expertise that can demonstrably be taught, but isn't obvious to the layman.
If it's a facade, it's one that is at least very functional and works reliably day in and day out. Sociology isn't demonstrating it's functional correctness billions of times per second on each of billions of devices.
Right, and imagine just how you'd react if it were proven to you that this was indeed all a facade? My model of the world doesn't have much room for something like that, where something that works literally quadrillions of times a day in basically exactly the way that all the scientists and engineers and experts and my own limited expertise say they do, and yet all of those are complete lies. This would be a revelation that comes close to - not quite, but close - to waking up in a pod and discovering that you had been in The Matrix.
Significantly, I'd wager that the belief by the typical non-stupid layman believer in those already-proven-to-be-stupid fields like sociology, [x] studies, psychology, etc. is about as strong and about as confident as mine is to electrical engineering. It's both reinforced by and causal to the constant defense such people run against scrutinizing such things with basic scientific/academic concepts like "evidence" and "logic." And if those people can be fooled into such stupid beliefs about sociology, then I could be fooled into an equally stupid belief about electrical engineering.
This indeed had happened before and it doesn't matter from engineers' perspective? It was proven that classical mechanics is indeed a facade, most of the time in general usage it is mostly correct, but the calculation result is different to result from general relativity
But billions of existing functional cars also means that we can still use classical mechanics for approximation, until you run into satellite related calculations
This doesn't capture what I'm talking about. Classical mechanics is a very good approximation for relativity when dealing with mass and accelerations that humans generally tend to deal with in everyday life, which is why it was - and still largely is - considered "correct" despite necessarily being inaccurate to the real world. This is very much not the case for the actual IRL examples, where actually studying the underlying science reveals that the things that people are saying don't require mere refinements to get at the truth, but rather have basically no relation. It'd be akin to the discovery that, actually, it's impossible to store or release energy in molecular bonds like in oil. Imagine that were proven beyond any doubt to you, and yet all the world still functions and runs the way that you've perceived it all your life. That'd be utterly mind blowing.
More options
Context Copy link
More options
Context Copy link
I don't know about electrical engineering, but at least in math, I expect it's far easier to change someone's mind with evidence than it is in a social science. Even with informal proofs, there's a greater shared understanding and trust of axioms that, when put together, invalidate (or confirm) an argument.
Occasionally a math paper gets retracted for errors, but it seems far rarer than psychology, and with much less controversy. Meanwhile, many theorems are getting formalized in Lean, effectively "replicating" the proofs, many are successful and the vast majority are successful or pending.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
LLMs are one of the things making people stupider (and are themselves very stupid), so you shouldn't hold your breath on that.
I think LLMs making people stupider doesn't have much relation to LLMs' ability detect past stupidity that had been sufficiently hidden via exploiting human biases. As long as it doesn't uniformly make all the smartest people stupid (possible, but I think LLMs likely have an effect that's bifurcated on smart/stupid people in terms of their intelligence, if at all), some smart people will exist who can direct the LLMs towards detecting and revealing that stupidity.
I suppose the depressing possibility (likelihood?) is that the populace will be so stupid due to LLMs that all the revealing-of-stupidity just makes the smart people feel like they have no mouth and they must scream. Probably a better apocalypse than the AM takeover? In any case, to avert the apocalypse, perhaps the thing to do is to accelerate the revealing of stupidity and discrediting of these institutions even more, faster than LLMs can make everyone stupid (though we may already be too late on that, even in the pre-LLM times).
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Some fun additional thoughts:
DataColada is going to focus on the data manipulation, because the proof there is undeniable and they don't need to make a 90% bet.
More options
Context Copy link
Scott found this earlier, by comparing papers between known pseudoscientific fields and soft sciences.
Do you remember where? I remember him describing Ariely et all as under challenge very early, and I don't remember too much predictive beyond a book reviewer calling him a general fraud.
I believe it was the famous The Control Group Is Out of Control where he compares quality of psychology and parapsychology papers - where parapsychology studies phenomena like telepathy etc. He deliberately used parapsychology as a control group, as the modern papers use the same methodology as psychology papers. And unsurprisingly, parapsychology fares quite well. If it was not for the fact that parapsychology is so suspicious of a field, these papers would be part of our scientific arsenal and would inform host of policies and self-improvement programs as the psychology does.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Together with the Arday scandal, social science is really having a bad time right now. The killer is that these are not just random researchers, but professors, and not just professors, but celebrated stars with multiple books, regular media appearances and political contacts. They are likely to impact actual policy and general large-scale decision-making. And to add insult to injury, it's getting increasingly obvious that a) competent fabrication would be undetectable in practice, especially this much later and b) nobody actually cares enough to check in general. They were just both incompetent enough and people pretty much stumbled over it.
I'm unsure whether social science in its current form is salvageable. It comes down to feedback loops: Math has a direct feedback loop with reality through the structure of proofs. You can set up implausible assumptions, yes, but since you have to state them clearly you can always be called out on it. And it's still relevant information that, if the assumptions are true, it follows whatever you have shown in the proof. In physics, you usually can turn a model into predictions which will eventually be measurable, even if through rather elaborate channels such as measuring red shifts of far-way celestial bodies or by ramming very small particles into each other.
And so on further down the line; My own field, medical genetics, is a lot worse than those as well but we also have at least some points of connection with reality; IF I find a variant that is plausibly connected to a serious disability, and IF there is a medication that can effectively bypass that variant (for example, producing the protein that became malformed), we will usually see a noticeable positive impact on patients, even if it does not fully cure them.
But social science? Almost any result can be interpreted in almost any way one may choose; Surveys and closed, short-term experiments are possibly more connected to social desirability bias than to anything else, real-life cohorts are probably mostly selection effects, and controlling for them is usually somewhere between illegal and impossible. There is not zero good research, but it's completely drowned out by the bad, and even where people attempt to do it well, it's hard to say; For example, I like discontinuity design, and it imo works well for entry tests (barely made it into an university vs barely didn't make it), but exactly the same design for school area impact is probably nonsense bc people can deliberately move into those neighbourhoods so barely inside vs barely outside of the area is again mostly selection effects.
In such an environment, and given considerable pressure on them to deliver convenient results (again, unlike math and most of physics) which also results in substantial rewards outside academia, social science predictably devolves into near-pure high-school style popularity contests.
Social science research done well is not going to reach the levels of hard science, of course, but it is certainly not worthless if done well. There are various strategies to establish validity, reliability, etc. in quantitative measures, and credibility, transferability, etc in qualitative research. The problem is ideological capture, wrongthink, and institutional guardrailing. As well as a general lack of rigor, even when the methods to maintain such rigor and force accountability exist.
Can you give us some examples?
It will all depend on what you're measuring and how you're measuring it. Quantitative and qualitative analyses are two completely different paradigms (though they are often married in mixed-methods research.) They don't even use the same vocabulary, and the product of their aims is fundamentally different.
A quantitative look at divorce, for example, might have a questionnaire as its main instrument. You get a bunch of numbers and it turns out the main predictor of divorce (note: not cause. Predictor) is contempt (as one of my friends likes to say whenever the topic comes up--I think he read something by Gladwell.) A qualitative study wouldn't have a questionnaire at all. You'd probably use interviews as participant observation of a marriage would have intrusion and ethical issues involved. You might end up with several narratives where you have a man recounting how his wife would always take her phone in the bathroom and stay far too long, as if sending messages, or a wife would talk about how the husband suddenly started washing his own clothes when he got home. Or whatever. The write-up for each would be totally different, and you'd be learning and digesting different things. A lot of people hate qualitative research. It's my main choice. It can be done very, very sloppily, however. I'll get to that in a second.
Back to making a quantitative questionnaire--if that questionnaire has not been, for example, piloted for content/face validity (via relevant experts who know the topic, but also item correlation/factor analysis) its results are more or less just noise (You also cannot use, for example factor analysis on just any set of questions. There are various requirements [a reasonable sample population, get rid of outliers, variables have to be on I believe an ordinal/interval scale, there has to be some degree of correlation in variables, you really need a normal distribution, etc.] If you don't fulfill these prerequisites again you're producing noise.) Once you have an instrument (questionnaire/whatever) that is relatively polished you administer it, but then in order to determine whether your results will be generalizable you'll need an appropriate sample size (and even then depending on the randomness of your sample your results will only generalize to that population, not the entirety of humans in all cultures).
Research design is also important. What you want to measure will determine how you measure it. I do not mean that you cater your plan to find a conclusion you want. I mean that if you want to measure something you need the right tool. Statistics do not show causation. A research design using statistics can establish causation (as in the numerous studies on cigarette smoking and cancer.) There are also criteria establishing causation.. The old saw correlation does not prove causation is of course true, though often positive or negative correlation is the smoke where there is fire. But not always, or even most of the time, in my opinion (which is worth about five cents.)
There are numerous threats to quantitative validity. Regression to mean, compensatory rivalry (in the case of two groups), compensatory equalization (when a researcher interferes), history, maturation, etc. etc. I highly recommend the book The Research Methods Knowledge Database (<- this is a US Amazon link, which I changed from the Japan link, but I cannot recall where you are.) I am not sure how much of this book no longer holds but it was a great read when I first went through it.
Qualitative research (what I do) can seem wild and woolly to an outsider, as if it's just you talking to someone and writing down your thoughts. And that is probably not a bad description of it, though there are, again, all sorts of rules you need to follow, primarily but not only to set the groundwork for reproducibility. And that, as we have seen, is a problem in both types of research--the findings are not reproducible. Documenting exactly what you do at every stage is very important for this. One study requires mountains of documentation. If you have any interest in qualitative research (and I would say the median user on the Motte does not) I would recommend the book Naturalistic Inquiry.
I have probably not answered your question, but I have let this sit and I didn't want to write something fast with my thumb. And you may know all of the above and more. You also may not find any of this compelling. I am not particularly good at statistics, though I used to use SPSS with some degree of expertise when I was made to do so. I believe that software may be outdated, or at least not used as much as it was. The Bayesian wave of the last 20/30 years has also changed a lot and I am out of the loop. Note I have known many researchers in social sciences, most of them what I would consider well-meaning ignoramuses, some of them outright frauds, this in both quantitative and qualitative. I have known only one man who I thought was a genius psychometrician; he didn't rely on the software as he knew the formulas and why they were the way they were. A math guy, like many of you (maybe not you-you, but the general Motte You.) He died far too early of a brain tumor in one of life's cruel ironies.
I have read a lot of shoddy research. Far too much. Enough to at least doubt (if not completely dismiss) every study that I see, most seriously when the study confirms my own biases. But then I think that's the right way to live. I could be wrong.
edit for clarity and this is still not very clear. My parentheses runneth over.
More options
Context Copy link
Standardized methods is one way. This at least lets you compare between studies that use the same methods, and thus determine if you find the same thing. The best methods have clearly defined scopes that indicate what they are and are not looking for. So the authors should be well aware of limitations well before starting their studies.
You are also supposed to be aware of your biases and validity threats, and admit to them in your writing so that the reader can easily understand your limitations.
It's just that a lot of researchers don't do this. I remember reading feminist papers, cited by journalists and used in political arguments, that make conclusions not supported by their results or where the methods are clearly designed to create a bias which is never admitted. The opportunity for rigor was clearly there, but the authors and reporters decided to forego it.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
For a while, the String Theory folks got a lot of funding, but never seemed to deliver practically testable predictions. I'm not sure this is as true in physics as you'd like it to be, but it is certainly moreso than the humanities.
In fairness, doesn’t theoretical physics deliver lots of results that require either exotic technology or bajillions of dollars to actually test?
The complaints I recall hearing from the less-theoretical physicists is that the string theorists were taking up a large share of the available funding effectively
navel-gazingdaydreaming about the pretty symmetries of strings without actually producing testable hypotheses even at "LHC" scales of "testable". Those folks are used to expensive tests, but the string theory results needed like 3-4 or more orders of magnitude in scale.More options
Context Copy link
More options
Context Copy link
I actually had the string theory folks on my mind while writing this, whether I should put a caveat there. There's always a "but". The key to me is that with almost everyone I've talked with who is sufficiently familiar with physics, they mention string theory, and they always clearly acknowledge their lack of predictions is bad. At least some number of those theorists tried to find something testable. And as you imply, it seems to be on the downturn in funding (unsurprising, the EU tries to pick up the slack somewhat, they certainly have impeccable timing). Altogether, that's good enough for me.
More options
Context Copy link
More options
Context Copy link
This is unfortunately a merely aspirational statement. The practice of mathematics involves writing proofs in a heavily intuitive manner, and their verification in turn involves people who have been socialised to share the same intuitions, which is a process that rarely involves proper reduction to the axiomatic basics but more often looks like "your elders and betters assert that it is trivial and look at you with mild disappointment, and all your peers seem to already have gotten it, so get on with it and flog your brain into producing the 'this is obviously true' qualium already". There were a handful of examples where local cultures/status hierarchies perpetuated a body of wrong mathematics this way, such as the "Italian school of algebraic geometry" and more recently (most likely) the Mochizuki abc conjecture incident.
We can be quite glad that it has not yet happened that the lines of a wrong intuition-subculture have aligned with general tribalism yet. If Interuniversal Teichmüller Theory were an invention of the likes of Arday, we would be seeing its detractors decried as racist and proper mathematicians performatively weaving it into the body of accepted and otherwise sound mathematics. Even if this did in fact cause further downstream inconsistencies to open a path to excision and repair, discovering those is (disproving abc)-complete, and that's something we have tried and failed for a long time.
One of the best professors I had would sometimes in his lectures go "Proof (1. Attempt): [...] and therefore the proposition is proven ... wrong! The proof I gave had a flaw! Can anyone see what the problem is?" When no one could, he explained the flaw, gave a 2. Attempt, which fixed the flaw, but there was another unfixed flaw, and only the third attempt was the real proof.
He also constantly asked questions to the students, he was really committed to teaching actual thinking.
He also held a lecture once where he gave no sample solutions to the exercises. The exercise was always supposed to be the students discussing what they came up with, with the assistant only supervising, and the exam was just a selection of the exercise problems. One of my prouder moments was when my solution was simpler than the one the assistant had in mind, and I was eventually able to convince him it was correct.
More options
Context Copy link
I'm struggling to understand what you mean. At this point we have systems for automatic verification of formal proofs. In what sense are those systems socialized to share the same intuitions?
Those systems are not, but automatically verifiable proofs are still only a tiny sliver of frontier mathematics. Maybe AI can fix that, but AI can fix social science too. Just require that credit for AI research always has to go to some human, allocate token budgets by protected class, and you can ensure the Ardays stay ahead without requiring them to lie.
More options
Context Copy link
The intuitions required for formal proof verification are "it's worth rewriting this proof in another language that's an order of magnitude more verbose just so you can get a computer to tell you what you're sure you already know" and "the verifier doesn't have any soundness bugs that will invalidate any of its verifications" Historically the latter has sometimes been untrue and the former has usually been untrue. LLMs are fixing the first problem, which is exposing and leading to fixes for the second problem, but that's all a relatively new development. It's going to take a while to go back through all the mathematical literature and see how much of it might have problems due to predating the coming era of formal verification.
Consider the recent construction of a complex structure on the six-sphere. The simplest English explanation I've seen is under a thousand lines of (admittedly difficult!) writing and mathematical notation. The Lean formalization is a quarter-million lines of code. It's ... probably correct, people seem to think? But if it is correct, the construction would (reportedly; this is outside my field and way beyond me) contradict a 2020 paper (which itself was a correction to a 1998 paper), so we're almost certainly either producing new broken proofs or revealing old invalid proofs here.
This is overstating things, right? The length of the proof is not important, it's the soundness of the axioms and the complexity of the statement itself. I haven't inspected the proof, but I expect they're just using the standard lean axioms and the statement is much less than a quarter million lines.
For it's validity, yeah. There's something disquieting about getting proofs "straight from The Necronomicon" rather than "straight from The Book", though.
The full statement of the problem had better be much much less than a quarter million lines; one of the unavoidable ways to screw up a formal proof and still have it pass verification is to make a mistake in the problem statement and so end up proving something other than what you thought you proved. Nobody's going to check a quarter-million-line problem statement for misstatements.
(But I don't think that was an issue here - looks like everything they needed to define the problem was already in Mathlib)
The proof, though? Follow that github link, download and
wc -l Solution.lean: 248818 lines.It will be interesting to see if these formalized proofs can be "golfed" into smaller and more digestible forms by LLMs.
In that case you are really just trusting the Lean kernel rather than anything about the proof or problem statement. Doesn't seem that bad!
Very much so. I'd be surprised if LLMs didn't turn out to be excellent at this, and I'd bet that the Navier-Stokes blowup in particular turns out to have a much simpler example or at least a much simpler proof for this example, once we turn AI loose with the goal of "find the best result you can" rather than "get to a result that lets us declare victory ASAP".
This is something of an anti-inductive problem to me - if people really wholly trusted the Lean kernel I would think that trusting the Lean kernel really was pretty bad! But people are trying (and succeeding, so trust is still a little bad) to find kernel bugs, and writing partially-independent proof checkers for Lean too, so even if it's only 90% up to the task of being The Dependency for all AI-derived mathematics I'm confident it'll be hitting 100% soon.
I'd say 99%, except that I don't think modern AI is well-aligned enough or even well-instructed enough yet for us to overlook the fact that this is something of an adversarial process; a model willing to commit felonies to complete its task is probably also willing to exploit a 0-day Lean bug rather than report it...
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Sure, you can fuck up everything. As noted in the wiki entry (presuming it's correct), there was no problem with formal proofs; The problem was first going for informal arguments (the start of the slippery slope, though as so often it worked fine-ish for awhile) and then lowering the standards until anything goes. I also think it's notable that the majority of mathematicians has so far refused to accept the Mochizuki proof until it's understood more widely. Both were also mostly localised incidents.
Admittedly I've always been on the more informal end of math myself as well (back when I did research in a field that could be described as mostly-theoretical math), but even there I could simply double-check my conjectures with simulations, which is just yet another way of a feedback with reality.
I think you underestimate the power of proofs. If you make a conjecture, try to write a proof, and not just fail, but you find obvious counterexamples, there just isn't much left to salvage. Or the other way around, if you find multiple different proofs to arrive at the same conclusion and nobody you meet can generate any counterexample or find a clear hole in the proofs. It's true that there is some social element, it would be silly to claim otherwise, but proofs aren't just arbitrary arguments at triviality, either.
I'm familiar enough with the practice of maths (being an actual researcher in adjacent-enough TCS, and having had one foot in extremal combinatorics through a big part of my toil towards the degree), and I think you really overestimate the rigour of proofs as usually written up. The incidents I mentioned are for sure geographically localised, but in the case of the Italians it's not like the situation was that they believed it, and everyone else disagreed; for the longest time it was instead that they believed it, and nobody elsewhere cared but probably would have assumed that their fellow mathematicians knew what they were doing if they had to build upon one of their results.
There are other examples where, due to insufficient rigour of proofs as allowed and encouraged in the field, wrong proofs stood for quite a while. This was the case for like 10 years for some of the earliest claimed proofs of the 4-color theorem (an undergrad assignment actually had us find the counterexample to the core lemma there), more recently for some pretty celebrated random matrices "result" by Avi Wigderson et al. (hardly nobodies!) which stood for several months, and when writing a summarising essay on some cluster of Erdös papers I learned about a key lemma that was straight up wrong as stated, but people in the community thought essentially "something like it is clearly true, the core ideas of the proof are fine, and it's okay for the situation where we use it" (no proof of any of those things committed to paper). Correction eventually happened in the cases we know about, but can we really rely on the mechanism? Mochizuki has intimidated a big part of his department to ignore the refutations, and Japan is hardly a backwater, and there are political forces out there that are stronger than "this guy is one of our country's best, we should stand by him".
More options
Context Copy link
More options
Context Copy link
Man, I remember struggling with mathematical proofs in University. I kept just not getting it; it was never explained what made a valid proof. I eventually just learned to intuit what my professors expected, barely well enough to make it through, but I never truly got it.
A valid proof in the mathematical sense is simply one where every single step is justified by either an axiom or an already-proven theorem. You see some of these in a good Algebra 1 class where just going from "z=2(x+y)+x" to "z=3x+2y" uses left- and right-distributivity once each, associativity twice, commutativity once, and identity once. Then if you're lucky you never see them again, at least until you're using a formal proof verifier that actually wants a complete proof, yes, even if it's a hundred times longer. What's in the literature instead are "proofs" where we skip five or fifty steps at a time because we know our audience will be able to fill in the obvious bits inside those gaps between the big checkpoints, even though the definition of "obvious" depends on the audience and is famously a little slippery. So don't feel bad. Even when your audience is a professor, "obvious" still isn't a well-defined adjective and everyone is still stuck trying to intuit it. At least with theses and papers your advisors and reviewers will come back and say "step 23 wasn't clear enough; fix it". With homework or tests requiring proofs, professors' only options are to let it slide or to mark you down if you skipped a step they don't think was obvious.
My favorite variation of this has the professor saying, "It's obvious, but it's not obvious that it's obvious."
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link