site banner

Culture War Roundup for the week of August 31, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

4
Jump in the discussion.

No email address required.

Dan Ariely Is In The News Again

[previous discussion here, hat-tip to TraceWoodgrains for bringing it to my twitter feed]

Data Colada Reports:

A new paper in Psychological Science reports a failure to replicate Study 2 of Ariely and Wertenbroch’s influential article entitled, “Procrastination, Deadlines, and Performance: Self-Control by Precommitment.” The original study, published in Psychological Science in 2002, found that people performed better on a set of tasks when each task had its own externally imposed deadline than when people set their own deadlines or faced a single last-day deadline for all tasks. The paper has had a lasting influence. It has been assigned reading in many economics and psychology courses, and has more than 2,100 citations on Google Scholar.

Because this paper has been so influential, it is worthwhile to take a close look at the original study to try to understand why it did not replicate. We did that. This post – and the next one – is about what we found.

For those who missed the first one, at least one of the studies in this data is fake. There's an orgy of suspicious problems with it: incredibly implausible reported numbers, ridiculously strong effect size, negative correlations for things that should have been strongly tied together, sample 'twins', inconsistent rounding where only three out of 360 survey values were not rounded to the nearest 5%, miscalculated statistics, means changing from a preprint without derived statistics changing with them. About the only thing missing is the text format mistake showing exactly what was copied and pasted in the Honesty trials; we should be glad there's not a spurious western blot that someone made its way in. DataColada doesn't even cover all of it: there's numbers in the underlying spreadsheet that don't match in the basic math sense.

There's also a bit of a buried lead, though. DataColada needed the raw data to do this analysis. That doesn't mean anyone needed that data to make a serious criticism. The errors here are not subtle. You can tell there's something fucky wucky with this paper just by looking at its charts, and there's only one for this study.

The experimental protocol described people receiving three papers with a hundred errors each, and were told that they would receive 10 cents for every correction, but would be penalized a dollar for every day they were late.

Condition Avg Errors Found (+$0.10) Avg Days Delay (-$1.00) "Earnings" (Errors / 10) - Delay (Errors / 10)
Evenly Spaced Deadlines 136.1 3.55 approx 20 10.06 13.61
Self-Imposed Deadline 107.3 7.9 approx 12 2.83 10.73
End Deadline 71.1 12.75 approx 4 -5.64 7.11

(errors and delay calculated from the underlying source data, but nothing within a round-to-10% value works, either, so this fails the MK1 Eyeball Test. These study's test involve machine-generated papers intended to be meaningless, so a best-case scenario of 70% errors found, while suspiciously even, is at least plausible even for MIT's best and brightest.)

The chart isn't, to be clear, close to a possible combination. The chart doesn't even have a negative value reported, period. And if you start thinking about what a negative value for earnings means -- any further work would have no value, and soon no plausible expected value -- and quickly many the numbers don't really make sense. At best, the trial used a different protocol than its own minimal results show, in a form clearly visible by just reading the paper. A flat rate for participation isn't impossible, but if it was included in the measurement, it should have been in the methods section. It wasn't.

But now we actually have e-mail conversations showing that wasn't the case, either: "There was no show up fee and we did not say anything about the fact that the payment rule could mean that subjects will lose money (and as far as I remember we never had to deal with this problem)". That's for a trial where, by the published protocol and data, a full 19 out of 60 participants had negative earnings, and even assuming a flat $10 participation fee needed to make the chart work, 8 participants would still have had negative earnings.

A fellow academic, Kyle Hyndman, received that e-mail in 2014.

Hyndman, to his credit, did forward the data and conversations to DataColada in 2023. And he was working on a replication attempt most of the time, albeit probably as a low-priority for a decade. He did, after his replication effort, publish in a footnote that : "In October 2024, at the request of the editors, we shared with Dan Ariely an analysis of the contents from the file purportedly for their Study 2 and asked for permission to include a summary of it in the paper. Dan Ariely denied our request, arguing, among other things, that the files we received may not be the actual data."

That does seem to be the defense, for another way this study rhymes with past scandals. Ariely can't remember the actual study protocol, can't be sure this data was what actually got published, and doesn't want a summary of the data that does exist being republished. It is quite possible no records of the experimental protocol exist to prove or disprove anything. I've got a search tool running through old MIT classifieds, and that's more a hope than a process.

There's a plausible, if disturbing, option that had the data fabrication been revealed before the replication failure and the separate scandals in other papers, it would have blown over. There's a plausible, if even more disturbing, option that even with those other data points, it will just blow over again. Maybe a retraction -- and to be fair, there's at least been a request for one -- probably Ariely doesn't get another tv show, probably not a slapped wrist, almost certainly no serious investigation by his school.

But that still leaves a lot of questions. Ariely has over a hundred other papers he's authored or coauthored listed on his own CV. It was plausible for the Honesty trial that no one else should have looked at Ariely's data, or looked at his conclusions with a skeptical eye. It remains plausible that, for two papers, his coauthors and peer reviewers didn't look at the paper with a skeptical eye. Two isn't feeling like a very real number here.

Trace mentioned this story with the opening line "I feel for the coauthor [Wertenbroch] here – he seems to have acted generally honorably and been caught off guard by having a liar for a collaborator." Before I read the paper, my gutcheck was "On one hand, it's nice to see Wertenbroch moving on it. After Gino, I'd seriously worried about the possibility Ariely had just been able to thrive so well because no one he was working with wanted honesty, either."

And then I saw that chart, counted on my fingers, and blinked.

Wertenbroch is, notably, CC'd on that 2014 e-mail where Ariely said that his study did not have any students with a negative balance, and had no participation fee to explain that 10 dollar offset. He could, plausibly, have not reviewed the data or the paper in a decade, and not realized the discrepancy even then. Leaves a bit of a question about what he does do, though. He could, plausibly, have not touched or looked at a single part of the study methodology or data in this entire paper.

In the Honesty trials from the previous scandal, there was an absolute mess where Ariely was responsible for Study 3, and there were serious questions about what, if any, exposure to the fraud the other authors might face. Gino seems to have had minimal responsibility of exposure to Ariely's original data... and separately produced some falsified data in a separate study in the same paper, and multiple other studies in other papers.

DataColada ends today's post with "In our next post, we will share analyses of the Study 1 data file [...]. That experiment is quite different. Our analyses are quite different. But our conclusions are quite similar." I'm working on writing up a post on the aftermath -- or lack thereof -- on the Hindawi scandal.

Eat at Arby's.

EDIT: while DataColada does not spell out the chart discrepancy, it is in their underlying R code. So the description here is more a dumbass noticing the same thing, not me finding something they missed.

Together with the Arday scandal, social science is really having a bad time right now. The killer is that these are not just random researchers, but professors, and not just professors, but celebrated stars with multiple books, regular media appearances and political contacts. They are likely to impact actual policy and general large-scale decision-making. And to add insult to injury, it's getting increasingly obvious that a) competent fabrication would be undetectable in practice, especially this much later and b) nobody actually cares enough to check in general. They were just both incompetent enough and people pretty much stumbled over it.

I'm unsure whether social science in its current form is salvageable. It comes down to feedback loops: Math has a direct feedback loop with reality through the structure of proofs. You can set up implausible assumptions, yes, but since you have to state them clearly you can always be called out on it. And it's still relevant information that, if the assumptions are true, it follows whatever you have shown in the proof. In physics, you usually can turn a model into predictions which will eventually be measurable, even if through rather elaborate channels such as measuring red shifts of far-way celestial bodies or by ramming very small particles into each other.

And so on further down the line; My own field, medical genetics, is a lot worse than those as well but we also have at least some points of connection with reality; IF I find a variant that is plausibly connected to a serious disability, and IF there is a medication that can effectively bypass that variant (for example, producing the protein that became malformed), we will usually see a noticeable positive impact on patients, even if it does not fully cure them.

But social science? Almost any result can be interpreted in almost any way one may choose; Surveys and closed, short-term experiments are possibly more connected to social desirability bias than to anything else, real-life cohorts are probably mostly selection effects, and controlling for them is usually somewhere between illegal and impossible. There is not zero good research, but it's completely drowned out by the bad, and even where people attempt to do it well, it's hard to say; For example, I like discontinuity design, and it imo works well for entry tests (barely made it into an university vs barely didn't make it), but exactly the same design for school area impact is probably nonsense bc people can deliberately move into those neighbourhoods so barely inside vs barely outside of the area is again mostly selection effects.

In such an environment, and given considerable pressure on them to deliver convenient results (again, unlike math and most of physics) which also results in substantial rewards outside academia, social science predictably devolves into near-pure high-school style popularity contests.

Social science research done well is not going to reach the levels of hard science, of course, but it is certainly not worthless if done well. There are various strategies to establish validity, reliability, etc. in quantitative measures, and credibility, transferability, etc in qualitative research. The problem is ideological capture, wrongthink, and institutional guardrailing. As well as a general lack of rigor, even when the methods to maintain such rigor and force accountability exist.

various strategies to establish validity, reliability, etc. in quantitative measures, and credibility, transferability, etc in qualitative research

Can you give us some examples?

It will all depend on what you're measuring and how you're measuring it. Quantitative and qualitative analyses are two completely different paradigms (though they are often married in mixed-methods research.) They don't even use the same vocabulary, and the product of their aims is fundamentally different.

A quantitative look at divorce, for example, might have a questionnaire as its main instrument. You get a bunch of numbers and it turns out the main predictor of divorce (note: not cause. Predictor) is contempt (as one of my friends likes to say whenever the topic comes up--I think he read something by Gladwell.) A qualitative study wouldn't have a questionnaire at all. You'd probably use interviews as participant observation of a marriage would have intrusion and ethical issues involved. You might end up with several narratives where you have a man recounting how his wife would always take her phone in the bathroom and stay far too long, as if sending messages, or a wife would talk about how the husband suddenly started washing his own clothes when he got home. Or whatever. The write-up for each would be totally different, and you'd be learning and digesting different things. A lot of people hate qualitative research. It's my main choice. It can be done very, very sloppily, however. I'll get to that in a second.

Back to making a quantitative questionnaire--if that questionnaire has not been, for example, piloted for content/face validity (via relevant experts who know the topic, but also item correlation/factor analysis) its results are more or less just noise (You also cannot use, for example factor analysis on just any set of questions. There are various requirements [a reasonable sample population, get rid of outliers, variables have to be on I believe an ordinal/interval scale, there has to be some degree of correlation in variables, you really need a normal distribution, etc.] If you don't fulfill these prerequisites again you're producing noise.) Once you have an instrument (questionnaire/whatever) that is relatively polished you administer it, but then in order to determine whether your results will be generalizable you'll need an appropriate sample size (and even then depending on the randomness of your sample your results will only generalize to that population, not the entirety of humans in all cultures).

Research design is also important. What you want to measure will determine how you measure it. I do not mean that you cater your plan to find a conclusion you want. I mean that if you want to measure something you need the right tool. Statistics do not show causation. A research design using statistics can establish causation (as in the numerous studies on cigarette smoking and cancer.) There are also criteria establishing causation.. The old saw correlation does not prove causation is of course true, though often positive or negative correlation is the smoke where there is fire. But not always, or even most of the time, in my opinion (which is worth about five cents.)

There are numerous threats to quantitative validity. Regression to mean, compensatory rivalry (in the case of two groups), compensatory equalization (when a researcher interferes), history, maturation, etc. etc. I highly recommend the book The Research Methods Knowledge Database (<- this is a US Amazon link, which I changed from the Japan link, but I cannot recall where you are.) I am not sure how much of this book no longer holds but it was a great read when I first went through it.

Qualitative research (what I do) can seem wild and woolly to an outsider, as if it's just you talking to someone and writing down your thoughts. And that is probably not a bad description of it, though there are, again, all sorts of rules you need to follow, primarily but not only to set the groundwork for reproducibility. And that, as we have seen, is a problem in both types of research--the findings are not reproducible. Documenting exactly what you do at every stage is very important for this. One study requires mountains of documentation. If you have any interest in qualitative research (and I would say the median user on the Motte does not) I would recommend the book Naturalistic Inquiry.

I have probably not answered your question, but I have let this sit and I didn't want to write something fast with my thumb. And you may know all of the above and more. You also may not find any of this compelling. I am not particularly good at statistics, though I used to use SPSS with some degree of expertise when I was made to do so. I believe that software may be outdated, or at least not used as much as it was. The Bayesian wave of the last 20/30 years has also changed a lot and I am out of the loop. Note I have known many researchers in social sciences, most of them what I would consider well-meaning ignoramuses, some of them outright frauds, this in both quantitative and qualitative. I have known only one man who I thought was a genius psychometrician; he didn't rely on the software as he knew the formulas and why they were the way they were. A math guy, like many of you (maybe not you-you, but the general Motte You.) He died far too early of a brain tumor in one of life's cruel ironies.

I have read a lot of shoddy research. Far too much. Enough to at least doubt (if not completely dismiss) every study that I see, most seriously when the study confirms my own biases. But then I think that's the right way to live. I could be wrong.

edit for clarity and this is still not very clear. My parentheses runneth over.

Standardized methods is one way. This at least lets you compare between studies that use the same methods, and thus determine if you find the same thing. The best methods have clearly defined scopes that indicate what they are and are not looking for. So the authors should be well aware of limitations well before starting their studies.

You are also supposed to be aware of your biases and validity threats, and admit to them in your writing so that the reader can easily understand your limitations.

It's just that a lot of researchers don't do this. I remember reading feminist papers, cited by journalists and used in political arguments, that make conclusions not supported by their results or where the methods are clearly designed to create a bias which is never admitted. The opportunity for rigor was clearly there, but the authors and reporters decided to forego it.