On February 6, 2025, Andrew Gelman (b. 1965) wrote about a Florida man whose cholesterol had begun collecting in yellow lumps under his skin after eight months of eating little but butter, cheese, and beef. Gelman put the question in his own title: how much am I to blame for this? He answered that he felt a share. Twenty years earlier he had promoted a psychologist named Seth Roberts (1953-2014) and his method of studying himself. Roberts spent his last months eating half a stick of butter a day to improve his brain speed. Reading back over the objections he had raised to Roberts in 2007, Gelman wrote, “In retrospect, I think I was too mild.”
The standard account of Gelman’s blog holds that he is the scourge of weak statistics, the man who takes apart the noisy study with the spectacular finding. The Roberts story runs the other way. Here Gelman is the promoter. His blog took an obscure journal article and delivered it to the New York Times. And here his criticism, delivered early, delivered often, delivered in print with the subject answering on the facing page, changed nothing at all.
They met in the early 1990s at Berkeley. Roberts came to the statistics seminar and the two of them talked about graphics. Berkeley was pushing faculty to teach freshman seminars, Gelman wanted to run one on left-handedness, and he asked the one psychologist he knew for a recommendation. Roberts volunteered. They taught it together and then had lunch regularly for years. Roberts told him about the experiments he ran on his own body: standing eight hours a day to sleep better, watching taped late-night monologues in the morning because life-sized faces lifted his mood, drinking unflavored sugar water to lose weight. He had spent close to ten years on the sleep problem before he turned to mood.
Gelman found the persistence admirable and the theory plausible. In April 2006 he reviewed The Shangri-La Diet on his blog, said he had known Seth for over ten years, and wrote that the case “seems pretty convincing to me,” adding that he had not tried the diet. Three months earlier he had told readers with New Year’s resolutions about happiness or weight to try Seth’s methods. In September 2005 he had written a defense of the whole approach, arguing that experiments that look easy had in fact cost Roberts years of disciplined measurement.
Roberts recorded what that did for him. In a May 2006 post called “The Ecology of New Ideas,” he traced the chain: his open-access self-experimentation article, then Gelman’s blog, then Alex Tabarrok at Marginal Revolution, then Stephen Dubner, then a Freakonomics column in the New York Times, then a book contract. He wrote it out again more formally in a 2012 article on the reception of his work, contrasting the coldness of his own profession with the interest he found outside it. His profession had no use for him. Gelman’s readers made him famous.
There is a small joke buried in the middle of this. In May 2007 Roberts interviewed Gelman about blogging, three parts over six days, the first extended account Gelman ever gave of what the blog was for. The man whose book the blog had built was asking the builder how the thing worked. Gelman told him the blog had started as a forum for his students, that the students never posted, and that he had developed an equanimous blog personality. In part three he mentioned a thousand readers a day and admitted he had no idea where the number came from.
Gelman’s first post on Roberts, March 16, 2005, praises the work and then asks: “What is the point where researchers should jump to a larger controlled trial?” Three weeks later, on April 8, he published an exchange with Roberts about control treatments and statistical power. In July 2006, defending himself against the charge that he was too hard on unconventional research, he cited Roberts as a case he had treated well, and said in the same breath that he thought replication plausible but could not be sure. In April 2006 he asked why Roberts had never tested the diet on rats, which would be cheap and would speak to the animal literature the theory rested on. Roberts answered that he and a collaborator had asked and been refused, because the committee thought the result impossible.
On April 23, 2007, Roberts reported that his balance improved on flaxseed oil compared with olive oil. Gelman’s first thought was measurement bias, since Roberts knew which oil he had taken and the balance test was hard. That August he proposed that Roberts have a partner assign the treatments blind, because expectation and noisy measurement can feed each other until an effect looks solid. Roberts replied that his findings had often surprised him and so could not be running on expectation.
Later in 2007 the two of them published the argument. “Weight Loss, Self-Experimentation, and Web Trials: A Conversation” ran in Chance, and the full text is on Gelman’s site. He grants a good deal: web trials can gather data cheaply, allow partial randomization, and study heterogeneity in ways a conventional trial cannot. He grants that clinical trials are often too small for subgroup claims and too slow to permit invention along the way. When Roberts argues that equalizing expectations can substitute for blinding, Gelman concedes he may be right and admits he does not know that literature.
Then he lays out the objections. No blinding. Protocol failures. Selection into the diet by people already unlike everyone else. Dropout. Measurement error. Motivation as an alternative route to both the weight and the mood results. Biases that cannot be assumed to cancel. He offers a design that would separate Roberts’s set-point theory from a duller explanation: give one group the oil away from meals and another group the oil with meals. And he tells Roberts he would believe the results a lot more if the treatments were blinded.
Roberts conceded that blinding would improve the oil experiment. He did not blind it.
The criticism kept coming. March 2009: Gelman says he is always skeptical when Roberts announces a new benefit, because Roberts may be hoping for the effect and then finding it, and because the reports reaching him come from people for whom the thing worked. June 2009: the acne anecdotes do not persuade him, and here is a cheap randomized volunteer study that might. January 2010: expectation could have a huge effect on Roberts’s own measurements, and here is why randomizing his particular design is hard and still worth doing.
On April 1, 2012, Gelman posted a randomized trial of the set-point diet, complete with an abstract from Nutrition. It was an April Fools fabrication. He noted that regular readers would know he had been waiting for this one a while. Seven years of asking had produced a joke, and the joke worked because the study did not exist.
Roberts answered everything. The next step after n of one should be n of one on somebody else, because a large study smuggles in assumptions nobody has tested. Self-experimentation costs nothing and permits treatments no grant would ever fund. Many things he tried failed, and some successes surprised him, so expectation cannot be doing all the work. He published his data, his graphs, and his R code. When a reader ran a reaction-time study that contradicted his soy theory, Roberts posted it with a link.
That transparency is part of the problem. Gelman would later write, citing his own paper on the subject, that honesty and openness are not enough. Roberts described exactly what he did. What he did could not support the conclusions he drew from it.
Then the pace changed. It had taken Roberts ten years to solve his sleep. By the last few years, everything worked. Acne, mood, reaction time, brain speed. In July 2018 Gelman put the point in one line: he started to let his ambition get ahead of him. Sleep hours and body weight can be measured. Brain function measured by quizzes you give yourself is another thing, and it does not take much unconscious bias to produce a clean result on a test you administer to yourself, about a treatment you invented, while an audience waits for the answer.
The last treatment was butter. Half a stick a day. A cardiologist in one of his audiences told him he was killing himself. Roberts answered that his own data beat epidemiology and its questionable assumptions.
He collapsed while hiking in Berkeley on April 26, 2014, at sixty. Occlusive coronary artery disease and an enlarged heart contributed to his death. His final column, published two days later, was called “Butter Makes Me Smarter.”
Gelman posted an obituary four days after. It praises the persistence and states the methodological failure without softening it: Roberts did little to reduce, control, or adjust for bias in his measurements, and he never systematically gathered data on other people. The comments filled with tributes. Gelman ended by saying it was good that Seth had found an online community that valued him.
Nine years later he took that back.
In a November 20, 2023 post, Gelman says his doctor told him his cholesterol was high and he needed to lose weight. He tried the Shangri-La diet for a few days, along with the eating less that Roberts always said the diet made easier. Then he thought: if the point is to eat less, why not just eat less. He dropped the sugar water and lost the weight anyway.
Had he stayed on it, he writes, he would be another testimonial. He would be telling you that only after switching had he been able to eat less without suffering. He would have been wrong, and there is no way he could have known.
A commenter named Juraj says you have abandoned a belief on the strength of a biased anecdote, which is the thing you spend your life criticizing. Gelman concedes the point and turns it into the argument: “Live by the anecdote, die by the anecdote.” Seth’s anecdote convinced him. His own anecdote unconvinced him. Neither one was evidence.
Roberts had a blog audience cheering every move, and Gelman writes that when people cheer your every move it becomes easier to fool yourself. In the comments the epidemiologist Sander Greenland observes that with death as your outcome you get no second trial, and Gelman agrees with him that the self-experimentation and the applause from admirers may have killed the man.
A woman named Wendy, who had known Roberts since the early eighties, wrote that the piece read as cruel innuendo about a kind man who could no longer answer. Alex Chernavsky, who runs a memorial site and has followed the diet since 2009, disputed the placebo account with fourteen years of his own weight data. Gary Wolf of the Quantified Self movement raised conditioning effects that neither placebo nor selection covers well. The bloggers at Slime Mold Time Mold pointed out that two of their three diet trials produced almost nothing, which is hard to square with the claim that anything works if you are ready. Gelman answered all of them and gave no ground on the main point. In a holiday open thread two years later, Chernavsky came back to say the Roberts posts had been unfair, and Gelman asked him to name one thing in them he thought was wrong.
In July 2017 Gelman posted that a bigshot psychologist, unhappy that his famous finding would not replicate, was scrambling furiously to preserve his theories. The psychologist was Fritz Strack (b. 1950), and Strack turned up in the comments to thank him for the promotion and to say he was quite happy. A graduate student in the same thread counted the evaluative words in the title, eight out of twenty-five, and asked Gelman to own that he had taken a shot. Gelman refused. He said he could not do this work if he were not free to say what he thought. The same refusal, on much heavier ground, is what Roberts’s friends met six years later.
Now the argument.
The comfortable position on public criticism in science says that ridicule fails on its target and succeeds on the audience. Amy Cuddy did not concede; thousands of graduate students learned what a forking path was. The joke is a teaching instrument aimed past the defendant at everyone watching the trial. That position lets the critic keep his jokes and his conscience.
The Roberts case tests it. Nothing about this criticism was mocking. It came from a friend of twenty years who had taught a class with the man and eaten lunch with him for years. It was specific, technical, and correct. It ran from March 2005 to January 2010 and beyond. It appeared in a peer-reviewed magazine with Roberts given equal space and the last word. It came with concrete, cheap, actionable proposals: use a partner to assign the oil blind, run the rats, run twenty volunteers on the acne question, split the oil groups by timing. Gelman conceded points to Roberts throughout, and in December 2009, when Roberts said Gelman had described his climate views unfairly, Gelman agreed on the spot and revised.
John Braithwaite drew the standard distinction in Crime, Shame and Reintegration: shaming that stigmatizes the person against shaming that condemns the act while leaving a road back into good standing. Gelman’s criticism is as reintegrative as criticism gets. The act was named, the person was kept, the road back was drawn on the map with the cost of the trip itemized. If reintegrative criticism works anywhere, it should have worked here.
It failed. Roberts answered every objection, conceded the small ones, kept the practice, accelerated it, and likely died of it.
Colin Wayne Leach and Atilla Cidam, in “When Is Shame Linked to Constructive Approach Orientation? A Meta-Analysis,” find that shame moves people toward repair when the damaged standing looks repairable, and toward withdrawal when it does not. Everything turns on what the criticized person believes is at stake.
Gelman was criticizing an experiment. He thought he was asking a colleague to blind one measurement, a fix costing a few weeks and one helper. Roberts heard something else. He had left mainstream academic psychology on purpose, after tenure, because the students cared about their lives and the publishable research did not. Self-experimentation was what he had instead of a career. Conceding that his balance results ran on expectation would not have cost him an experiment. It would have cost him the second half of his life, retroactively, and left him a man who had walked away from a respectable position to fool himself for twenty years in public.
Nothing about that is repairable. So the courtesy in Gelman’s criticism could not reach it, and neither could sarcasm, and neither could a better argument. The size of the concession decides the response. Cuddy had a book, a talk with tens of millions of views, and a public identity built on one finding. Wansink had a lab, a directorship, and a career. Roberts had his second act. In each case the correction on offer was small and what it implied was total.
What moved the Wansink case was an institution with the power to act, arriving after reproducible demonstrations of error had accumulated past the point a provost could ignore them. Roberts had no institution over him. He had retired from Berkeley, moved to Tsinghua, and answered to a message board. Once you strip out the employer, the journal, and the tenure committee, you can see in his case what criticism accomplishes when the person being corrected has everything riding on the answer.
The audience half of the standard position holds up better, and gets darker. Hezhi Chen and colleagues find that people who watch a third party punish a transgressor revise their sense of what is normal and change their own behavior accordingly, which is the evidence for the claim that the joke teaches the gallery. But Xiaoyu Ge reports that taking part in online shaming lowers the participants’ own moral sensitivity and raises their willingness to join the next one. Applied to a blog with two hundred thousand comments, that says the gallery is being trained, and what it is trained to do is show up for the next defendant.
Gelman spends his working life asking what selects the evidence a person sees. Why this study and not that one. Who sent it. What made it visible. In February 2025 he pointed that question at his own archive and answered that he bore some share of responsibility for a stranger in Florida with cholesterol under his skin, because in 2005 he had made an obscure psychologist’s method famous. He is the reason the audience existed. He built the cheering section he later blamed for the death, and then he apologized for having built it, to a man he never met, about a diet he never recommended.
Almost nobody does this. The usual move after a promotion goes bad is silence, and the archive makes silence easy, since nobody is going to read your 2005 posts. Gelman went back and read them.
Roberts loved Brian Wansink’s (b. 1960) work. He praised the research design on his blog in 2006 and again in 2007, and in August 2019 Gelman reproduced that enthusiasm to show how completely the bottomless soup bowl had been believed before anyone looked. The man Gelman promoted vouched for the man Gelman helped bring down, and both of them fell for the same reason, which is that a friendly audience and a flexible measurement will find you whatever you are looking for.
And in December 2024, describing an athlete who experiments on himself with more care, Gelman called him a sane Seth Roberts. The friend has become a type, a shorthand, a character in the blog’s permanent cast. Every long-running blog does this to the people in it.
In March 2026 Gelman wrote that computer scientists now occupy the position economists held in 2000 and Freudians held in the 1950s, and that the trouble starts when the gurus and their hangers-on begin to believe their own hype. He had used that phrase before. He used it in 2023, about Seth Roberts, whose hype he supplied.
