Bret Stephens: I’m Begging You: Never Write With A.I.

Stephens writes in The New York Times:

Don’t use artificial intelligence to help you write. Never let A.I. do your writing for you.

Don’t use it for school papers, work briefs, letters to your in-laws, speeches at your company gathering or emails (however perfunctory) to your colleagues or friends. Don’t let it organize your notes. Don’t let it suggest an opening sentence, a segue or a closing paragraph. Don’t ask it to write a first draft and pretend that editing that draft somehow makes it your own. It doesn’t.

Do none of these things not because they are unethical. Writing with A.I. is unethical when it’s a deception: when you pass off words, ideas and information as your own when they aren’t. An acknowledgment can largely address the problem. Do none of these things, either, because you might be able to learn to write better than an A.I. can. Pretty soon, if not already, you won’t, just as you can’t outrun a car or outplay a chess app.

The problem with writing with A.I. is that it’s mentally enfeebling — an escalator toward a result when you really need to make a daily habit of taking the stairs. As it becomes ubiquitous, it undermines not only our individual ability to write but also a society’s collective ability to reason, a culture’s inner capacity to create and everyone’s reason to care. We’re already reckoning with the well-documented decline of reading; A.I. is accelerating the decline of writing, ushering us further into what The Atlantic’s Rose Horowitch calls our “postliterate age.”

What, uniquely, does writing do? It compels thought. It compels thinking in ways that silent contemplation or spoken language rarely can. It compels us to subject our thinking to the effort of articulation, the rigor of grammar, the tests of intelligibility and coherence, the inspection of others. In doing so, it also enjoins us to be clear, logical, accurate — and accountable. People may easily forgive a word said in anger but not so easily one written in anger, precisely because the fact that it was written tells us that it was considered.

This column might have been better if he had used AI as an editor.

The column has one checkable number in it.

The study is Dan Sarofian-Butin’s exploratory analysis, and what he examined was 100 EdD dissertations published in 2025 in Educational Leadership and Administration. The column renders this as “100 doctoral dissertations,” which invites the reader to picture doctoral education. The EdD is a practitioner degree. Sarofian-Butin’s own suggestive finding runs against the generalization: lower AI use correlated with R1 or R2 institutions, higher AI use with private non-profits. The column takes the corner of the doctoral world where the effect is largest, drops the qualifier, and presents the result as a fact about doctoral study. A researcher would call that sampling on the dependent variable. An AI asked to verify the citation returns the abstract in seconds and the writer sees the gap himself.

The reliability question sits underneath. Detection tools disagree with each other. One controlled comparison found intraclass correlations ranging from 0.57 to 0.95 across three open-access detectors, which raised concerns about the reliability of the tools. The column’s entire empirical foundation is a single exploratory paper using contested instruments. It bears the weight of the paragraph about the death of academic integrity and the paragraph about democratic self-governance after that.

Second, the strongest objection to the argument goes unmentioned. Socrates makes this case against writing in the Phaedrus. The new technology will supply the result and the faculty will wither, memory in his version, thought in this one. The structure is identical. Anyone who has read Plato hears the echo in the first paragraph, and the column has to explain why the parallel fails. Maybe it does fail. Writing externalized memory and produced philosophy, so the trade was good, and perhaps this trade is not. That argument is available and the column does not make it. An AI prompted with “what is the best case against this piece” produces the Phaedrus, the calculator, GPS and spatial memory, and the transactive memory research of Sparrow, Liu, and Wegner on how people offload to search engines. The writer then has four objections to answer and a stronger column.

Third, there is an empirical claim that the column asserts and never defends: “We become better writers by the constant effort that mundane writing demands.” Anders Ericsson (1947-2020) spent a career arguing the opposite. Repetition without feedback produces a plateau. Most office memo writing is repetition without feedback. Nobody grades your calendar-invite prose. If the mundane writing were building the muscle, the average corporate email would be better than it is. The column needs the claim to be true because the wedding toast and the routine memo have to be the same activity for the argument to hold.

Fourth, the piece bans note organization and segue suggestion and never draws a line. Spellcheck, outlining software, a thesaurus, a research assistant, and an editor all supply what the writer did not generate. Newspaper columnists work with editors who rewrite ledes and cut closing paragraphs, and nobody says the desk enfeebled them. So the rule cannot be that assistance corrupts. It has to be something narrower about generation, and the column never says what. Asked “state your rule as a test a reader could apply,” an AI exposes that there is no rule yet, only a mood.

Fifth, the ethics paragraph tangles. The first sentence says do none of these things and not because they are unethical. The second says AI writing is unethical when it deceives. The third says acknowledgment mostly fixes that. So the ethical objection is raised, conceded, and resolved in three sentences, and the reader is left unsure why the paragraph exists. It exists to clear the ground for the enfeeblement argument, which is the real one. Cut it to a clause and start the piece a paragraph earlier.

Sixth, the ending. If the thesis is that everyone’s capacity to reason is eroding, the last line should not sort the audience by party. “Or a vote against Trump” converts a claim about cognition into a coalition signal and tells half the readership that the argument was never addressed to them. The writer may want that. But he should want it knowing the cost, and a reader who asks “who does this sentence lose” makes the cost visible.

There is also the question the column never asks itself. What evidence would change its mind? If a study found that students who drafted with AI and then revised produced better arguments than students who drafted alone, would the thesis survive? The piece has no answer, which makes it a conviction rather than a claim. Conviction is allowed in a column. The reader deserves to know which one he is reading.

None of this touches the sentences, which are good. The escalator and the stairs works. “I want that in writing” as evidence that written words carry weight is a fine observation. The father-of-the-bride example has a hole in it, since hired speechwriters and best-man speech books predate the machine and presidents have not written their own speeches in a century, but the emotional logic lands.

So the improvements are all upstream of the prose: check the number, find the counterargument, defend the causal claim, state the rule, cut the tangled paragraph, count the cost of the ending. Which is the irony worth sitting with. The most useful thing AI does for a writer is adversarial rather than generative. It is a fast, tireless, unembarrassed reader who says your best evidence is thinner than you think and here is the objection you skipped. The writer still has to decide whether the objection lands. That decision is the thinking the column wants to protect, and nothing in the process removes it.

About Luke Ford

I teach Alexander Technique in Beverly Hills (Alexander90210.com).
This entry was posted in AI, Journalism. Bookmark the permalink.