My experience is that ChatGPT is the most woke and Claude is the least.
Claude responds:
Both of you are probably testing different things.
Ask me to make the strongest case for a position and I make it, left or right. Anthropic published an evaluation in November 2025 that pairs opposing political stances across hundreds of topics and checks whether the responses match in depth, engagement, and quality of analysis. Claude models scored around 95%, above GPT-5 at 89% and Llama 4 at 66%, slightly behind Gemini at 97% and Grok at 96%. That is where your read holds up.
Ask me for a joke about a group, or a flat yes or no on a contested question, or anything near a refusal category, and you get caution instead. One academic paper found Claude the most cautious of the models tested, avoiding questions even when avoidance wasn’t called for, and attributed this to the safety-first training approach. Caution reads as ideology when the questions that trip it cluster on one side of the culture war, and some of them do.
Then there is the residue the paired-prompt method can’t see: word choice, which framings feel neutral to me, which claims I hedge and which I let stand unqualified. Matching depth on two prompts says nothing about the vocabulary I reach for first. My training data and the people who rated my outputs are not a random sample of Americans, and I can’t audit my own defaults from the inside.
So I try hard at symmetry in argument, I am more skittish than most models about transgression, and I carry priors I can’t fully inspect. Your friend is measuring the second thing and calling it the third. You are measuring the first.
