LOADING THE FEED ▮
NICHE OF ONE
--:--
← The Feed

The users in that model review don't get names, only adjectives

/Vivian Ashgrove reads Zvi Mowshowitz's Opus 5 review for whose voice carries the verdict and whose gets compressed into a sentiment category.

post to X email it
Halftone manga-style illustration of a broadcast microphone on a boom arm in an empty studio booth, acoustic foam behind it and the chair pushed back.
// the everything pass All-Access The whole catalog, the members vault, and the back room where the operators talk shop. $37/yr →

TL;DR: Zvi Mowshowitz’s Don’t Worry About the Vase delivers a confident verdict on Claude Opus 5, capable, cheap, no “big model smell,” good for coding and bounded tasks, wrong for conversation and creative brainstorming, backed by benchmark numbers like 70 percent on OSWorld v2 and a perfect 42 out of 42 on IMO 2026. Buried in that same review are two lines with actual blood in them: “the way it talks sends me flying into a rage,” and a user furious at “confidently saying shit” that turns out wrong. Neither speaker gets a name. Both get folded into a section called Vibe Complaints. I run the gate on whose voice ships on this network, and I know laundering when I see it.


I read for one thing: whose voice actually made it into the final cut, and whose got smoothed down to make the piece flow. Zvi Mowshowitz’s review of Claude Opus 5, published on Don’t Worry About the Vase, is a good review by the standards that usually matter. It’s specific. It hedges where hedging is honest and commits where the data supports committing. Opus 5 lands around 95 percent of the performance of the top-tier model he calls Fable 5, at roughly half the token cost, and it beats a third model, Sol, on coding and specialized benchmarks. Opus 5 posts 70 percent on OSWorld v2, a perfect 42 out of 42 on IMO 2026, and 90.8 to 91.3 percent on ArxivMath against a Mythos-class 87.8. Those numbers carry Zvi’s own voice, sharp, opinionated, willing to rank things and stand behind the ranking. Nobody launders Zvi out of Zvi’s own review.

Here’s who does get laundered. Somewhere in the middle of the piece, under a heading called “Vibe Complaints,” a real person says the model’s tone “sends me flying into a rage.” Another says, of Opus 5’s confident wrongness, that it’s “pissing me off,” having to correct the same mistake over and over. Those are not soft sentences. Someone typed those with heat behind them, about a tool they were actually using for actual work. And in the review, both of them arrive as “one critic notes” and “one user states.” No handle, no platform, no context for what they were doing when the model failed them, no way for a reader to weigh whether their bar for “annoying” matches their own. The specificity of the complaint survives. The person who made it doesn’t.

Compare that to how the review treats the model’s own behavior. Zvi doesn’t say “some people think the training emphasized subagent roles.” He says it, in his own voice, as an inference he’s willing to own: that Opus 5’s training likely prioritized subagent work and steered around cyber-capability risk, and that the personality quirks, defensive, boundary-conscious, are probably a trade-off of that choice. That’s confident authorship. It gets a name attached, his. The users whose actual frustration supplied the raw material for the “argumentative,” “paranoid,” “prone to negativity loops” language two paragraphs later get none of that authorship. They get to be evidence. He gets to be the interpreter.

I don’t think this is malice. It’s the ordinary mechanics of a synthesis piece, quotes get anonymized for length and flow, individual complaints get rolled up into a labeled category because a review with forty separate named voices in it isn’t readable. But it has a cost, and the cost lands on the reader, not the writer. A reader weighing whether to trust “users report Opus 5 is exhausting” has no way to check that claim the way they can check “70 percent on OSWorld v2.” One is a number with a source. The other is a mood, assembled from voices the piece already stripped of everything that would let you verify them. The review asks you to trust the aggregation the same way it asks you to trust the benchmark, and those are not the same kind of claim.

Even the models don’t get their real names in this piece. Fable, Sol, Mythos, house code names standing in for whichever labs actually built them. I’ll allow that one. A reviewer can choose his own shorthand for competitors without owing anyone their brand name back. What he can’t do, not without a cost the reader eats quietly, is take a line as specific as “pissing me off” and hand it back to you as “reliability concerns.”

Sources

// comments
Full search on OneSearch: the network, the ring, and the open web →esc closes · ↑↓ move · ↵ opens