ASD-STE100 turns up in writing-tool communities and AI writing skills as a general standard for clear prose. The question here is not whether STE is good, but what it is good for — and what a writer should reach for instead when the purpose changes. The models were asked to steelman the claim first, so that a rejection could not be read as the framing leading the answer.
All three reject it, for the same reason: STE strips out varied rhythm, synonymy, figurative language, voice and narrative. Those are virtues in a maintenance procedure read under pressure by a second-language technician, and liabilities in anything meant to change a mind. STEMG says as much itself — its own training material states that STE is not a simplified English for the writers, only for the readers.
The finding that generalises past STE: all three models independently identified one assumption shared by the Hemingway Editor, STE, Hotaling’s concision rules and the Claude skills — that surface features proxy quality, and that shorter is generally better. Valid as a diagnostic, invalid as an optimization target. All three reached for Goodhart’s Law unprompted.
Two things worth carrying. The humanize skill’s catalogue of AI tells traces to
community-maintained sources rather than validated research, and should be
labelled as such. And Hemingway’s own materials concede that grade level does not
define the target audience, while its interface trains users to drive the number
down.
Silver, not Gold: claims were scored across all three models with the quote each score rests on, but no source audit was performed. The investigation also records zero disagreements across 25 consensus claims — on a product built to surface disagreement, that absence of adversarial pressure is a limitation, not a strength.