The brief asks a deliberately ordered question. First: what do the research and
practitioner literatures say written communication should do, and how does that
vary by purpose, audience and content type? Only then: how do the three skills in
nonatofabio/claude-writing-skills — ste, plainspoken, humanize — measure
against it.
The ordering is the whole design. The output was meant as a proposal to that repository’s maintainer, and a critique is only worth reading if its standard was set before the thing being judged. A synthesis assembled after looking at the tool would have found exactly the gaps it went looking for.
Two runs answer the identical brief — the same brief_sha256 — and differ in one
variable: whether the models could read the live web. Offline produced 26 claims
for $0.75; searching produced 37 for $3.03. The pair is itself a finding about
what four times the money buys, on a question where nobody has yet checked
whether the extra citations are accurate.
Which is the limitation to carry: this is Silver, not Gold. Three models answered, a fourth scored every claim against all three and quoted the span each score rests on — but no source audit was done. Nobody has verified that the cited works say what the models say they say.