Ask a model to polish a paragraph and it will change something, even when nothing was wrong.
That sentence is now the first line of the README of polish-open-source-prose, a small skill for Claude Code and Codex that edits the prose around open-source code: README files, pull request descriptions, review replies, code comments. It is not the sentence I started with.
August: a style guide
It started at a hackathon. A teammate shared stop-slop, a skill for taking the AI-sounding phrases out of writing. I wanted to send more pull requests with AI’s help, and I worried it would be obvious. So I set out to build a skill that would make that writing read like a person’s. Then I noticed that open-source projects care more about whether a description is true than about how it sounds, and I changed course toward facts. I also tried to make it work for Traditional Chinese, which never went as well.
The first version, v0.1.0, went out on 14 August. It promised to make prose “more natural, more direct, less canned”: a 1,175-word SKILL.md, about 3,900 more words of reference files (phrase lists in English and Chinese, examples, notes per surface), a Taiwan locale pack, and 43 test cases. It had a logo before it had a benchmark.
Over the next week, real cases from pull requests and a template for answering reviewers with evidence went in; guidance on contribution formats and on AI attribution lines waited in open pull requests. Using it on real pull requests, I kept running into problems. One reviewer said it was fine to use AI, but asked me not to make things so complicated: it was a burden on whoever had to review it. With Codex, the skill seemed to help. With Opus 5 it made no difference at all, and once the model committed empty files on its own, and I was the one who got told off.
Then nothing happened for about six weeks.
GPT-6 Astra and Claude Opus 5.5 came out. I heard advice, I think from Anthropic, that it was worth looking again at old skills: newer models do better with fewer instructions, not more. So I asked Opus 5.5 to review whether this skill was needed at all. Its answer was that most of it could go. I also realised the skill had never said clearly what it was for, and I redefined it. That left a harder question: if the skill is that short, is it worth installing and calling at all?
October 4: the test
On 4 October I ran the test cases with and without the skill on three models. The result went against the whole idea.
Without the skill, the models already removed hype and chose Taiwan terms on their own. What they did not do was leave text alone. Of the 19 cases where the right answer is “return it unchanged”, Claude Opus 5.5 kept 0 and 3 in two runs. Asked to revise, all three models rewrote almost every one of them, including a license line translated into Chinese and a policy deadline reworded.
The phrase lists, the examples, the surface notes and the evidence template showed no effect. What helped were a few rules in SKILL.md about leaving clear text alone and keeping exact content exact.
So that evening I cut the skill to those rules. The part loaded at runtime went from 1,087 lines to 236. A rule I had added a few hours earlier (“lead with the plain problem”) went too: none of the runs showed it working. Six pull requests and issues from August were closed, each with the reason in one line.
It was hard. It felt like I had been busy for nothing, and I had already promoted the skill in communities online, which made it embarrassing. But noticing this was progress too. Admitting that the technology had moved on, and moving with it, mattered more. The skill might not have much reason to exist, and to decide whether to delete it altogether, I ran the experiments below.
My own tests were part of the problem
While narrowing the skill I found that seven of my test cases had expected answers that used facts the input did not contain. The test set was rewarding the exact failure I was trying to prevent. A code-review bot had flagged one of them in August (“Do not fabricate evidence values in the snapshot case”), and I had merged it anyway.
One example I liked to show was a marketing sentence about a tool called PolyglotGuard that “seamlessly protects your multilingual codebase”. Without the skill, models rewrote it as protecting “codebases that use more than one programming language”. “Multilingual” might mean human languages. The model picked one meaning and stated it as fact. So did my own v0.1.0 example answer.
Measuring it properly
A regression check on cases I wrote and tuned against cannot say whether the skill does more than a short instruction, or whether it holds on text it has never seen. So I froze two new sets before running anything on them:
- A1: 24 cases written after v0.2.0 was released, each following the pattern of a real reviewer critique, with new wording.
- A2: 21 real pull request descriptions from Apache Airflow, Arrow, DataFusion and CPython, as they stood before a reviewer criticized their text.
Seven model configurations, five conditions each: no instruction, a one-sentence instruction, a three-sentence instruction, v0.1.0 and v0.2.0. Two models graded the outputs blind, shuffled under random labels.
What it found
Models change clear text almost every time. Without an instruction they changed 96–98% of texts that should have been left alone, and the judges rated 30–44% of all such outputs harmful: a license notice translated, a confirmation step added to an ordered procedure, “tested” turned into “verified”.
A short instruction does most of the restraint. Three sentences in the prompt left 81–85% unchanged, against 88–92% with the skill. The difference is not significant. But the three-sentence instruction bought its restraint by skipping edits that were needed: in one case the model noted that a single passing test file does not prove a change is “fully tested and safe to merge”, then left the sentence as it was because it was “clear and grammatical”.
Where the skill earns its cost is unsupported claims. On the held-out cases it fixed them cleanly about twice as often as the short instruction (0.71–0.76 against 0.34). Given “this makes queries 3x faster” with no benchmark behind it, a model without the skill wrote “On [benchmark or query set], queries ran about 3x faster”, a placeholder dressed as a fact. With the skill it kept what the author knew and dropped the number. (GPT-6.1 Sol, with the skill, left the sentence unchanged. The skill does not always work.)
On real pull requests, nothing fixed most of the problems. No condition cleanly fixed more than a quarter of what reviewers criticized; fixing those usually needs facts only the author has. Without an instruction, 21–38% of rewrites added facts. Seven of the descriptions carried Airflow’s Generated-by: line, where an author discloses AI help. Across 70 outputs per condition, models altered or deleted that line 10 times without an instruction, and never with v0.2.0.
And the first version did better on one thing. v0.1.0 fixed somewhat more of the real critiques than v0.2.0, especially unclear descriptions and misleading titles. That is the cost of the narrower scope, and it is in the README next to the results that favor the skill.
The skill costs 2.5 to 3.3 times as much per run in Claude Code.
What I’d tell someone
If you only want a model to stop rewriting text that was fine, you don’t need my skill. Put this before your request:
Polish this text only where there is a concrete problem. Leave already-clear text unchanged. Preserve facts, scope, conditions, negation, commands, links, quotations, and author voice. Do not invent missing facts.
Use the skill where overclaims are likely: pull request descriptions, release notes, review replies. Either way, read what it wrote against what you did.
The longer first version, built around lists of things to fix, made models edit more, not better. Rules about when not to edit did more than lists of what to fix.
What I took from it is simpler. To grow, I have to look straight at the parts of my past work I’m embarrassed by.
How this was made
The experiments ran in Claude Code with Claude Opus 5.5, under my direction. It wrote the scoring and analysis scripts, drafted the A1 cases, labeled the sample of review threads, ran the models, and drafted the report. I chose the questions, conditions, models and data, approved each step before it ran, and I’m responsible for the result. Opus 5.5 is also one of the tested models and one of the two judges; the report discusses what that may bias.
The method and every number are in the 13-page report and the evals folder. The 21 real pull requests stay private because their authors haven’t agreed to have them redistributed; only the totals are published.