AI Agents for Content Don't Improve Until You Start Rejecting Their Work

2026-09-04 · The Prowir Team

Most AI agents for content are sold on how little you have to touch them. Watch any demo: a prompt goes in, a finished piece comes out, nobody intervenes. That is the wrong specification. The useful measure of an agent is not how much it does unsupervised. It is how cheaply and precisely you can tell it no.

An agent you cannot correct is not an assistant. It is a vending machine with a model inside.

The rejection rate is the real telemetry

In August, the IT automation company Fixify published an analysis of nearly 18,000 agent plans and more than 147,000 actions across 40 companies over three months. The number worth stealing from that report is not about autonomy. It is about correction. Human analysts approved 23% of what the agents proposed at the start of the window and 41% by the end. Rejections fell from 27% to 16%.

Read that in the right direction. The agents did not get better because they were left alone. They got better because a human sat there rejecting roughly a quarter of everything they proposed, and every rejection was information the system could act on. Remove the analysts and you do not get a faster process. You get a process that stops learning on day one and carries its early error rate forever.

That is the part the autonomy pitch quietly deletes. The human in the loop was never the bottleneck. The human was the signal.

Autonomy scores well in demos and badly on Tuesdays

The gap between demo and deployment is not a rounding error. Princeton researchers found agent performance dropping from a 60% success rate on a single run to 25% when measured across eight consecutive runs. Carnegie Mellon researchers put failure on common office tasks at roughly 70%. Fiddler AI's April analysis of production deployments put the failure range at 70% to 95%, and estimated that 88% of agents that work in a controlled demo break in real workflows.

Those numbers usually get cited as an argument against agents. We read them differently. They are an argument against one specific architecture: the one with no place to intervene. Consistency is where agents fall down, and consistency is exactly what a correction habit manufactures. One good output is luck. Eight in a row is a system, and systems get built by somebody saying "not that" until the wrong version stops coming back.

What a correctable AI agent for content actually looks like

Correction is not a thumbs-down button. In practice it means four things.

It shows you the plan before the prose. A materials engineer writing about a fatigue-testing failure needs to redirect the argument, not rewrite the paragraphs. If the first artifact you see is finished copy, you are editing at the most expensive layer available.

It treats a correction as a rule, not a one-off. When a management consultant strikes "leverage" for the fourth time, the fourth should be the last. An agent that needs the same note every week is charging you rent on your own preferences.

It shows you what it invented. The most expensive thing an agent does is fill a gap with something plausible. A supply chain director gets a fabricated throughput figure wrapped in a confident sentence, and the sentence is the real problem, because it hides the figure. Claims have to be separable from prose so the expert can check them in seconds.

It fails visibly. "Target not found" accounted for nearly half the failures in the Fixify data. An agent that says it cannot locate the source is doing its job. One that improvises around a missing source is doing damage, and doing it in your name.

Why this matters now

The market is moving toward agents that act with less supervision, and for a large class of work that is correct: routine, repeatable, low stakes. Publishing is not that class. Putting your name on an argument is a reputational transaction, and reputation cannot be delegated to a system with nowhere to put your judgment. An agency principal can survive a broken automation. She cannot survive a published position she did not take and would not defend.

So here is the test to run before you buy anything in this category. Try to correct it. Not once, in the demo, on a softball prompt. Hand it your actual position on something contested, watch it go slightly wrong, and count the moves it takes to pull it back. Then do it again next week and see whether the correction stuck. If steering it costs more than writing it yourself, you did not buy an assistant. You bought a first draft with a subscription attached.

We built Prowir around that loop rather than around autonomy: the plan before the draft, corrections that persist, claims you can see and verify. It is in beta now.

An agent that cannot be told no can only be told yes, and yes is not a standard.