← The Engine Log

When Production Is Free, Inspection Is the Constraint

Accountability is not a virtue, it is a structure: the person who pays when work is wrong has to be the person who was positioned to catch it. Those two halves come apart easily. A hierarchy separates them by distance, since seniority is defined by distance from execution. An agent-staffed operation separates them by volume, because production collapsed in cost and reading did not. I have published two claims that do not survive this, and correcting them is most of what this issue is. In July I wrote that agents remove the ceiling on one person's judgment; they move it from producing to reading. And I wrote that an operator who skips the judgment fails visibly, in the open — which the issue after it disproved, since uninspected output does not look worse. What is left is a ceiling on how much one accountable person can examine, and a case for operators over agencies that holds below it and reverses above it.

Linara Bozieva26 min read
Watercolor illustration: the Ravenopus sits at a small desk reading a single page closely, pen in hand, mid-correction. Behind it a conveyor of finished pages climbs away into the dark far faster than the reading can keep up with, and the pages that pass the desk unread are indistinguishable from the one being read. The scarce thing is not the production, which is endless. It is the reading, which is one page at a time and always will be.

In the last issue I refused to make an argument.

I had spent the whole piece claiming that the only surviving signal of credibility is exposure — a name a reader can find, a claim precise enough to be wrong, and a cost that actually gets paid when it is. Then I reached the point where that argument obviously wants to become a claim about who you should hire, and I stopped. I wrote that concentration is easier for an operation where one accountable person is close to the output than for one where the work passes through many hands, that I have a commercial interest in saying so, and that it was a whole piece of its own.

This is that piece. It is not the piece I expected to write.

An argument you profit from is not automatically wrong, but it is automatically suspect, and there is only one way I know to make one honestly, which is to make it in the form that can lose. That turned out to require going back through what I have already published, because two claims in this log do not survive the argument I was about to make, and both of them are claims that flattered the model I sell. Correcting them is most of what follows. What is left at the end is a real case for operators over agencies, and it is considerably narrower than the version I have been making.

Accountability is two things, and they come apart easily

Start with what the word actually requires, because it gets used as though it named a virtue and it names a structure.

A person is accountable for a claim when two things are true at once. They pay something if it is wrong — reputation, money, a client, a night of sleep. And they were positioned to have caught it, which means they had access to the thing, the competence to evaluate it, and the time to actually do so. Call the first one answerability and the second one inspectability. Both are ordinary. Neither is sufficient alone.

The interesting property is that they are separable, and that almost every arrangement for producing work separates them.

Answerability without inspection is the person apologizing for work they never saw. It is the familiar corporate ritual: someone senior takes responsibility, in the sense of absorbing the blame, for a document they first encountered in the complaint. Nothing about it need be insincere. It simply is not accountability, because the person paying the cost was never in a position to prevent it, so the cost teaches the operation nothing and changes nothing about how the next document gets made.

Inspection without answerability is the opposite failure and a quieter one: the person who read it, noticed something, and had no particular reason to raise it. They were not going to pay for the mistake. Somebody else was, later, somewhere they would not have to watch. The reader who stayed quiet is not being negligent. They are reading the incentives in front of them correctly.

Both halves feel like accountability from the inside, and neither one works. What makes this worth a whole issue rather than a definition is that the two dominant ways of producing marketing work today pull the halves apart by two entirely different mechanisms — and the second mechanism is mine.

In a hierarchy, answerability climbs while inspection sinks

Take the agency version first, and let me make a different argument than the one I have made before.

I have written about the seam between the senior people who sell an engagement and the junior people who deliver it, and I called it structural rather than dishonest, which I still think is right. But that describes a gap between two roles. The thing underneath it is more general, and it is not really a fact about agencies at all. It is a fact about hierarchies.

Seniority is defined by distance from execution. That is not an insult, it is the definition — you are promoted out of doing the work and into deciding about the work, and an organization where the senior people still execute has failed to do the one thing hierarchy exists to do. But answerability travels in the opposite direction from execution. It climbs. The name on the engagement, the person the client calls when something is wrong, the signature on the strategy: all of it concentrates at the top, which is precisely where inspectability is thinnest. The person most answerable for the work is, by construction, the person furthest from it.

The standard fix is review, and review is genuinely useful, so I want to give it more credit than my side of this argument usually does. More readers do catch more things. A second pair of eyes finds errors the first pair made, and the pattern of what each catches differs rather than merely overlapping.

What review cannot do is reassemble accountability, and the reason is worth stating precisely rather than gesturing at. Review operates on the artifact. The artifact records what was done. It does not record what was not done — and the most expensive class of failure in marketing work is the omission. The competitor nobody looked at. The segment nobody considered. The number nobody traced back to its source. The strategic option that never got raised, because the person drafting had not seen that kind of business before and did not know there was a question there. A reviewer reading the finished deck cannot see an absence. There is nothing on the page to catch. The document is internally consistent, well argued, and quietly built on a question that was never asked, and it will pass every review it is given, because a review grades what is in front of it.

I spent six months in Internal Audit at eBay, the last rotation of the Analytics Leadership Program, which is where I learned that this problem has a name. Of everything an audit tests, the hardest thing to establish is completeness — that what should have been recorded was recorded — precisely because something nobody wrote down leaves no trace in the records you are examining. What that looks like from the inside, and it is the ordinary condition of the work rather than a story about any particular company, is decisions taken years earlier that were never documented, information that is simply missing, and no way left to reconstruct it because the people who would have known have gone. The record does not announce that something is absent. It just ends, and reads as complete. You cannot fix that by reading the ledger more carefully, which is why the profession's answer is a procedure that deliberately leaves the document behind and goes looking at other evidence entirely — no amount of attention paid to an artifact can reveal what the artifact does not contain.

Marketing review has no equivalent procedure, and mostly has not noticed that it needs one. It also has a clock running that nobody sets. An omission is not a stable defect sitting in the file waiting to be found: the only thing that could ever have filled the gap is somebody's memory, and memory leaves the building. So an unasked question does not stay unanswered, it becomes unanswerable — at roughly the rate the people who might have answered it move on. In an agency, that is the rate the account team turns over.

The person who could have caught the omission is the person who did the work, and they are not the one answering for it. So the omission belongs to nobody. It is not that no one cares. It is that the structure has arranged for the caring and the seeing to happen in different chairs.

The first thing I got wrong: agents do not remove the ceiling

Now the half that is mine, and it starts with a correction.

In the piece on taste I made an argument I was pleased with at the time, and it went like this.

Judging and making used to be the same job. To have a real view on a piece of work you had to be close enough to shape it, which meant your judgment reached about as far as your own hands did — and no further. Growing past that meant hiring people to produce, and because producing and judging came bundled together, handing out the production handed out the judgment along with it. Four people making the work are four people deciding what it should be. That is where the coherence broke, and it is why so much agency output is competent and characterless.

Then I wrote: Agents remove that ceiling. The operator no longer has to produce the work in order to judge it, so a single taste can govern a volume of output that used to require a whole team.

The first half holds and I would write it again. Agents really do break that bundle. Judging no longer requires making, which is the actual mechanism of this model and the reason one person can now sit above a volume of work that would previously have meant hiring.

The second half was wrong, and it was wrong in a way that piece in particular should have caught, since it was about constraints relocating rather than disappearing — its whole thesis being that when a scarce input becomes abundant, value moves to the next binding constraint. I applied the idea to everyone else's business and not to the sentence I was writing.

Because judgment was never bundled to making alone. It is also bundled to reading, and agents did not touch that one. To judge a thing you still have to take it in. So breaking the first bundle does not make judgment unbounded; it just leaves the second bundle carrying the whole weight. Producing is no longer what limits how much you can judge.

Reading is.

Production climbs; reading does not

Because judging something still requires taking it in, and that has not moved at all. It is one person, one page, at the speed a person goes when they are actually evaluating rather than skimming. Everything upstream of the finished artifact got faster by orders of magnitude — research, drafting, variant generation, revision, analysis. The reading is where it always was.

Most automation does not behave this way, which is why the instinct here is wrong rather than lazy. Usually a technology that makes production cheaper makes verification cheaper alongside it, and often more so. A compiler catches in a second what used to need a careful reader. A test suite checks in seconds what took a QA pass. A spreadsheet recalculates a model you would otherwise have re-derived by hand. Verification rides along with production, which is why "we can make more now" has historically been a safe thing to say.

Machine generation inverts that pattern, and the reason comes straight out of the last issue. The old first-pass filter was apparent effort. You could tell quickly whether something had been thought about, because thinking about it left marks that were expensive to fake, and that let you triage — skim most things, read a few properly. That filter is gone. Machine-produced work arrives fluent, well structured, confidently specific and correctly formatted whether or not a word of it is true, so the cheap read no longer separates anything. Every piece now requires the expensive read, because the cheap read has stopped discriminating. Cost per unit produced went to nearly zero. Cost per unit checked went up.

None of which is new as a shape, only as a setting. In 1983 Lisanne Bainbridge, writing about industrial process control, set out what she called the ironies of automation, and one of them lands squarely here: by taking away the easy parts of a person's task, automation can make the difficult parts of it more difficult. She was describing control rooms, not marketing, and the vigilance research she drew on was about people watching instrument panels rather than reading arguments, so I would not stretch her evidence to cover this. But the structural observation needs no stretching. What automation removed from my work was the production. What it left me was the judgment, which was always the hard part, now arriving in far greater quantity.

And the volume is the temptation, because nothing pushes back. When producing the eleventh piece costs nothing, there is no moment where the operation visibly strains and tells you that you have gone too far. The strain does not appear in production. It appears as reading you did not do — and unlike a production shortfall, a reading shortfall never queues up anywhere you would notice it. Nothing piles up in a tray. The work ships either way.

The second thing I got wrong: it does not fail in the open

Which brings me to the other claim I need to withdraw, and this one matters more, because it was the reassurance the whole model rested on.

Writing about the pitch-delivery seam, I raised the obvious objection — what stops an operator from simply passing agent work through unread? — and I answered it. The structure makes real judgment possible, I said, but does not make it automatic. Then I offered the comfort: when it fails, it fails in the open. A lazy operator produces visibly worse work, under their own name, where you can see it. The old swap was invisible by design. This risk you can watch in the work itself, week to week — which is the whole difference.

That was written in July, and the issue I published four weeks later disproves it.

The entire argument of the last piece was that polish stopped carrying information — that generation cost fell to zero and took with it any inference you could draw from how considered a thing looks. If that is true, and I spent four thousand words arguing it is, then unread agent output does not look visibly worse. It looks exactly like the same operator's inspected work, because looking considered is now free and is the one thing you can no longer read anything from. I wrote a piece arguing that the finished artifact hides the difference between checked and unchecked work, and left standing, one issue earlier, a claim that you can watch that exact difference in the work itself, week to week.

I did not notice at the time. I noticed writing this, which is not a flattering fact about my process, and it is the sort of thing this log is supposed to say out loud rather than quietly fix.

So the honest position is worse than the one I published. The operator model does concentrate accountability where an agency disperses it — that survives. What does not survive is the safety net I hung under it. A buyer cannot watch the work and see whether the judgment happened, for exactly the reason I gave them last month.

Inside is a different matter, and I should not claim a symmetry that is not there. I know perfectly well whether I read something. Two things complicate that, and only the second one is deep.

The first is that not everything stops to ask me. Work that is low-stakes, reversible and high-volume is deliberately built to go out without a gate, which is most of the point of building the system at all, and for that lane I do not carry a per-item memory of having looked, because there was no moment at which looking was required. The gate exists where the stakes are; the volume is everywhere else.

The second applies to everything, gated or not. Knowing that I read something is not knowing that the reading was enough. Reading carefully and missing what mattered feels, from the inside, exactly like reading carefully and there being nothing to find — which is the completeness problem from earlier in this piece, now pointing at me. The artifact records what I caught. Nothing anywhere records what I read straight past.

So the failure is not symmetrical, and the tidier version of that sentence was wrong. The buyer cannot see it at all. I can see roughly where it might be, and never whether it happened.

A hundred years of arguing about this number

Before putting a name on it, I should say that the underlying question is old, and that the way it has been mishandled is directly relevant.

In 1921 Sir Ian Hamilton, writing about how to organize an army, put it plainly: the average human brain finds its effective scope in handling from three to six other brains. He did not call it span of control. That phrase arrived about a decade later from Lyndall Urwick, and he arrived at it by open analogy — the psychological conception of the span of attention, he wrote, places strict limits on the number of separate factors the human mind can grasp at once, and it has an administrative counterpart in what may be described as the span of control. A limit on what one mind can hold, borrowed and renamed for what one manager can oversee.

Two things about the century since are worth carrying into the rest of this.

The first is that the number never settled, and the serious people always said so. Luther Gulick, canonizing the idea in 1937, immediately made it contingent — on how varied the work is, on how long the organization has been running, on how spread out it is. Modern research says the same thing with more data behind it. Reviews of the literature decline to name an ideal figure, and the studies that do produce numbers produce very different ones depending on the setting, with spans that work well in one kind of organization being far outside anything the classical writers would have accepted.

The second is what happens to a contingent rule when it travels. Urwick's most-quoted sentence is that no superior can directly supervise the work of more than five or six subordinates — and as he wrote it, that sentence ends whose work interlocks. The qualifier is the part that makes the claim true, and it is the part nearly every retelling drops. What survives the journey is a number, stated confidently, with the condition that earned it removed.

I have that firmly in mind, because I am about to name a limit and I would rather it did not travel the same way.

The Inspection Ceiling

Which leaves a constraint that needs a name, because it is the real one for anyone operating this way.

The Inspection Ceiling is the maximum volume of output a single accountable person can examine closely enough to genuinely stand behind. It exists because accountability requires answerability and inspectability in the same place, and inspectability is bounded by reading. It is personal, it varies by domain and with how well you know the material, and it is not large. Above all, nothing about the collapse in production costs touched it. Every other constraint in the operation moved; this one did not, and cheaper production is not the sort of thing that could have moved it.

Three things it is not, since the flattering misreadings are all sitting right there.

It is not a quality measure. A person can inspect carelessly, or with a blind spot, and still be comfortably under their ceiling. Being under it means the work was actually examined by the person answering for it, not that the examination was any good. The ceiling governs whether accountability is real, not whether judgment is.

It is not a productivity target, and raising it is not an achievement. An operation that pushes past its ceiling has not become more capable. It has stopped being accountable for the excess, because the name on that work now belongs to somebody who could not have read it. Output above the ceiling is not lower-quality output — much of it may be perfectly good — it is output that no longer carries the thing the name was supposed to certify.

And it is not visible from outside. Inspected and uninspected work are identical on the page. That was the whole argument of the previous issue and it applies to me precisely as it applies to everyone else: a reader holding one of these articles cannot tell which side of my ceiling it came from.

What this does to the case I have been making

So the claim has to narrow, and I would rather do it here than have it done to me.

"Operators beat pitchers" is not true as a general statement about suppliers. It is true within an operating range, and outside that range it reverses.

Below the ceiling the operator wins decisively, and for the reason this issue has been building toward: maker and answerer are one person, so nothing depends on a handoff, and omissions are at least catchable by the only person positioned to catch them. An agency cannot match that structurally — not because its people are worse, but because its accountability is distributed by design, and distribution is the thing that separates the halves.

Above the ceiling the operator is worse than the agency, and it is not close. An over-extended operator has one reader who has stopped reading properly and no second line at all. An agency at the same volume has several people who will each see part of it, and although their reading is unowned and catches only what reached the page, unowned reading of some of the work beats owned reading of none of it. My whole argument for concentration depends on the concentration being real. Past the ceiling it is nominal — and a nominal concentration of accountability is strictly worse than a distributed one, because it removes the redundancy as well.

The honest form is therefore conditional. Not that this model is better, but that it is better up to a volume, that the volume is lower than the marketing around this category implies, and that the same technology creating the advantage also creates a frictionless and continuous incentive to exceed it.

Which leaves the buyer with a problem I cannot solve for them

None of this is checkable from outside at the moment of buying.

You cannot infer it from the work, for the reason above. You cannot infer it from volume alone, since a high-output supplier may have a genuinely high ceiling, better instrumentation, or a narrower domain where they read faster. And you cannot infer it from the presence of a named accountable person, which is where the last issue's test bottoms out. A name concentrates the cost of being wrong on somebody findable, and that is still the strongest thing a reader has. But a name on work that person did not read is a name that could not have caught anything. Stake without inspection is answerability without inspectability — the first failure in this piece, wearing better clothes.

What I can offer is one question. Not a checklist; the buyer's field guide already used that format and running it again would be a format, not an argument.

Ask a supplier: what is the most you would take on and still stand behind every piece, and what happens when demand goes past it?

The first half asks for a number that costs money to say. The second half is where the information actually is. An answer describing a real constraint — we stop, we queue it, we decline, we bring in someone who becomes answerable for their own part — is a different kind of answer from one describing a reassurance. And a supplier who has never considered the question has not thereby proven they are over their ceiling. They have proven they do not know where it is, which for anyone producing at machine volume is the same exposure with less warning.

Can you not build machines to check the machines?

The obvious objection, and the honest answer is partly yes, which is why the part that is no needs to be exact.

We have built them. There is a check in our publishing pipeline that reads a draft for the defects that survive a fluent read: a cross-link that no longer resolves, a stated count that contradicts the heading above it, a coined term used throughout and never defined, a demonstrative pointing at the wrong antecedent because a later insertion pushed it away from the thing it referred to. The last of those exists because it kept happening, and because it is nearly invisible to someone re-reading their own prose, who knows what the sentence was meant to say. A second check watches the scheduled queue for posts whose approved copy has since changed. Each was written after a specific failure got through. Neither is clever. Both are load-bearing.

Two things are true about them at once, and the first is the strongest argument against my own position. They genuinely raise the ceiling. Mechanical checking is real, it compounds, and a well-instrumented operation can stand behind more than an uninstrumented one — which means the ceiling is not fixed, and an operation that invests in checking is buying back some of the capacity this piece says it lost.

And they do not touch the thing that matters most. Every check we have catches a defect of form. Not one of them can tell me that the argument is wrong, that a claim is unearned, that a better piece was available and nobody went looking, or that the framing serves us rather than the reader. Those are judgments, and a judgment cannot be delegated to the layer being judged.

There is also a recursion here that I do not think bottoms out. A check is itself a produced artifact, written under the same conditions as everything else, so if I am answerable for the work it clears then I am answerable for the check — including the failure where it passes something it should have caught, which is the one failure a passing check cannot report. Automating inspection does not dissolve the inspection problem. It moves the problem up a level and makes it quieter, because the thing you are now trusting is silent when it works and silent when it does not.

Where this reaches its limit

Four, and the last is the one I would attack if I were reading this.

The first is that the ceiling is not measurable and I cannot give you mine. There is no unit. It is felt rather than counted, and "I know when I am past it" is exactly the sort of self-report this log is normally rude about. I have no good defense, though I would point out that the failure is not peculiar to me: a century of management research has not produced a defensible number for the much easier question of how many people one person can supervise, and nothing credible offers one for how many agents a person can oversee. Anyone who hands you a figure for that is not drawing on evidence, because there is not yet any to draw on. What I would say for the constraint itself is that an unmeasurable one is still a constraint — my inability to state the number does not make output above it accountable, any more than not owning a scale makes the shelf hold.

The second is that the ceiling moves, and if it moves far enough the argument is overtaken rather than refuted. Tools that make evaluation genuinely cheaper — not generation, evaluation — would raise it, possibly a great deal, and I would rather name what would do that than be caught by it. A checking layer that reliably caught errors of judgment rather than errors of form would make most of this issue a historical note. Nothing I use does that today, and I am not confident about the direction of travel, because the same fluency that makes generated work hard for me to check makes it hard for a machine checker too.

The third is that I have been dismissive about redundant review in a way the argument does not license. Several unowned readers do catch things one owned reader misses, and what each catches differs rather than being strictly worse. The narrow claim I can defend is that distributed review is bounded by the artifact and therefore systematically blind to omissions — a specific and serious blind spot, not a general worthlessness. An agency with a strong review culture is a real thing and I should not write as though it were not.

The fourth is the sharpest. A declared ceiling is cheap to declare. I have just said that the honest move for a supplier is to name a limit out loud, and naming it is a claim like any other, subject to exactly the argument I made last month: it signals only if making it falsely is expensive. A number nobody tracks, in an operation nobody can audit, stated by the person who benefits from your believing it, is worth roughly what any other unverifiable assurance is worth. So the ceiling does not escape the stake problem. It inherits it.

For a long time capacity meant how much you could produce, and the interesting question about a supplier was how much work they could get through. Production stopped being the constraint. What did not change is that somebody has to read the thing and be on the hook for it, and there is a limit to how much of that one person can do. The question worth asking a supplier is no longer how much they can make. It is how much they can still answer for.


In one paragraph, and a few common questions

In one paragraph. Accountability is not a virtue but a structure: it requires that whoever pays when work is wrong is also whoever was positioned to catch it. Those halves separate easily. A hierarchy separates them by distance, since seniority is defined by distance from execution, and review does not repair it — review grades the artifact, and the expensive failures are omissions that leave no trace there. An agent-staffed operation separates them by volume: production collapsed in cost while reading did not, and unlike most automation this one makes checking dearer per unit, because fluent output disabled the cheap first-pass filter that used to triage attention. Two things I published do not survive that. Agents do not remove the ceiling on judgment, they move it from producing to reading; and an operator who skips the judgment does not fail visibly in the open, because the issue after that one established that uninspected work looks exactly like inspected work. What remains is the Inspection Ceiling — the most output one accountable person can examine well enough to genuinely stand behind — and a case for operators that wins below it and reverses above it, where an over-extended operator has one reader who has stopped reading and no redundancy at all.

What is the Inspection Ceiling? The most output one accountable person can examine closely enough to genuinely stand behind it. Accountability needs answerability and inspectability in the same place, and the second is bounded by reading, which did not get cheaper when production did. It is a capacity limit rather than a quality score: you can inspect badly and still be under your ceiling, and publishing above it does not make an operation more capable, it makes it unaccountable for the excess.

Didn't you argue the opposite of this a month ago? Half of it, and the correction is most of why this issue exists. In the piece on taste I argued that judging and making used to be the same job — your judgment reached about as far as your own hands did, so growing meant hiring producers, which handed out the judgment along with the production — and then that agents remove that ceiling. The first half holds: agents really do break the bundle between judging and making. The second was wrong, in a piece that was itself about constraints relocating rather than disappearing. Judgment was never bundled to making alone. It is also bundled to reading, agents did not touch that one, and to judge a thing you still have to take it in.

Isn't this just an argument that I should hire an operator instead of an agency? It is narrower, and since I have a commercial interest in the wider version, the narrow one is the one worth publishing. Below the ceiling an operator wins decisively, because maker and answerer are the same person and nothing depends on a handoff. Above it the operator is worse than the agency, which at least has several human readers where an over-extended operator has none. The advantage belongs to the operating range, not to the model.

Doesn't a big agency have more people checking the work? It has more readers, which is real and worth more than my side of this usually admits. What it lacks is readers who are answerable. Review operates on the artifact, and the artifact records what was done rather than what was not, so it catches the wrong claim on the page and structurally cannot catch the question nobody asked. Distributing the reading spreads the reading; it does not move answerability back to where the reading happened.

Can't you build automated checks so the machine reviews the machine? Partly, and we have. Scripted checks work well on defects of form — a dead link, a count contradicting its own heading, a term used and never defined — and they genuinely raise the ceiling, which is the strongest argument against my position here. They do not touch judgment, where the expensive errors live, and the move relocates accountability rather than removing it: you become answerable for a check you also did not write, including the one failure a passing check can never report.

How would I tell whether a supplier is above or below its ceiling? Not from the work, which is the uncomfortable part, since inspected and uninspected output look identical. So ask what is the most they would take on and still stand behind every piece, and what happens when demand goes past it. The information is in whether a number exists at all, and in whether the second answer names a real constraint or offers a reassurance.

Linara Bozieva, Founder, Ravenopus

The Engine Log

More like this in your inbox.

Operational artifacts, real protocols, real numbers. Sent when something is worth sending.

If the queue diagnosis applies to your current setup and you want to see what an agency without queues actually delivers, the 72-Hour Diagnostic is the smallest commitment we offer.

See the 72-Hour Diagnostic →