← The Engine Log

How to Buy Marketing in the Agent Era

Every buyer of marketing runs a checklist they never wrote down: how big is the team on my account, how many deliverables a month, what's the turnaround, how much of the senior person's time do I get. Those questions all meter production, which is the one input that just stopped being scarce. The claim here is stronger than "they're outdated." A proxy that loses its scarce input does not decay to noise — it reverses, and starts pointing at the partner who kept the most of the old cost structure. This is the Proxy Inversion, and it comes with four replacement questions a buyer can run in a single call.

Linara Bozieva17 min read
Watercolor illustration: the Ravenopus stands between two market stalls holding an old brass measuring instrument on a chain. The left stall is stacked to the ceiling with identical crates and the instrument's needle swings hard toward it. The right stall holds one small closed box on an otherwise empty counter, and the needle barely moves for it — though the thing the Ravenopus actually came to buy is the box. The instrument is not broken; it is working perfectly, and pointed at the wrong quantity.

Every buyer of marketing services runs a checklist, and almost nobody has written theirs down.

It comes out in the questions asked on the second call. How big is the team on my account. How many deliverables a month. What is the turnaround. How much of the senior person's time am I actually getting, and who does the work when they are on another account. Buyers ask these without deliberating, the way you check the mileage on a used car, and the questions are not arbitrary or stupid. They are the residue of a market in which producing the work was the expensive part. When making the thing is what costs money, everything that meters the making — people, hours, units shipped, how long you wait — is a decent stand-in for how much capability you are buying.

That is the part that changed, and I want to make a claim about it that is stronger than "the old questions are outdated."

They have not weakened. They have reversed.

What the checklist was actually measuring

A proxy earns its place by satisfying two conditions at once. It has to correlate with the input that actually determines the outcome, and it has to be expensive to satisfy. Miss the first and it measures nothing. Miss the second and everyone maxes it out, so it stops telling you who is who.

Headcount on the account cleared both bars for as long as anyone buying today has been buying. So did deliverables per month, hours logged, and turnaround time. Production capacity was the binding constraint on what a marketing partner could do for you, people were how you bought capacity, and nobody could conjure a second strategist on Tuesday. The numbers were costly, so they were credible, and they tracked the scarce input closely enough that a buyer could use them and be roughly right.

Both conditions are now gone, and it is worth separating them because they fail in different ways.

The second one fails loudly. Deliverable counts are trivially satisfiable now. An operation running agents can produce forty assets in a month at very close to the effort of producing four, which means any supplier facing a buyer who meters output can simply give them output, in volume, immediately, forever. A metric that everyone can max out on demand is not a low-quality signal. It is not a signal. The buyer who keeps grading on production will get production, and will learn nothing about the supplier in the process — including, and this is the expensive part, nothing about whether any of it should have been made.

The first condition fails quietly, which is why it does more damage.

The Proxy Inversion

Here is the step I think most of the discussion skips.

When the scarce input behind a proxy becomes cheap, the intuitive expectation is that the proxy decays toward noise. Its correlation with outcomes drifts to zero, the number stops meaning anything, and a buyer who keeps using it is merely wasting a question. Annoying, but harmless.

That is not what happens, and the reason is that the variance does not disappear. Suppliers still differ on headcount and hours and account-team size — often by a lot. If those differences no longer track capability, they have to track something else. And what they now track, mostly, is how much of the old cost structure a given supplier has retained.

Take two suppliers producing comparable work for you. One carries twelve people on your account; the other carries a much smaller number of people plus a system that does the production. Under the inherited checklist the first scores higher on nearly every line and can defend a higher price, because the buyer's instrument reads capacity and the first has more of it. But capacity was not the constraint on the outcome. What actually varies between those two is how much coordination the buyer is funding — the status meetings, the internal reviews, the handoffs between the person who decided and the person who executes, the interval in which your work sits in a tray waiting for a specific human to become free. That interval is not anyone's craft and it is not anyone's labor. It has never been valuable to the person waiting either. It is the queue, and the old checklist now prices it as a feature.

That is the Proxy Inversion: a buying signal that has lost its scarce input does not go silent, it changes sign, and starts recommending the supplier who has most preserved the conditions that made the signal necessary in the first place.

There is an honest defense of the old proxy and I want to concede it in full, because it is the strongest thing a buyer can say back to me. Headcount does still buy something real. It buys redundancy, coverage, and availability: a bench when someone is sick, a person reachable at nine at night, continuity when the individual you trust moves on. Those goods did not stop existing because production got cheap, and a buyer who needs them is right to pay for them. But notice that the thing being bought is insurance, not capability — and insurance should be priced, disclosed, and tested as insurance. The failure is not paying for coverage. The failure is reading a coverage number as a quality number, which is what the inherited checklist does automatically.

Four questions that meter the right thing

The remedy is not to be more skeptical of the answers to the old questions. It is to ask questions whose answers are still expensive to fake. These four are the ones I would run, and they fit inside one call.

1. What gets cheaper for you next year, and what happens to my invoice when it does?

This meters the pricing structure, which is where a supplier's incentives are actually written down. If you are paying for reserved time, then falling costs are a margin event for them and a non-event for you: the structure requires holding the price and absorbing the difference, whatever anyone intends. The shape that does not create that conflict is a flat monthly fee, where a drop in the cost of production shows up as more work against the same number rather than as a slowly negotiated discount. Listen for whether they can describe what their own structure does under falling costs without being asked twice. A supplier who has never thought about it is telling you their pricing was inherited, not designed.

2. Show me something you killed.

The single highest-yield question on this list. A portfolio is a record of what got approved, and in an era where making the work is nearly free, a full portfolio is nearly free to accumulate — it is production, and production no longer discriminates. What remains scarce is the judgment to reject competent work, and rejection leaves no artifact unless someone kept it deliberately. So ask for it: a campaign they took to completion internally and never shipped, a concept the client liked that they argued against, a channel they turned down. Then ask what the tell was. The good version is specific, slightly uncomfortable, and costs them something to tell you. The answer to be wary of is a version of "we don't really kill things, we test everything" — that is not taste, it is judgment outsourced to an experiment you will be paying to run, and most buyers cannot afford to run enough of them to substitute for a point of view.

3. How will we find out we were wrong, and how long will that take?

This meters two things at once: whether the operation closes its loops, and whether it measures honestly when it does. The unserious answer arrives as a dashboard promise. The serious one names the comparison — a holdout, a geo split, a staged rollout — or concedes there will not be a clean one for this particular channel and says what they will do instead. That concession is a good sign, not a bad one. Almost every marketing decision is read without a control arm, and a supplier who tells you that upfront is more likely to be reading their own results honestly than one who never mentions it. Push the second half of the question too, because the speed of the loop is itself the asset: a partner who finds out they were wrong in nine days and one who finds out in nine weeks are not the same purchase, and no scope of work will show you the difference.

4. Did you make this, or did you approve it?

You already know the answer to the older version of this question. The people who show up to win the account are not the people who service it, the industry has a name for it, and most buyers stopped asking years ago because the answer never changed anything. The gap between the pitch and the delivery is structural in any organization that separates business development from execution, and no assurance in the room ever fixed it: the assurance costs nothing to give, and the person making it is not there when it is broken.

What changed is that the claim is now falsifiable. An operator-led model's entire pitch is that the seam is closed — the person you assess is the person who does the work. That is a far bigger claim than the old one and it is just as cheap to say, and because buyers were trained by experience that this question has no useful answer, almost nobody checks it. It is the least-tested claim in the new model.

There is also a second version of the gap that belongs to this era specifically, and that I would want checked if I were the one buying. In an agent-staffed operation the seam does not disappear. It relocates — it stops running between senior and junior people and starts running between the operator and the output. Someone can hand you work their system produced and that they never seriously examined, and it will look exactly like work they made. Judgment can be skipped as easily as it can be delegated, and the artifact does not record which one happened.

That is what the walk-through is for. Ask the person in front of you to open something they are presenting as their own and walk you through a decision inside it — why they cut that section — and then ask what the version before it looked like and what changed. Do not read the first answer. Anyone competent can narrate a finished rationale, and presenting other people's work is a professional skill that good account leads have for a reason. What does not survive the second question is the history: the approach tried first, the constraint that killed it, the person in the room who disagreed. Makers carry that residue because they lived it, and none of it gets written into a brief. So read the direction the answers move as you push. Someone who did the work gets more specific. Someone who was briefed on it gets more general, and starts describing the strategy instead of the object open in front of you.

The ChatGPT-seat test

One more, and it is the fastest.

Most suppliers now say they use AI, which makes the claim useless on its own. The question that separates them is: what breaks if I take the AI away?

Take it away from an operation that has added tools to an existing structure, and that operation slows down. Drafts take longer, research takes longer, the org chart is otherwise unchanged, because the tools sit at the end of a process that was designed around people producing work. Nothing structural goes missing — a step simply got expensive again. And I do not want that read as dismissal: a good tool in a good process is a real gain worth paying for, and a chat-facing assistant doing serious work is not a lesser thing than a pipeline. But it is an accelerant on the old structure, and it inherits every wait state that structure had.

An operation where a function is genuinely staffed by agents answers differently, because the function does not exist without them. There is no manual fallback to slow down to. Whether that is better for you depends on the work, but it is a different thing you are buying, with different failure modes, and it costs you one question to establish which. Ask the follow-up too: where in the process does it sit — at the end, drafting, or as the process itself, with a named owner and something written down that says what it checks before anything reaches you.

What a retainer should actually cost

I do not think anyone can give you a defensible rate card right now, and I would treat one as evidence against the person handing it to you, because it was built from labor inputs that no longer dominate the cost.

What can be said is which shapes are broken on their face. Pricing per deliverable meters the thing that just became free, and it now creates an incentive to flood you, which is easy and looks like value. Pricing per headcount or per hour meters reserved time, which is the queue itself — you are literally buying the interval you are trying to escape, and the supplier's revenue rises with the amount of waiting they create. Both were sensible when production was scarce. Neither survives contact with a market where it is not.

What is left is a flat monthly fee for an accountable operation. The question that actually matters is where a buyer takes their share of the falling costs. One way is to take it in price: renegotiate the fee down each year on the grounds that the work got cheaper to produce. The other is to take it in scope: the fee holds, and what it covers keeps growing — more surfaces, more tests, shorter cycles for the same money this year than last. I have an obvious commercial interest in recommending the second, so here is the argument rather than the assertion. What got cheaper was production, and production is not what you are buying. A discount argued from "AI made this cheaper" therefore comes out of the judgment and the accountability, which did not get cheaper — and the only lever a supplier has to absorb it is to put less thought into your account. You win the negotiation and lose the engagement.

The flat fee is the shape I use, and I will not pretend it is comfortable for a buyer: it is harder to justify on a spreadsheet, because the thing you are paying for — judgment, accountability, and the rate at which the operation corrects itself — has no unit. The correct response to that discomfort is not to demand a fake unit. It is to buy a small, bounded piece of the judgment before you commit to any of it, which is the entire reason a paid diagnostic exists as a product category, and to keep the commitment period short enough that being wrong is survivable.

Since it would be cheap of me to write a buyer's guide and hide my own answers to it: Ravenopus runs a flat monthly retainer starting at twenty thousand dollars, three-month minimum, one to two new engagements a month, with a one-time fifteen-hundred-dollar diagnostic as the way in. Those numbers are on the site, so they are checkable — and the numbers are the least of it. Run every question above at me exactly as you would at anyone else.

Where this reaches its limit

A guide like this is gameable, and it will be gamed. Everything above is a question, and questions can be rehearsed — this piece is public, and a supplier who reads it can prepare a good answer for "show me something you killed" the way candidates prepare a weakness. The defense is the same in every case and it is why each question above asks for an object: request the artifact, not the account of it. The rejected concept, the actual measurement plan, the file they made. Answers are cheap to produce now. Artifacts with a history are not.

"Cheaper" is the wrong takeaway, and it is the most likely one. A buyer who reads all of this and concludes they should simply pay half will find someone who meters every one of the same wrong things at a discount — the same purchase with a smaller invoice and no better odds. Judgment and accountability are what is actually being bought, and they are the two things the inherited checklist never had a line for, which is exactly why they are the two easiest to negotiate away without noticing.

If you need capacity rather than judgment, none of this applies. There are real buyers whose strategy is settled, whose positioning is decided, and who need hands, coverage, and a name to call on a Friday afternoon. For them, capacity is the thing being bought, the old proxies measure it correctly, and they should keep using them. The inversion only bites when a buyer is paying for judgment and grading the supplier on production — which, in my experience of watching these decisions get made, is most of them, and almost none of them on purpose.

The thing to hold onto is that the old checklist was never wrong. It was a good instrument, honestly built, pointed at the quantity that used to matter. It is still working. Nobody moved the needle; the market moved out from under it. What a buyer has to do now is decide what they came to measure, and then find out whether the instrument in their hand has ever been able to see it.


In one paragraph, and a few common questions

In one paragraph: Buyers of marketing evaluate suppliers with an unwritten checklist — team size on the account, deliverables per month, hours, turnaround, senior time — and every line of it meters production, which was a fair proxy while producing the work was the expensive part. It is not any more, and the failure is worse than obsolescence. A proxy needs to correlate with the scarce input and to be expensive to satisfy; production metrics now fail both, and because suppliers still vary on them, the leftover variance measures how much of the old cost structure each one retained rather than how much capability it has. So the signal reverses and begins recommending the supplier funding the most coordination, which is the queue, priced as a feature — the Proxy Inversion. The honest defense of headcount survives but narrows: it buys redundancy, coverage, and availability, which are real, and should be priced and tested as insurance rather than read as quality. What replaces the old checklist is four questions whose answers are still costly to fake — what happens to my invoice when your costs fall, show me something you killed, how will we find out we were wrong and how fast, and did you make this or approve it. That last one is not the old bait-and-switch check, which every buyer already assumes the answer to; it tests the claim an operator-led model actually makes, that the seam is closed, and it catches the newer seam that runs between an operator and their own output. Fastest of all, and costing a single question, is what breaks if I take the AI away. On price, no defensible rate card exists, but per-deliverable meters the free thing and per-hour meters the wait, leaving a flat monthly fee — where the buyer's share of falling costs is better taken in scope than in price, because a discount argued from cheaper production comes out of the judgment and the accountability, which never got cheaper. The limits are real: a public guide gets rehearsed, so ask for artifacts and not answers; "cheaper" is the wrong conclusion, because judgment and accountability never got cheap; and a buyer who genuinely needs capacity rather than judgment should keep the old proxies, which measure capacity correctly.

What is the Proxy Inversion? What happens to a buying signal when the scarce thing it stood in for becomes cheap. Production metrics can now be maxed out by any supplier on demand, so they stop discriminating, and their remaining variance measures retained cost structure rather than capability. The signal does not go quiet — it points the other way.

How do I tell an AI-native operation from an agency with a ChatGPT seat? Ask what breaks if you take the AI away. A seat makes an unchanged process slower; an agent-staffed function has no manual fallback to slow down to. A seat is a genuine gain, not a scam — it is just an accelerant on the old structure, and you should know which one you are buying.

What should a retainer cost in 2026? Nobody can hand you a defensible rate card, and one is evidence against whoever does. Per-deliverable meters the thing that became free; per-hour meters the queue. A flat monthly fee is what is left, with your upside as rising output against a fixed number — tested first with a small bounded purchase rather than a formula.

Isn't a small operation just concentrated risk? Yes, and it is the honest cost of the shape. Ask what is documented, what runs without the operator in the room, and what happens to your account if they are unreachable for two weeks. A buyer who needs guaranteed coverage should weight that heavily and pay for it deliberately.

Shouldn't AI make marketing cheaper for me? The production half already did, and it is priced in. Deciding what is worth making and being accountable for the outcome did not get cheaper, and that is what the retainer buys. Look for more output and faster correction against a flat fee rather than for a smaller invoice.

What if I just need execution? Then ignore all of this. If the strategy is set and you need hands and coverage, capacity is what you are buying and the old proxies measure it correctly. The inversion only bites when you pay for judgment and grade on output.

Linara Bozieva, Founder, Ravenopus

The Engine Log

More like this in your inbox.

Operational artifacts, real protocols, real numbers. Sent when something is worth sending.

If the queue diagnosis applies to your current setup and you want to see what an agency without queues actually delivers, the 72-Hour Diagnostic is the smallest commitment we offer.

See the 72-Hour Diagnostic →