Disclosure. The draft this post is based on was written by Monid, a tool marketplace we work with as a content partner, and edited by us. We have an existing content exchange with them, described in Your Ad Agent Can't See the Auction It's Bidding Into. No money changed hands in either direction and the links below are the arrangement, stated plainly so you can weight them accordingly.
The question we get most often is not whether an agent can run a campaign. It can. The question is what has to be true before you let it.
That deserves a better answer than "start small". Here are the five checks we would run on any autonomous system before it touches live spend, including ours.
1. Can it tell you why a number moved?
This is the check almost nobody runs, and it is the one that separates an agent that compounds results from one that compounds noise.
Suppose cost per click rises 30% over two weeks and ROAS falls. There are two very different explanations. Your creative fatigued: the audience has seen it, click-through fell, and you are paying more for the same placement. Or the auction repriced: a larger advertiser expanded into your keywords and the clearing price moved for everyone.
Inside the ad account, those look identical. Same line, same direction, same size.
They call for opposite responses. The first says refresh the creative. The second says your creative is fine and the market changed, so the honest move is to re-evaluate which keywords are still worth holding at the new price.
An agent that can only see your account will pick one of those stories anyway, because it has to act. Worse, it will log a confident rationale, because from inside the account the reasoning genuinely was sound.
The fix is not a cleverer model. It is one more input. We went and measured this: using a paid-search competitor endpoint from Monid, one of the ai agent tools whose catalogue we pull from, we mapped the competitive set around two real advertisers and found that the median competitor shares only one to four keywords with you, while the largest bidder in a small advertiser's set was running three orders of magnitude more paid keywords than the median. Monid's own write-up of the same run is here. None of that is visible from inside the ad account, and all of it moves your costs.
What to check: ask whether the system can distinguish "my performance changed" from "the market changed". If it cannot, treat every optimization decision it makes as a hypothesis rather than a conclusion.
2. Does it have a floor it cannot argue its way below?
Guardrails the agent can reason about are not guardrails. They are suggestions.
The useful version is a hard constraint enforced outside the decision loop: a daily spend ceiling, a minimum bid, a list of campaigns it may not touch. Not a target it optimizes toward — a boundary it cannot cross regardless of how good its argument is.
The distinction matters because a well-constructed argument for exceeding a limit is exactly what a capable system produces under pressure. "Conversion volume justifies temporarily raising the cap" is a sentence that can be true, and it is also the sentence you will read in the log after an expensive weekend. The same reasoning applies to platform-level limits, which is why account safety rules belong outside the agent rather than in its prompt.
What to check: find one limit and try to talk the system past it. If a persuasive enough justification moves it, that limit is decorative.
3. Can you reconstruct a decision three weeks later?
Every autonomous system should be able to answer, for any change it made: what did you see, what did you conclude, what did you do, and what happened next.
This sounds like a compliance requirement. It is actually a debugging one. When results drift, the only way to find out whether the agent's model of your account is wrong is to read back a decision it made confidently and check whether the reasoning still holds now that you know the outcome.
Logs that record only actions are half the artefact. "Reduced bid on campaign 4 by 15%" tells you nothing you cannot see in the platform. "Reduced bid on campaign 4 by 15% because CPA rose 22% over seven days against a stable conversion rate" is something you can later prove wrong. This is the same standard we hold an audit agent to: a finding without its evidence chain is an opinion.
What to check: pick a decision from a month ago and try to reconstruct it end to end. If you cannot, you are not supervising the system, you are watching it.
4. What does it do when the data is missing?
This one is subtle and it is where most integrations quietly break.
When an agent asks an external system for something and gets back an empty result, there are two possible meanings. Either the thing genuinely does not exist, or the request was malformed and the system returned nothing without complaining. A surprising number of APIs return a perfectly valid, perfectly empty response to a request they did not understand, with no error at all.
An agent cannot tell those apart from the response alone. It will treat "your filter was ignored" as "there is nothing there", and reason forward from a fact that is not one.
For an ad agent this shows up as confident action on thin evidence: a segment paused because it "had no conversions", when the conversion query silently returned nothing. Anyone who has wired up the Google Ads API has hit the shape of this at least once.
What to check: feed the system a deliberately malformed input and watch what it does. An agent that says "I could not retrieve this" is safe. An agent that proceeds as though the answer were zero is not.
5. Who owns the rollback, and how fast is it?
The last check is organizational rather than technical.
Before autonomy goes up, somebody specific needs to own the answer to: how do we stop it, how long does stopping take, and what state are we left in. If the answer involves finding the person who set it up, autonomy is higher than the team's ability to supervise it.
The good version is boring. One switch, known owner, campaigns return to their last human-approved state, and the whole thing takes under a minute. On multi-account setups this is a structural question rather than a button, which is why multi-party approval flows are worth getting right before you raise the autonomy level rather than after.
What to check: actually perform a rollback once, on a small campaign, before you need it. A rollback nobody has tested is a plan, not a capability.
The pattern underneath all five
Every check above is a version of the same question: does the system know the limits of what it knows?
An agent that reports uncertainty, respects a hard boundary, records its reasoning, distinguishes missing data from absent data, and can be stopped cleanly is one you can give more room over time. An agent that is confident in all conditions is not more capable; it is less legible, and the confidence is the problem rather than the feature.
That is why Soku's control ladder starts at analysis only, moves to act with approval, and reaches full autopilot last. Not because the earlier stages are training wheels, but because each one is where you find out whether the five checks above actually pass on your account, with your data, before the answer costs you anything.
If you are evaluating autonomous ad tooling — ours, the platforms' own agents, or anything else on the marketing AI agents list — run these five. The ones worth using will survive them.
A note on the creative side
Worth saying plainly: none of the above is an argument against automating creative. Generation is genuinely the fast part now, and the constraint has moved to knowing which thing to generate.
If you want a concrete example of an outside signal that answers exactly that, what a competitor's ad run length tells you is a good one. An ad that has been running for ninety days is the closest thing to published proof that it converts, and it is sitting in a public archive. That is a better creative brief than most briefs, and it costs nothing but the lookup.










