Every marketing service in 2026 claims AI, so the label no longer tells you anything. What does: 12 questions across four areas — accountability, humans, transparency, and economics — each with what a good answer sounds like and the red flag that should end the meeting. This checklist works on any provider, whether or not you ever talk to us.
The context that makes evaluation urgent: 54% of small businesses already use AI marketing tools and another 27% plan to within 12 months, according to a U.S. Chamber of Commerce survey compiled by Capsule (2026). Among small businesses using AI, 91% report revenue increases, per Salesforce research in the same compilation. Adoption is everywhere — but adoption is not the same as outcomes, and the gap between the two is exactly what these questions probe.
That gap has a practical shape. Two services can quote the same price, claim the same AI, and produce results that differ by multiples — because the difference lives in structure: who owns outcomes, who reviews the work, what you can see, and how the incentives point. None of that appears in a sales deck, which is why you have to ask.
Accountability: who answers when the number stalls?
These three questions establish whether anyone is on the hook for results. They come first because every other answer is decoration if nobody owns the number.
1. Who owns the outcome? AI can produce infinite activity, which makes it easier than ever for a provider to look busy while nothing moves. A good answer names a metric and a person: leads, qualified pipeline, CAC — and who is on the hook for it. Red flag: the outcome named is an output — posts published, impressions served, campaigns launched.
2. What happens when a metric stalls? Every growth program hits a plateau; what distinguishes providers is the documented response. A good answer describes a process: diagnose within a defined window, present the analysis, reallocate effort. Red flag: 'we would discuss it on the next monthly call' — a stall discovered in week one and addressed in week five costs you a month of spend.
3. Can I leave monthly? Commitment terms reveal how a provider expects to retain you. A good answer is month-to-month by default, with a clear offboarding path. Red flag: a 6–12 month minimum justified by 'results take time'. Results do take time — but a provider confident in month three should not need you locked in through month nine to prove it.
Humans: who is actually behind the AI?
AI is now table stakes; the human layer is where services actually differ. These three questions find out whether the humans are real, senior, and structurally involved — or a photo on the website.
4. Who reviews AI output before it ships — names and roles? Unreviewed AI output is where generic content, factual errors, and off-brand claims come from. A good answer gives you actual names, roles, and where they sit in the workflow — the way we list ours on /faq. Red flag: 'our team reviews everything' with no names, or review that turns out to be a spelling pass rather than judgment about whether the work is right.
5. How much senior time is in my tier? Cheap tiers everywhere run thinner human review — the honest question is how much thinner. A good answer quantifies the difference between tiers in strategist attention and review depth, in writing. Red flag: every tier claims the same 'full senior oversight' at wildly different prices. One of those tiers is mispriced, and it is not the expensive one.
6. What does the AI do, and what do humans decide? A provider who cannot draw this line crisply either does not know or does not want you to. A good answer is specific: AI drafts, analyzes, and produces variants; humans set strategy, approve what ships, and decide what changes. Red flag: vague gestures at proprietary AI doing everything — or, equally, AI as pure garnish on what is actually a traditional agency's manual work sold at AI-era prices.
Transparency: what do you actually get to see?
These questions test whether you can verify anything the provider claims without asking permission. The pattern to look for is visibility by default, not by request.
7. Do I see work before it ships? You are accountable for what goes out under your brand, whoever produced it. A good answer describes an approval flow with defined turnaround, plus a sensible default for routine items. Red flag: publishing without any approval path, or its mirror image — everything requires your sign-off, which quietly makes you the bottleneck they will later blame.
8. Live metrics or monthly slides? Reporting cadence is the difference between steering and archaeology. A good answer is a dashboard you can open any day, showing work shipped and results attributed — the standard we hold ourselves to with the Scale AI-hub. Red flag: a polished monthly PDF as the only window, which gives a provider four weeks of cover per underperforming decision.
9. Do I own the accounts and assets? Ad accounts, analytics, domains, content, audience data — everything should live under your ownership with the provider as an invited user. A good answer is an unhesitating yes, all of it, from day one. Red flag: accounts run under an agency umbrella you would lose by leaving. That converts switching costs into a retention strategy, and prices your exit into every month you stay.
Economics: what does the price actually buy?
The final three questions price the relationship over its whole life, not just the first invoice. Cheap entry pricing with opaque scaling is the oldest trick in services.
10. What exactly does the price include? AI-era services vary wildly in what a monthly fee covers. A good answer is a published list per tier — channels, volume, strategy cadence — like the one on our pricing page. Red flag: scope discovered by request, where every second ask turns out to be an add-on. That is not a subscription; it is an estimate with a monthly minimum.
11. What costs extra, and is ad spend mine? The clean structure: ad budget flows from your card to your accounts, and the service fee is the service fee. A good answer states this plainly and lists extras in writing before you sign. Red flag: ad spend routed through the provider — especially with percentage-of-spend pricing, which quietly rewards them for spending more of your money, not better.
12. How does pricing scale as I grow? You are not just buying this month; you are buying the price curve. A good answer ties price to defined tiers with published thresholds, so you know today what doubling volume costs. Red flag: 'we will figure it out as you grow' — which reliably means renegotiation from the weak side of the table, after your data and workflows already live inside their system.
How to run the evaluation
Send the 12 questions in writing before any demo, and hold answers to a simple bar: names, numbers, and processes are answers; adjectives are not. On the call, drill where the written answer was thinnest. Ask for evidence over claims — real client examples with attributable results, like the ones we publish in case studies — and weigh how the provider handles the questions themselves. One that welcomes scrutiny will behave the same way when a campaign stalls.
A simple scoring method keeps the comparison honest: two points for a specific written answer, one for a partial one, zero for adjectives. Providers cluster fast — most score well on economics and fall apart on accountability or transparency. Do not let averages hide that: a high total with a zeroed group is a structural no, because the groups do not compensate for each other.
Score all four groups, not just economics. A cheap service that fails accountability is not cheap — unowned outcomes cost more than fees — and how the categories trade off against agencies and platforms is a decision we mapped in agency vs. platform vs. GaaS.
Where to start
Start with whoever you already pay: your current agency, tool stack, or freelancer bench, scored against all 12 before you evaluate anyone new. That baseline turns vendor selection from a sales conversation into a comparison. Expect the exercise to take under an hour per provider and to disqualify most of a shortlist — that is the point: the cost of a wrong choice is not the fee, it is two quarters of stalled pipeline with a switching cost at the end.
For our part, Scalehackerlab publishes its own answers — reviewers, tiers, ownership, and what is included — on /faq and /pricing, and the model behind them is explained in what is Growth-as-a-Service. And if you want a second opinion on where your current setup stands, the free growth assessment returns a strategy document in 48 hours, no credit card required.
