Almost every AI automation agency can show you something impressive in a thirty-minute call. Very few can show you something that has been running in a client's business for six months without them touching it. That gap is what this checklist is for. Work through these seven in order — the first one alone will remove most of your shortlist. Note that this is about choosing a firm to build and run systems for you, which is a different decision from vetting a consultant for an advisory engagement; if you are earlier than that, the questions are different.
1. Ask for proof of production, not a demo
This is the single highest-signal question you can ask, and it is one sentence: show me three systems you have shipped in the last six months that are still running.
Note what it rules out. A demo proves the technology works, which was never in question. A case study written by the agency proves they can write. A system still in production six months later proves it survived real data, real edge cases, and a client who stopped paying attention to it — which is the only thing that matters.
A good answer is specific and slightly awkward: what it does, what broke early on, what it took to fix. An agency that has genuinely shipped will have war stories. One that has not will stay at the level of capability and enthusiasm.
2. Check they integrate rather than replace
Integration has become the qualifying test rather than a nice-to-have. If a proposed system does not connect to the CRM, inbox, and scheduling tools your team already lives in, adoption fails regardless of how good the automation is.
Be direct about it: ask which of your existing systems they will connect to, and which they would want you to replace. An agency that proposes replacing your stack is either solving a genuine architecture problem or selling you a bigger project. Make them tell you which.
Ask specifically about the system you would least like to touch. Every business has one — the old CRM, the spreadsheet the finance team guards. How they answer that question tells you whether they have worked with real businesses or only clean ones.
This is the kind of system we build as an AI automation agency for US businesses — scoped to the process, not sold as a seat licence.
3. Watch how they scope before they quote
A number produced before anyone has looked at your systems is a guess, and guesses are corrected upward later. The sequence you want is: diagnostic, then scope, then price.
What good scoping looks like in practice:
- They ask about volume and failure cost, not just what you want built.
- They ask to see the actual process, not a description of it.
- They tell you which parts they would not automate, and why.
- They raise data quality early, because that is what drives cost and timeline.
- The proposal names what is explicitly out of scope, not only what is in.
4. Require a baseline and a target in writing
Ask what number will move, by how much, and how it will be measured. If they cannot answer, the engagement has no definition of success — which means it cannot fail, and it also cannot be defended when the budget is reviewed.
Lack of measurable success criteria is one of the most reliable predictors of a project that stalls after the pilot. The baseline has to be captured before work starts, because once automation is running the old number is unrecoverable and any estimate made afterwards will flatter the result.
The specific commitment to look for: an agency willing to write down a target, in the proposal, with a date.
5. Settle ownership before you sign
Establish who owns the accounts, the automations, the prompts, and the data — and get it in writing at kickoff, when it costs nothing. Raised in month four it becomes a migration project.
The failure here is rarely malicious. It is faster for an agency to build inside their own platform account and their own API keys. Everything works until the relationship ends and you find the system your operations depend on lives somewhere you cannot reach.
Specify that every account is created under your billing and your domain, with the agency added as a user. If an agency resists that, you have learned something important early.
6. Ask what happens after go-live
An automation is not a finished object. Models change, APIs get deprecated, and the process it automates will drift. Ask what ongoing support covers and what it costs, before the build rather than after.
Two questions cut through the vagueness. First: if this breaks at 7am on a Monday, who notices — you or us? A system with no monitoring will fail silently, and silent failure in an automated process is far more expensive than an obvious one. Second: what does support actually include — how many hours, what response time, and what counts as new work rather than maintenance?
A retainer quoted with no definition beyond the word support is unpriced. Make them define it.
7. Start with a paid pilot, not a programme
The most effective way to de-risk this decision is to make the first engagement small enough that being wrong is survivable. A scoped pilot of 30 to 90 days, with agreed proof points and a cap on the first build, tells you more than any amount of reference-checking.
It also reveals the things references never do: whether they communicate when something slips, whether estimates hold, and whether the documentation is real. Those are the qualities that determine whether a long engagement works, and you cannot assess them from a proposal.
Resist the pitch for a large upfront commitment with vague deliverables. An agency confident in its work will accept a small first project, because it expects to earn the second one.
The disqualifiers
Any one of these is worth ending the conversation over.
- Demo-only proof — nothing running in production they can point to.
- A fixed quote given before anyone has looked at your systems.
- Systems only they can log into, described as managed for you.
- No monitoring, so failures surface when a customer complains.
- Over-scoping — a six-figure programme proposed for a problem you described in one sentence.
- No discussion of what happens to your data, or where it is processed.
Once an agency passes these checks, the next question is budget — what an AI automation agency actually costs covers the real 2026 ranges by pricing model.
Key takeaways
- Ask for three systems shipped in the last six months that are still running — it removes most shortlists in one question.
- Integration with your existing stack is the qualifying test, not a feature; adoption fails without it.
- Diagnostic first, then scope, then price. A quote before anyone has looked at your systems is a guess.
- Require a written baseline and target — no success metric is the most reliable predictor of a stalled project.
- Start with a paid 30-90 day pilot; it reveals communication and documentation quality that references never will.
Related services
AI Workflow Automation
We architect and build intelligent workflow systems that connect your tools, execute decisions, and run processes around the clock — without any human input.
Explore automation infrastructureCRM Automation
We turn your CRM into an intelligent revenue engine — automated pipelines, smart follow-ups, AI scoring, and reporting that shows what to do next.
Explore crm intelligence