To evaluate an AI consulting firm, ask for a live client system in production (not a demo), confirm in writing that you own the code and IP, get a defined post-launch support and maintenance model, and verify they run formal evaluation before go-live. A firm that answers all four concretely, with named references, is worth shortlisting. One that deflects on any of them is selling slides, not systems.
The AI consulting market in 2026 is crowded with firms that can produce an impressive demo in a week and a polished deck in a day. Very few can take a defined business problem, deploy a system that works on your real data, and hand it over so your team can run it. The gap between those two groups is where most failed AI projects live. The questions below are designed to surface that gap before you sign, not after. They are the same questions we would ask if we were the client.
For a structured starting point, see our AI consulting service and the way we scope engagements; this guide is written to be vendor-neutral so you can apply it to anyone, including us.
How to Use These Questions
Do not send these as a written questionnaire. Vendors will draft perfect answers. Ask them live, in a call, and listen for specificity. The signal is not whether they say "yes," it is whether they can produce a concrete example, a name, a number, or a contract clause on the spot. Vague answers to specific questions are the single most reliable predictor of a project that stalls.
Green Flags vs Red Flags at a Glance
| Signal | Green Flag | Red Flag |
|---|---|---|
| Proof of work | Live client system you can see or reference | Demos, prototypes, slideware only |
| IP and code | "You own everything, it is in the contract" | "We retain the platform / it runs on ours" |
| Pricing | Fixed scope or clear time-and-materials with a cap | Vague estimates that balloon mid-project |
| Evaluation | Named metrics and a pre-launch acceptance test | "We will know it works when we see it" |
| Team | The people on the call do the work | Senior closers, offshore juniors deliver |
| References | Two or three clients who will take your call | "Confidentiality prevents introductions" |
| Failure honesty | Will tell you what went wrong on a past project | Claims a perfect track record |
| Tooling | Current models and frameworks, named | Generic "we use the latest AI" |
The 12 Questions
1. Can you show me a working system in production, not a demo?
Why it matters. A demo proves nothing except that the firm can build a demo. Production systems handle messy real-world inputs, edge cases, integration failures, and load. The distance between a demo that works on three curated examples and a system that works on ten thousand real ones is where most of the actual engineering lives.
What a good answer looks like. They show you a live client deployment, or arrange a reference call where the client confirms the system is running and delivering value. They talk about specific problems they hit in production and how they fixed them. If everything they show is internal, sandboxed, or "almost ready," treat it as a prototype shop, not a deployment partner.
2. Who owns the code, the data, and the IP after delivery?
Why it matters. This is the clause that quietly traps companies. If the firm retains the codebase, or your system runs on their proprietary platform, you do not have a system, you have a dependency. Switching costs become a renewal weapon.
What a good answer looks like. Unambiguous: you own the source code, the models or configurations built for you, and all your data. It is written into the contract before work starts. Any hesitation, or a "platform fee" structure that means leaving them breaks your system, is a red flag. We cover the build-versus-buy ownership tradeoff in detail in AI agency vs in-house team.
3. What does post-launch support and maintenance actually cost?
Why it matters. AI systems are not static software. Models get deprecated, APIs change, your data drifts, and accuracy degrades if no one is watching. A system shipped and abandoned is a liability with a six-month fuse.
What a good answer looks like. A defined maintenance model with a real number attached: a monthly retainer, a support SLA, or a documented handoff so your team can maintain it. A common, reasonable benchmark is 15 to 25 percent of build cost per year for ongoing support. Firms that have not thought about maintenance have not deployed enough systems to know it is the hard part.
4. How do you test and evaluate before go-live?
Why it matters. "It seems to work" is not a launch criterion for a system that will talk to your customers or touch your money. Without a formal evaluation framework, you are shipping on vibes, and you will find the failure modes in production, in front of users.
What a good answer looks like. They describe an evaluation harness: a test set, accuracy or quality metrics relevant to your use case, edge-case and adversarial testing, and a defined acceptance threshold you both agree to before launch. For retrieval and knowledge systems, they should mention grounding and hallucination checks. See what is a RAG pipeline for the kind of evaluation that matters on knowledge-based systems.
5. Who, specifically, will be doing the work?
Why it matters. The classic agency bait-and-switch: senior experts win the deal, junior staff deliver it. In AI specifically, the difference between a senior engineer and a junior one is the difference between a system that handles edge cases and one that breaks on them.
What a good answer looks like. They name the people, their experience, and how much of the build those specific people will do versus oversee. Bonus signal: the people who will build it are on the sales call. If you never meet the delivery team before signing, assume you will not work with the people who impressed you.
6. How do you stay current with a field that changes monthly?
Why it matters. The model, framework, and cost landscape in 2026 looks nothing like 2024. A firm running on 18-month-old architectures will build you something already obsolete, often at higher cost and lower performance than current approaches.
What a good answer looks like. They name current models, current frameworks, and recent shifts that changed how they build. They can explain a decision they reversed because the tooling improved. Generic claims of "using the latest AI" without specifics mean they are not close to the field.
7. Can you give me two or three client references I can actually call?
Why it matters. Logos on a website are not references. A firm with real production work has clients who will vouch for it. A firm that hides behind blanket confidentiality usually has thinner delivery experience than its marketing implies.
What a good answer looks like. They offer references proactively, ideally in your industry or use-case category, and the references confirm the system shipped, works, and the firm was straight to deal with. Ask the reference one question: "What went wrong, and how did they handle it?" The answer tells you more than any case study. Review their published work too, such as our case studies.
8. What happens when the AI gets something wrong?
Why it matters. Every AI system makes mistakes. The mature question is not "will it be perfect" but "what is the blast radius when it is not." A firm that promises zero errors either does not understand the technology or is lying to you.
What a good answer looks like. They talk about guardrails, human-in-the-loop for high-stakes actions, confidence thresholds, fallback behavior, and logging that lets you trace what happened. They scope the agent or model so that a wrong answer is recoverable, not catastrophic. Comfort discussing failure is a sign of real deployment experience.
9. How do you scope and price the work?
Why it matters. Open-ended pricing on an unclear scope is how a project doubles in cost halfway through. You need to know what you are buying and what triggers more cost before you commit.
What a good answer looks like. Either a fixed price against a tightly defined scope, or transparent time-and-materials with a cap and clear change-order rules. They push back on vague requirements and help you define scope, because a firm that has delivered before knows that fuzzy scope is the number one cause of overruns. Compare structured options on our AI consulting plans page.
10. How will you integrate with our existing systems?
Why it matters. AI rarely lives alone. It connects to your CRM, your data warehouse, your support desk, your auth. Integration is usually where timelines slip, because real systems have undocumented quirks. A firm that treats integration as an afterthought will surprise you.
What a good answer looks like. They ask detailed questions about your stack early, flag integration risks before you raise them, and have done similar integrations before. They distinguish clearly between the AI work and the plumbing, and they budget realistically for both.
11. What data do you need from us, and how do you handle it?
Why it matters. AI quality is bounded by data quality, and data handling carries real legal and reputational risk, especially in regulated industries or jurisdictions with strict data-residency rules. A firm casual about your data is a firm that will eventually cause an incident.
What a good answer looks like. They are specific about what data they need and why, they raise privacy, residency, and compliance proactively, and they can explain where your data lives during development and whether it is ever used to train shared models. In regulated contexts they should know the relevant frameworks without you teaching them.
12. What happens if we want to leave or bring this in-house?
Why it matters. The healthiest vendor relationship is one you are free to exit. If leaving means your system stops working, you were never a client, you were a hostage. The exit terms reveal whether the firm is built on results or on lock-in.
What a good answer looks like. A clean answer: documentation, source code, deployment access, and a knowledge-transfer process so your team or a successor can take over. Good firms treat handoff as a deliverable and are comfortable with you eventually running things yourself, because they win the next project on merit, not on captivity.
How to Score the Answers
Run every shortlisted firm through all 12 questions and score each answer as concrete, vague, or evasive. You are not looking for twelve perfect answers; you are looking for a pattern. A firm that is concrete on proof of work, IP ownership, evaluation, and exit terms (questions 1, 2, 4, and 12) has the operational maturity that matters most. A firm that is vague or evasive on those four is high-risk regardless of how strong the rest of the pitch sounds.
Two questions are non-negotiable. If a firm will not put IP ownership in writing (question 2) or has no real evaluation process before go-live (question 4), stop there. Those two failures alone predict the most expensive outcomes: a system you cannot leave, and a system that does not work.
Frequently Asked Questions
What is the single most important question to ask an AI consulting firm?
Ask to see a working system in production with a reference you can call. Everything else is secondary, because a firm with live deployments has already been forced to answer the hard questions about integration, evaluation, support, and failure. Demos and decks prove only that a firm can market; production systems prove it can deliver.
How much should AI consulting cost in 2026?
A focused build, deploying a single well-scoped AI system, typically runs in the low tens of thousands to low six figures depending on complexity and integration depth. Budget another 15 to 25 percent of build cost annually for maintenance. Be wary of both extremes: quotes that seem too cheap usually mean a thin prototype, and open-ended pricing usually means scope that will balloon. For a structured view of how engagements are priced, see our AI consulting plans.
Should I hire an AI consulting firm or build an in-house team?
For most companies, a firm is faster, cheaper, and lower-risk in the first year, and the right choice unless AI is the core of your product. Build in-house only when the AI system itself is what you sell, or when your scale justifies a dedicated team. The full tradeoff, including a hybrid handoff model, is covered in AI agency vs in-house team.
How do I know if an AI consulting firm is a ChatGPT wrapper?
Ask question 4 and question 8. A thin wrapper has no real evaluation process and cannot explain its failure handling beyond "the model decides." A serious firm talks specifically about retrieval, grounding, guardrails, testing, and integration, and can show production systems that prove it. If their differentiation is access to a model anyone can buy, you are paying a markup for an API key.
If you are running this evaluation now, Iyara Labs is happy to answer all 12 questions on a call, on the record. You can also see how we structure engagements on our plans page or get in touch to discuss your use case.
