Buyer guide
12 questions to ask an AI agency before you hire one
Most AI vendor calls are decided by whoever demos best, which selects for the wrong thing. These twelve questions, from Digiton's enterprise engagements in Europe, select for a firm that ships.
The twelve questions, and what a bad answer sounds like
Read the table first. Then take the four or five that matter most for your situation into the call, because asking twelve in one hour turns a conversation into an interrogation and you learn less.
| # | The question | What a good answer contains | What should worry you |
|---|---|---|---|
| 1 | Show me a system you have in production right now. Who uses it and how often? | A named system, a rough user count, a usage frequency, and what it does when it fails | A demo video, a prototype, or a case study whose outcome is "increased efficiency" |
| 2 | Who gets paged when it breaks at 3am, and what is your incident process? | A named person or rota, an alerting path, and a worked example of an incident they handled | "You can email support" or a promise to add monitoring later |
| 3 | Who owns the code, the prompts and the evaluation set at the end? | You do, in writing, including the evaluation set and the golden dataset | Ownership of the code but not the prompts, or an evaluation set that is "part of our methodology" |
| 4 | What is the scope tied to numerically, and who measures it after launch? | One number, a baseline measured before the build, and a named person inside your organisation who reads it monthly | A number with no baseline, or a measurement plan that starts after go-live |
| 5 | How do you handle GDPR erasure when the data is inside embeddings and agent memory? | A specific mechanism: source-linked chunk IDs, a re-index path, and a documented retention rule for agent memory | "We delete the record from the database", which leaves the embedding and the conversation history intact |
| 6 | Do your deployed systems meet Article 50 transparency duties, which have applied since 2 August 2026? | Yes, with the specific mechanism: disclosure on first contact, machine-readable marking of synthetic media, deepfake notices where relevant | Any answer that starts by telling you the AI Act deadline moved to December 2027 |
| 7 | What is your agent governance model: scoped identity, least privilege, human approval on irreversible actions, audit log, kill switch? | Five concrete answers, ideally with the audit log schema described | A single sentence about "guardrails" with no mechanism underneath it |
| 8 | What does the handover look like if we bring this in-house in year two? | A named handover package: runbook, evaluation set, architecture notes, and a fixed number of shadowing weeks | Visible discomfort, or an offer to "discuss that closer to the time" |
| 9 | Which parts of this are your own code and which are a vendor you resell? | A clear split, with the vendor named and the licence terms disclosed | A platform diagram with no vendors on it |
| 10 | What did the last engagement that went badly look like, and what did you change? | A real story with a specific mistake and a specific process change | "We have not had one", which is either untrue or means they have not shipped much |
| 11 | How do you evaluate model output, and can I see the golden dataset? | A named evaluation approach, a set of cases drawn from real failures, and a rerun that happens on every change | Manual spot-checking, or evaluation that happened once before launch |
| 12 | What is your answer to a client security questionnaire asking for model provenance and data residency? | A written answer sheet they already have, covering which models, hosted where, trained on what, and who can read the logs | A promise to "look into it" for a question their other clients have certainly already asked |
1. Show me a system you have in production right now. Who uses it and how often?
This is the question that ends most conversations honestly. A firm that has shipped will answer in specifics inside thirty seconds, because the system is a thing they think about. A firm that has not will reach for a demo, a pilot, or a client who cannot be named for reasons that turn out to be vague.
Push once on frequency. Daily use by a team of forty is a different engineering problem from a system three people open in a quarter, and only one of them has met real failure modes.
2. Who gets paged when it breaks at 3am, and what is your incident process?
An AI system is a running service. It drifts, it costs money per call, and it fails in ways a normal web application does not, because a model provider can change a response shape or deprecate a version with a few weeks of notice. Somebody has to own that.
The useful follow-up is to ask for the last incident they handled and what changed afterwards. A firm that has operated anything will have one. The answer also tells you whether they run their own software or only build other people's.
3. Who owns the code, the prompts and the evaluation set at the end?
The code is the easy part and almost everybody concedes it. The prompts and the evaluation set are where lock-in actually lives. An evaluation set built from your own historical failures is the single most valuable artefact of the engagement, because it is what lets the next team change anything without breaking what already works.
Get it named in the contract as a deliverable, not as a by-product.
4. What is the scope tied to numerically, and who measures it after launch?
Hours removed, cycle time, error rate, queue length, revenue. Pick one. The baseline has to be measured before anything is built, because retrofitting a baseline afterwards produces whatever number the firm needs it to produce.
Then ask who reads it in month six. If the answer is the agency, the number is marketing. If the answer is somebody in your own operation, it is a metric.
5. How do you handle GDPR erasure when the data is inside embeddings and agent memory?
This is the question that separates people who have shipped a retrieval system in Europe from people who have read about one. A deletion request has to reach every derived copy: the vector store, the cached retrieval context, the agent's conversation memory, and any evaluation set built from production traffic.
Most builds handle the first and quietly miss the rest. The engineering detail is on GDPR-compliant RAG, and it is worth reading before the call so you can tell a real answer from a confident one.
6. Do your deployed systems meet Article 50 transparency duties, which have applied since 2 August 2026?
The deadline that moved is not this one. Regulation (EU) 2026/1744 pushed Annex III high-risk obligations to 2 December 2027, and it left Article 50 exactly where it was. Transparency duties went live on 2 August 2026 and they bind today, on ordinary systems, including the chat assistant somebody put on a website last spring.
A firm that answers this question with the December 2027 date has misread the regulation in the direction of its own convenience. There is also a second date to know: systems placed on the market before 2 August 2026 have until 2 December 2026 for content marking. Full timeline on EU AI Act readiness for enterprises.
7. What is your agent governance model: scoped identity, least privilege, human approval on irreversible actions, audit log, kill switch?
An agent with a shared API key and write access to production is an incident waiting for a bad afternoon. Each of the five has a specific shape. Scoped identity means the agent authenticates as itself and not as a service account somebody else also uses. Least privilege means read access by default and write access granted per action. Human approval covers anything irreversible, which includes sending mail, moving money and deleting records.
The audit log is the one to press on, because it is the artefact an auditor reads. Inputs, retrieved context, model version, output, approver. More detail on AI agent governance.
8. What does the handover look like if we bring this in-house in year two?
The answer matters much less than the reaction. A firm confident in its work will describe the handover cheerfully, because it has done a few and because clients who leave well come back. A firm whose commercial model depends on you never leaving will handle the question differently and you will feel it.
Ask for the handover package to be a deliverable in phase one, produced as the build goes, rather than a document somebody writes in a hurry at the end.
9. Which parts of this are your own code and which are a vendor you resell?
There is nothing wrong with reselling. There is something wrong with not being told, because a resold component brings its own pricing changes, its own outage history, and its own exit problem. If forty percent of the delivered system is somebody else's platform, your renewal negotiation is with them and not with the firm you signed.
Ask what happens to your data inside the vendor's system, and where it sits.
10. What did the last engagement that went badly look like, and what did you change?
Everybody who has delivered enough has a bad one. The useful signal is whether the lesson turned into a process. "We now measure the baseline before the build because we could not prove the value on that project" is a good answer. "The client changed their mind" is a firm blaming a client to a prospective client, which tells you what they will say about you.
11. How do you evaluate model output, and can I see the golden dataset?
Without an evaluation set, nobody can tell whether a prompt change made the system better or quietly worse on the cases that matter. The set should grow from production: every correction a human makes becomes a case, so the system is measured against its own history rather than against a benchmark somebody published.
Ask when it last ran. Ask what the score was. A firm that cannot answer is shipping on vibes.
12. What is your answer to a client security questionnaire asking for model provenance and data residency?
Enterprise procurement now asks this regardless of regulation, and the question has moved down-market fast. Model provenance means which model, from which provider, hosted in which region, and whether your data touches training. Data residency means where the vectors, the logs and the conversation history physically sit.
A firm selling into enterprises has been asked this a dozen times and has a document. A firm that has not been asked is telling you something about its client base. The vendor-side view is on AI vendor due diligence.
The three questions that changed in 2026
Nine of the twelve above would have been on this list two years ago. Three would not.
Article 50 became a live obligation. Transparency duties under the EU AI Act applied from 2 August 2026 and were untouched by the Omnibus deferral. Disclosure when a person is talking to an AI, machine-readable marking of synthetic media, deepfake notices. This binds ordinary systems, not only high-risk ones, and most vendors have not updated their answer.
The December 2027 deferral became a way to spot who reads primary sources. Regulation (EU) 2026/1744 entered into force on 27 July 2026 and moved Annex III standalone high-risk obligations from 2 August 2026 to 2 December 2027, with Annex I product-embedded systems going to 2 August 2028. It is not eighteen free months. Classification is the prerequisite for every later obligation, and a firm that treats the deferral as a reprieve will hand you the same problem in a year with less time to fix it.
The security questionnaire arrived early. Model provenance and data residency questions used to appear at contract stage in large enterprises. They now appear in the first procurement pass at mid-market firms, because those firms are being asked the same questions by their own customers. A vendor without a written answer sheet costs you weeks.
How to score the answers
Give each question three points for a specific answer with a mechanism in it, one point for a plausible answer with no mechanism, and zero for deflection. Questions 1, 2, 4 and 6 are the load-bearing ones, so double them.
Above 32 out of 48 and you are talking to a firm that has shipped. Between 20 and 32, you are talking to a good team who have not yet operated anything, which is fine for a first build if you have the internal capacity to run it afterwards. Below 20, you are buying a pilot.
One caution about the scoring. A firm can rehearse answers, and the better ones will have. What cannot be rehearsed is the follow-up: ask for the incident, the number, the audit log schema, the evaluation score from last week. Specifics are the whole test.
Where our own answers sit
Digiton's content and citation systems produced 365 Copilot citations across 30 days from 249 indexed pages, measured in Bing Webmaster Tools, which is what a production content system looks like when somebody keeps it running. The same applies to the client work: we build and then operate, which is why questions 2, 8 and 11 are on this list rather than the easier ones.
If you are working through a shortlist, how to vet an AI agency covers the reference-checking side, and enterprise AI consulting in Europe sets out what production delivery looks like once the shortlist is down to two.
Frequently asked questions
What should I ask an AI agency before hiring them?
Start with a system running in production today and the name of the person who gets paged when it breaks. Then ask who owns the code, the prompts and the evaluation set at the end, what the scope is tied to numerically, how they handle GDPR erasure across embeddings, and whether their deployed systems meet Article 50 transparency duties. A firm that cannot answer those six is selling a pilot.
What is the single most revealing question?
Who gets paged when it breaks at 3am. Building a working agent has become the easy half of this work. Keeping one correct through model deprecations, schema changes at the other end of an API and eighteen months of drift is where engagements succeed or quietly stall, and a firm that has never operated anything will answer with a support email address.
Did the EU AI Act deadline move, and does it change what I should ask?
It moved for one category only. Regulation (EU) 2026/1744 entered into force on 27 July 2026 and pushed Annex III standalone high-risk obligations to 2 December 2027 and Annex I embedded systems to 2 August 2028. Article 50 transparency, the general-purpose AI obligations and the Article 5 prohibitions did not move. Ask about Article 50, because that one binds today.
How do I check whether an agency really has systems in production?
Ask for a named system, a rough user count and a usage frequency, then ask what it does when it fails. A firm that has shipped answers in specifics inside thirty seconds. Follow up by asking for the last incident they handled and what changed in their process afterwards, because anyone who has operated software has one and anyone who has not will improvise.
Should the agency own the prompts and the evaluation set?
No, you should, and it should be written into the contract as a deliverable. Code ownership is conceded by almost everybody. Prompts and the evaluation set are where lock-in actually lives, because an evaluation set built from your own historical failures is what lets any future team change the system without breaking what already works.
What is a fair way to score the answers?
Three points for a specific answer with a mechanism in it, one for a plausible answer with no mechanism, zero for deflection, and double the weight on production evidence, incident ownership, the scope metric and Article 50. Above 32 of 48 means the firm has shipped. Below 20 means you are buying a pilot. Then ignore the score and press on one follow-up, because specifics cannot be rehearsed.
Related
Ready to put AI to work?
Book a discovery call and we will map the highest-value AI agents and automations for your business.
Book a discovery call