Every agency in London now offers AI. Very few can show you what happens when the model is wrong, and that is the only part that matters once a system is live in front of staff or customers.
Ask about evaluation before you ask about models
An assistant is only as trustworthy as the test set behind it. On our own retrieval work we hold a fixed set of real questions with known correct answers, and every change is scored against it before release. If a prospective partner cannot describe how they measure quality, they are shipping vibes. Ask to see an evaluation run, not a demo.
Ask who is allowed to see what
Retrieval systems fail on permissions long before they fail on language. If a document is restricted in your file store, the assistant must respect that restriction for every user, every time. Permission modelling is unglamorous engineering work and it is where most proofs of concept quietly break when they meet a real organisation.
The eleven questions
- —What does your evaluation set look like, and who wrote the correct answers?
- —How do you handle a question the system should refuse to answer?
- —Where does our data go, and is it used for training anywhere?
- —How are document permissions enforced per user?
- —What is the fallback when the model is unavailable?
- —How do you log what the system said, so we can review it later?
- —What does this cost per month at our actual volume, not at demo volume?
- —Who owns the prompts, the pipeline and the index?
- —How will you tell us this is the wrong use of AI?
- —What happens after launch: who monitors quality drift?
- —Can we speak to a client running this in production?
The useful answer to most AI briefs is a narrower system than the one being asked for.
Where AI has genuinely paid for itself
The wins we see in London are dull and specific: finding the right clause in a decade of documents, triaging inbound enquiries so a person answers the right ones first, turning an audit trail into an evidence pack, drafting from a house style rather than from nothing. Broad assistants impress in a boardroom and disappoint in a workflow.
We build AI systems from the Mayfair studio, alongside the engineering practice that has to run them afterwards. If you are shortlisting an AI agency in London, put these questions to all of us.
