When a team says it wants to train AI on company data, it can mean three very different things: letting a model retrieve approved documents at answer time, fine-tuning a model on examples of a specialist task, or training a new model. Those routes have different data, infrastructure and governance requirements. Treating them as one purchase is how promising pilots become expensive risks.
Begin with the task, not the archive
Do not begin by moving every file into a training folder. Define one bounded task, the person accountable for it, the input they may use and the output they need. Record how the task works today so the AI version has a baseline. If nobody can agree what a correct result looks like, more data will not solve the problem.
Decide what is allowed to leave the room
Classify the material before it reaches a model. Public, internal, confidential, personal and regulated records need different treatment. Check ownership, consent, retention, provider training terms, data location and the access rights of each user. Remove duplicates and obsolete policies before they become confident but outdated answers.
- —Who owns every dataset and example?
- —Which users may retrieve each document today?
- —Can a person request correction or deletion?
- —Where are prompts, outputs and logs retained?
- —What must never be sent to an external model?
- —Who approves the answer before it affects a customer, employee or regulated decision?
Retrieval, fine-tuning and foundation training are not interchangeable
Retrieval is usually the first route for company knowledge because the source can be updated and cited without changing the model. Fine-tuning is useful for a repeatable specialist behaviour when you hold enough high-quality examples and can test on examples the model never saw. Training a foundation model requires a defensible dataset, substantial compute, research capacity and a reason existing models cannot meet the requirement.
The right question is not how much company data can be added. It is how little approved data is needed to prove one useful result.
Build the evaluation set before the model
Create representative questions, difficult edge cases, refusal cases and known correct answers before development begins. Keep some examples out of the build process. Score source support, factual correctness, missing information, review effort and safe refusal. Re-run the same set when the model, prompt, data or permissions change.
Choose the deployment from the risk
An approved business workspace may suit low-risk internal drafting. A controlled integration may suit company knowledge where the provider terms and access model are acceptable. Restricted workloads can justify self-hosting or an isolated environment. M3CA Sovereign AI provides a UK private-compute route for organisations that need their models and data away from the public internet.
Leave a system people can operate
The final deliverable is not only a model. It is an operating guide: owners, approved sources, user permissions, evaluation results, review steps, monitoring, incident handling, deletion and rollback. Train the people who will challenge the system, not only the people excited to use it.
M3CA helps teams choose between retrieval, fine-tuning and private deployment, then builds and evaluates the narrowest useful workflow. Bring one task and a safe sample to an AI consultancy session in our Mayfair studio.

