15 May 2026 · 8 min
Not every problem needs a custom model. Sometimes an API call is right; sometimes it is a trap.
The most expensive AI mistakes are almost never technical. They are architectural, made in the first week, before a line of code is written. A team builds a custom model where a single API call would have done the job, burning months on something a vendor already offers. Or the reverse: a team leans on a per-token API for a high-volume, privacy-sensitive workload that was screaming to be self-hosted, and watches the bill and the compliance risk climb together. Both mistakes come from answering "build, buy or fine-tune" by instinct instead of by analysis.
There is no universal right answer, and anyone who gives you one without asking about your data is selling something. The answer depends on three things: how sensitive your data is, how much volume you are running, and how specialised your task is. Get clear on those and the decision usually makes itself.
For a general task, at low or moderate volume, on data that is not sensitive, a frontier API is almost always the right first move. It is the fastest path to something working, it requires no infrastructure, and you get the capability of the best models with none of the operational burden. Building anything custom in this situation is premature optimisation — solving a cost or privacy problem you do not yet have, at the expense of the speed you do need.
We will happily tell a client that the answer is a simple API integration and there is no project for us here. That is often the correct and honest recommendation, and pretending otherwise to manufacture work would be a disservice.
The calculus changes when two things are true: the task is narrow and repeated, and the data is something you would rather not send to a third party. For a well-defined job you run millions of times on sensitive inputs, a fine-tuned small model often matches the frontier API's quality on that specific task, at a fraction of the running cost, while keeping every record in-house. Here the volume justifies the up-front effort and the privacy profile justifies the control.
This is the case teams most often miss, because the API feels easier at the start. The API is easier at the start — and then the volume arrives, and the per-token cost that looked trivial in the demo becomes the largest line in the budget.
Some problems are not solved by any single model call, however good. Multi-step workflows, retrieval over your own proprietary knowledge, and regulated decisions that need documentation and oversight all require real engineering around the model: orchestration, retrieval layers, guardrails, evaluation and monitoring. The model is a component; the system is the product. Underestimating this is how a promising prototype fails to become a reliable feature.
The mistake here is the opposite of premature optimisation — it is assuming that because the model is capable, the system around it is trivial. It rarely is, and the gap between a working demo and a dependable system is exactly that surrounding engineering.
The single most useful exercise is to plot cost against volume for each option before deciding. An API has near-zero fixed cost and a linear per-request cost; a self-hosted model has a real fixed cost and a marginal cost near zero. These two lines cross somewhere, and where they cross depends entirely on your numbers. Below the crossover, the API wins; above it, self-hosting wins, and the gap widens with scale.
Teams routinely estimate this crossover point wrong, almost always placing it higher than it really is, because the API bill is invisible until it arrives while the hardware cost is visible up front. Plotting it honestly, against real projected volume, is what turns this from a guess into a decision.
Sometimes the cost curve is not even the deciding factor. If your data genuinely cannot leave your control — for legal, regulatory or contractual reasons — then self-hosting a model may be the only viable option regardless of volume, because the alternative is not "a bit more expensive" but "not allowed." In those cases the privacy requirement sets the architecture, and the cost analysis operates within that constraint rather than driving it.
Being clear about which factor is actually binding — cost, volume or privacy — is half the decision. They do not always point the same way, and knowing which one governs prevents optimising the wrong thing.
The honest recommendation is frequently not one of the three but a combination. A fine-tuned small model handles the high-volume common case cheaply and privately, while a frontier API is called only for the rare, hard cases that need it. Retrieval grounds a bought model in your own data without training anything. The strongest architectures mix approaches deliberately, using each where it is strongest, rather than forcing the whole problem into a single mould.
Our approach is to benchmark the honest options against your actual data and constraints before recommending anything, because the architecture that is cheapest on paper is frequently not the cheapest in production, and the approach that is fashionable this quarter is not necessarily the one your problem needs. Build, buy and fine-tune are not ideologies; they are tools with different cost, privacy and effort profiles. The right choice is whichever your own measurements support — and the only way to know is to measure before you commit, not after.
Book a 30-minute call. We will tell you honestly whether we can help.
Book a call