Private cloud for AI is a data gravity decision
The business case says security and compliance. The actual driver is usually that the data is already somewhere, and moving it is expensive.
Every business case I have read for keeping AI workloads on private infrastructure leads with security and compliance. Almost none of them are actually about security and compliance.
The real driver is usually simpler and less presentable: the data is already somewhere, there is a lot of it, and moving it is expensive in ways that do not fit neatly on a slide.
What gets written down, and what decided it
Security is the reason that survives contact with a board. It is legible, it maps to obligations someone in the room already owns, and nobody argues against it.
“Moving this would take four months and cost more than the project” is the reason that actually decided it. That one invites questions about whether the estimate is right, whether it could be phased, whether the vendor would discount the transfer. So it appears further down the document, if at all.
The distinction matters because the two reasons lead to different architectures. If the driver is genuinely regulatory, the question is which controls satisfy the obligation, and the answer may well be a public cloud region with the right certifications. If the driver is gravity, no amount of control design changes anything — the data is where it is.
Gravity is three separate costs
Worth separating, because they have different fixes:
- Egress. Charged per gigabyte, and the figure people quote is usually the one-time transfer. The recurring cost of a pipeline that keeps pulling from the same source is the one that surprises.
- Time. A large estate moves over months, not weekends, and the source system carries on producing during it. Cutover planning is most of the work.
- Permission. Some of it legally cannot move, or cannot move without a review nobody has budgeted for. This is the one that turns a migration into a two-year program.
Only the first is a number people put in a model. The second and third are what actually stop projects.
What you give up
The honest version of this argument has to say what private infrastructure costs you, because it is not nothing.
Elasticity. Training is bursty. Inference is not. Buying for the peak means paying for idle capacity between peaks, and buying for the average means the training runs queue.
Model access. The frontier models are API-only. An architecture that keeps everything inside your perimeter is an architecture that uses open-weight models, and the capability gap is real even as it narrows.
Pace. Managed services ship features you would otherwise build. That is a genuine cost, paid quarterly, and it compounds.
The test worth applying
If the data could move tomorrow at no cost, would you still keep this on private infrastructure?
A yes means the driver is genuinely control or obligation, and the design should follow from that. A no means it is gravity, and the honest architecture is hybrid: keep the data where it is, move the smallest possible thing, and be explicit that the boundary is drawn by economics rather than by policy.
Both are defensible. Only one of them is usually what the document says.
Working on something this touches?
Start a conversation