Enterprise AI has reached an inflection point. Based on some of the discussions I’ve had with leaders recently, a clear consensus has emerged: the challenge is no longer gaining access to powerful models, but deploying them with control, resilience and economic certainty. Public cloud remains an important part of the enterprise architecture, yet it cannot satisfy every requirement for sensitive data, proprietary workflows or regulated operations. Sovereign AI therefore is not an alternative to cloud; it is the control layer that enables organisations to place each workload where it delivers the greatest value.
Beyond data residency and security, local deployment can also improve performance and resilience. Locating compute closer to the data—whether in the data centre, at the edge or in an air-gapped environment—can reduce latency, limit exposure to network disruption and support workloads that cannot depend on continuous external connectivity. The appropriate architecture will differ by workload: some applications belong in public cloud, some on premise or at the edge, and many across a governed combination of environments. For organizations, the strategic question is no longer “cloud versus local”, but which combination provides the required control, performance, resilience and cost predictability.
The proof is in the performance
To give organisations a more reliable basis for those placement decisions, we have worked with NVIDIA to benchmark our Large Language Model deployments. The objective is not simply to publish peak figures, but to show how model choice, hardware configuration and deployment architecture influence the performance enterprises can expect in operational environments.
The strategic lesson from our benchmarking is that sovereignty and enterprise-grade performance are not opposing choices. A well-designed deployment can provide predictable scaling across locations, retain operational control over sensitive workloads and tune model performance to the accuracy, latency and throughput requirements of a specific use case. Those trade-offs—not a single headline result—should guide infrastructure decisions.
- Predictable Scalability: Deliver consistent AI performance across your organization, from the data center to the edge.
- Operational Control:Reduce dependence on external connectivity and keep the most sensitive data within the infrastructure and governance boundaries defined for the workload.
- Efficiency: Achieve high-precision results optimized for your specific industry requirements.
Explore our benchmarking whitepapers here:
Private GPT v1.7: More connected, more scalable
Private GPT v1.7 is designed for the next stage of enterprise adoption: moving from isolated conversational use cases to AI that can work securely with operational systems and scale across distributed organisations. Its significance lies less in adding another assistant and more in creating a governed bridge between proprietary data, enterprise applications and local AI services.
Key features include:
- Model Context Protocol (MCP) Connectors:Model Context Protocol connectors allow Private GPT to connect consistently with databases, ERP platforms, wikis, ticketing systems and more. This shifts the role of the assistant from retrieving knowledge to participating in governed workflows. The opportunity is significant, but so is the need for clear permissions, auditability and human approval when an AI system can trigger actions rather than simply return information.
- Trusted Provider Model:The model combines a central Core for governance and management with distributed Satellites that provide secure, local access. This gives multi-location enterprises and service providers a way to standardise policy and operations while keeping selected workloads close to users and data. For SMEs, it also creates a path to managed sovereign AI without requiring each organisation to build and operate a dedicated GPU estate.
Expanding beyond: Introducing OSAI
Private GPT addresses an important set of enterprise use cases, but no single solution can become the foundation for every AI workload. OSAI extends the proposition into a modular sovereign AI ecosystem developed in Europe, allowing organisations to assemble capabilities around their own data, governance requirements and operating model rather than adopting a fixed, vertically closed stack.
With a vast range of building blocks to choose from, OSAI enables organisations to start with a defined business problem and add capabilities as their requirements evolve. The breadth of the solution is intended to provide choice across models, deployment patterns and services rather than requiring every workload to conform to one fixed architecture.
Crucially, OSAI is modular by design. Partners can extend the core solution with specialised capabilities, enabling organisations to tailor solutions to complex requirements without committing every AI workload to one model, one deployment pattern or one provider.
Agentic infrastructure management: Introducing Mamoru
As AI moves from generating content to taking action, infrastructure management becomes a critical test of trust. Mamoru applies agentic automation to complex, multi-vendor IT environments while retaining human authority over consequential actions. Automated oversight and remediation can reduce operational burden, but expert validation remains part of the control model. This combination of machine speed and human accountability is designed to keep AI-powered infrastructure stable, compliant and aligned with business priorities.
The road ahead
The next phase of enterprise AI will not be defined by one model or one deployment destination. It will be defined by an organisation’s ability to combine cloud, local and edge resources under consistent governance; connect AI safely to business systems; and scale without surrendering control of data, operations or economics. OSAI is our answer to that challenge: a modular foundation for turning selected AI use cases into governed business capabilities.
The practical next step is to identify one high-value workload and assess it against four questions: where must its data remain, what level of latency and resilience does it require, which systems must it connect to, and where must a human retain approval? Explore our benchmark whitepapers for deployment evidence, register for the Fsas Technologies Tech Community for upcoming technical deep-dives, and meet us at NVIDIA GTC Berlin to discuss how these principles can be applied to your AI roadmap.