Skip to main content
Register

Why Run Inference on Exoscale?

Start Fast, Scale on Your Terms

Start Fast, Scale on Your Terms

Prototype with pay-per-token access to a curated model catalog, then move to dedicated GPUs once you need guaranteed capacity and predictable latency. Both run on the same platform, so there is nothing to migrate.

European Data Sovereignty

European Data Sovereignty

Run inference in Exoscale's European and Swiss locations, with control over where your workloads and data are hosted. This helps you meet your data residency and GDPR obligations.

Built for Production

Built for Production

Both inference options are engineered for stable, predictable operation, with 99,95% SLA. Open interfaces and an OpenAI-compatible API keep your workloads portable and your integrations simple.

Inference Options

Choose the inference option that matches how much control you need and how your traffic looks. Dedicated and On-demand Inference share the same sovereign infrastructure and OpenAI-compatible API, with Batch Inference coming soon.

Dedicated Inference

Dedicated Inference

Reserve dedicated NVIDIA GPUs for any open-source or custom AI model, public, gated or private, from Hugging Face. Get an OpenAI-compatible endpoint with full isolation, per-second billing, and the option to scale to zero when idle. Zero infrastructure operations required.

Discover
On-Demand Inference

On-Demand Inference

Call a curated catalog of production-ready models through a pay-per-token API. No GPU sizing, no capacity planning, no infrastructure to manage. Ideal for rapid prototyping, variable traffic, and teams who want AI capabilities without running dedicated GPUs.

Discover
Batch Inference

Batch Inference

Process large volumes of AI requests asynchronously for workloads that do not require real-time responses. Ideal for data enrichment, document processing, evaluations, and scheduled jobs. Coming soon.

Contact us

What You Get With Either Option

Complete Your AI Stack

Pair Inference with the rest of Exoscale’s AI infrastructure and cloud services to build a complete, sovereign platform.

GPU Cloud Computing

High-Performance GPUs

Accelerate AI training, fine-tuning, and volume inference using powerful NVIDIA GPUs. Scale from a single GPU to multi-GPU setups without complex orchestration. Perfect for machine learning, data processing, 3D rendering, inference, and scientific computing.

Discover
Managed Vector Databases

Managed Vector Databases

Essential tools for modern AI, powering Retrieval-Augmented Generation (RAG) and semantic search workloads. Fully managed PostgreSQL with pgvector and OpenSearch-based vector search.

Discover
Support Plans

Support Plans

Get the help you need to run your infrastructure with confidence through flexible support plans that provide expert guidance and guaranteed response times (SLA), ensuring our experts are there when you need them most.

Discover

Trusted by Engineering Teams Across Europe

Running AI inference in production needs a dependable partner. Our engineering and support teams help organizations across Europe deploy, scale, and move between Dedicated and On-demand Inference on Exoscale’s sovereign, sustainable cloud platform.

Contact us

Frequently Asked Questions

What is the difference between Dedicated Inference and On-Demand Inference?

Dedicated Inference reserves NVIDIA GPUs for you. You get full isolation, predictable latency, and per-second billing, and you can run any open-source or custom model from Hugging Face. On-demand Inference gives you pay-per-token access to a curated catalog of ready-to-use models, with no GPU sizing or capacity planning needed. Both share the same OpenAI-compatible API.

Can I move from On-Demand to Dedicated Inference as my usage grows?

Yes. Both options run on the same sovereign infrastructure and expose the same OpenAI-compatible API, so moving a workload from On-demand to a dedicated GPU deployment does not require rewriting your integration.

When Should I Choose On-Demand Inference?

On-demand Inference is ideal for rapid prototyping, applications with variable or unpredictable traffic, and teams that want to use production-ready AI models without sizing GPUs or managing infrastructure. You pay per token and can access models through an OpenAI-compatible API.

What is Batch Inference for?

Batch Inference is designed for high-volume jobs that do not need an immediate response, such as document processing, data enrichment, evaluations, and scheduled workloads. It is coming soon on Exoscale.

Can I use my own AI models on Exoscale?

Yes, with Dedicated Inference. You can deploy any public, gated, or private model from Hugging Face, including models from your own organization’s Hugging Face account, on dedicated GPU infrastructure.

How is Inference priced?

Dedicated Inference is billed per second for the GPU time you reserve. On-demand Inference is billed per token, so you pay only for what you consume, with no upfront GPU cost. Both are transparent, with no hidden fees.

Is Inference on Exoscale GDPR compliant?

Yes. All Inference workloads run in European and Swiss data centers, so your data never leaves European jurisdiction. This applies to both Dedicated and On-demand Inference.

Does Inference support open standards?

Yes. Both Dedicated and On-demand Inference expose an OpenAI-compatible API, so you can connect your existing tools and libraries without proprietary lock-in.