Start Fast, Scale on Your Terms
Prototype with pay-per-token access to a curated model catalog, then move to dedicated GPUs once you need guaranteed capacity and predictable latency. Both run on the same platform, so there is nothing to migrate.
Prototype with pay-per-token access to a curated model catalog, then move to dedicated GPUs once you need guaranteed capacity and predictable latency. Both run on the same platform, so there is nothing to migrate.
Run inference in Exoscale's European and Swiss locations, with control over where your workloads and data are hosted. This helps you meet your data residency and GDPR obligations.
Both inference options are engineered for stable, predictable operation, with 99,95% SLA. Open interfaces and an OpenAI-compatible API keep your workloads portable and your integrations simple.
Choose the inference option that matches how much control you need and how your traffic looks. Dedicated and On-demand Inference share the same sovereign infrastructure and OpenAI-compatible API, with Batch Inference coming soon.
Reserve dedicated NVIDIA GPUs for any open-source or custom AI model, public, gated or private, from Hugging Face. Get an OpenAI-compatible endpoint with full isolation, per-second billing, and the option to scale to zero when idle. Zero infrastructure operations required.
DiscoverCall a curated catalog of production-ready models through a pay-per-token API. No GPU sizing, no capacity planning, no infrastructure to manage. Ideal for rapid prototyping, variable traffic, and teams who want AI capabilities without running dedicated GPUs.
DiscoverProcess large volumes of AI requests asynchronously for workloads that do not require real-time responses. Ideal for data enrichment, document processing, evaluations, and scheduled jobs. Coming soon.
Contact usYour data remains in Europe, ensuring GDPR compliance and sovereignty. Our data centers use renewable power and efficient cooling to reduce environmental impact significantly. Built in Europe, for Europe.
Pay per second for the GPU time you reserve with Dedicated Inference, or per token for what you consume with On-demand Inference. Transparent billing, no hidden fees.
Both Dedicated and On-demand Inference integrate easily with your existing AI tools. No vendor lock-in and a simple transition to our sovereign cloud infrastructure.
When you need it, Dedicated Inference runs on fully isolated infrastructure, guaranteeing maximum privacy and performance for your workload.
Get 24/7 support with a 30-minute response time, handled directly by Exoscale engineers in Europe.
Run your entire AI stack on Exoscale: GPUs, vector databases, inference, compute, storage, networking, and more, all fully integrated, cloud-native, and sovereign.
Pair Inference with the rest of Exoscale’s AI infrastructure and cloud services to build a complete, sovereign platform.
Accelerate AI training, fine-tuning, and volume inference using powerful NVIDIA GPUs. Scale from a single GPU to multi-GPU setups without complex orchestration. Perfect for machine learning, data processing, 3D rendering, inference, and scientific computing.
DiscoverEssential tools for modern AI, powering Retrieval-Augmented Generation (RAG) and semantic search workloads. Fully managed PostgreSQL with pgvector and OpenSearch-based vector search.
DiscoverGet the help you need to run your infrastructure with confidence through flexible support plans that provide expert guidance and guaranteed response times (SLA), ensuring our experts are there when you need them most.
DiscoverRunning AI inference in production needs a dependable partner. Our engineering and support teams help organizations across Europe deploy, scale, and move between Dedicated and On-demand Inference on Exoscale’s sovereign, sustainable cloud platform.
Contact usDedicated Inference reserves NVIDIA GPUs for you. You get full isolation, predictable latency, and per-second billing, and you can run any open-source or custom model from Hugging Face. On-demand Inference gives you pay-per-token access to a curated catalog of ready-to-use models, with no GPU sizing or capacity planning needed. Both share the same OpenAI-compatible API.
Yes. Both options run on the same sovereign infrastructure and expose the same OpenAI-compatible API, so moving a workload from On-demand to a dedicated GPU deployment does not require rewriting your integration.
On-demand Inference is ideal for rapid prototyping, applications with variable or unpredictable traffic, and teams that want to use production-ready AI models without sizing GPUs or managing infrastructure. You pay per token and can access models through an OpenAI-compatible API.
Batch Inference is designed for high-volume jobs that do not need an immediate response, such as document processing, data enrichment, evaluations, and scheduled workloads. It is coming soon on Exoscale.
Yes, with Dedicated Inference. You can deploy any public, gated, or private model from Hugging Face, including models from your own organization’s Hugging Face account, on dedicated GPU infrastructure.
Dedicated Inference is billed per second for the GPU time you reserve. On-demand Inference is billed per token, so you pay only for what you consume, with no upfront GPU cost. Both are transparent, with no hidden fees.
Yes. All Inference workloads run in European and Swiss data centers, so your data never leaves European jurisdiction. This applies to both Dedicated and On-demand Inference.
Yes. Both Dedicated and On-demand Inference expose an OpenAI-compatible API, so you can connect your existing tools and libraries without proprietary lock-in.