AI & Computing – Serving language models productively

LLM inference: tokens served, Keeping costs under control

Self-hosted language model endpoints with predictable latency and predictable costs – as an alternative to token-based pricing from third-party APIs.

Why centron

Inference as a separate infrastructure

Once a certain volume is reached, in-house inference outperforms any token API – in terms of cost, latency and data protection.

Dedicated inference GPUs

vLLM or TGI runs your models on RTX A4000 or RTX 6000 Ada – without neighbours, with constant latency.

Cost per hour instead of tokens

From €92.59 per month: With high volumes, this quickly pays for itself compared to token prices.

No prompt leaves the house

Customer data in prompts remains on your German infrastructure – eliminating the GDPR concern that plagues many AI projects.

Scales with usage

Multiple inference replicas behind a load balancer in Kubernetes scale in response to the volume of requests.

API-compatible

Your own OpenAI-compatible endpoint

vLLM and Co. support the OpenAI API format – existing applications can switch to your own inference system via an endpoint URL. Models are stored in S3, deployments are managed by Kubernetes, and Advanced Monitoring keeps an eye on latency and utilisation.

from €92.59 per month (billed by the hour)In-house inference rather than token-based billing
Calculate costs using the price calculator
  • OpenAI-compatible – Drop-in via endpoint change
  • Constant latency – dedicated GPUs
  • Model Registry – Weights versioned in S3
  • Observable – Metrics & alerts included
Recommended modules

The right centron products

Customers typically implement this use case using these building blocks – which can be combined and expanded at any time.

Cloud GPU
Ab
92,59 € / Monat
NVIDIA performance
  • RTX A4000 from €92.59 per month
  • Quadro RTX 6000 from €170.83 · A100 from €489.47 · RTX 6000 Ada from €858.19 per month
  • Dedicated, not shared
  • No minimum term
Kubernetes
Ab
29,99 € / Monat
Container orchestration
  • AutoScaler included
  • Traffic at a fixed price
  • CI/CD-ready
Advanced Monitoring
Ab
4,99 € / Monat
Full visibility
  • Real-time metrics
  • Customised alarms
  • External checks
In a nutshell

How much does LLM inference cost at centron?

LLM inference on your own GPU infrastructure: Run language models with high performance and in compliance with the GDPR – NVIDIA A100 & RTX 6000 Ada from €92.59 per month. The cornerstone of the service is Cloud GPU from €92.59 per month – billed by the hour, with no minimum contract term. This is supplemented, as required, by Kubernetes and Advanced Monitoring. All data remains in Germany: our own data centres in Hallstadt near Bamberg, certified to ISO 27001 on the basis of IT-Grundschutz and BSI C5:2020 Type 1. New accounts receive a €200 starting credit.

Packages and prices
Building blockPrice
Cloud GPUfrom €92.59 per month
Kubernetesfrom €29.99 per month
Advanced Monitoringfrom €4.99 per month
AI & Compute FAQ

Frequently Asked Questions

At what point does it become more cost-effective to carry out your own inference rather than use API providers?

As a rule of thumb: as soon as your token bill consistently exceeds the cost of a suitable GPU instance, or if data protection requirements preclude the use of external APIs. An RTX A4000, starting at €92.59 per month, can already handle quantised models in a production environment – check your volume in the price calculator.

How much does it cost to get started?

The starting prices are deliberately low: Cloud GPU from €92.59 per month, Kubernetes from €29.99 per month and Advanced Monitoring from €4.99 per month. New accounts receive a €200 starting credit valid for 60 days – you can calculate the cost of your specific configuration transparently using the price calculator.

Can this be implemented in a way that complies with the GDPR?

Yes – that is the key advantage of in-house inference: with API services from third countries, prompts – and thus often personal data – leave the EU. If the model runs on centron GPUs, inputs, context and outputs remain entirely within Germany and are not reused for training purposes. The data centres are certified to ISO 27001 on the basis of IT-Grundschutz; for ccloud³ / Managed Cloud there is an unrestricted BSI C5:2020 Type 1 attestation. Details can be found in the Trust Centre.
Read more

More on AI & Machine Learning

Get started for free

Sign up and receive €200 credit at centron within your first 60 days.

This promotional offer applies to new accounts only. Available exclusively to businesses.

Jetzt loslegen Sales kontaktieren