Self-hosted language model endpoints with predictable latency and predictable costs – as an alternative to token-based pricing from third-party APIs.
Once a certain volume is reached, in-house inference outperforms any token API – in terms of cost, latency and data protection.
vLLM or TGI runs your models on RTX A4000 or RTX 6000 Ada – without neighbours, with constant latency.
From €92.59 per month: With high volumes, this quickly pays for itself compared to token prices.
Customer data in prompts remains on your German infrastructure – eliminating the GDPR concern that plagues many AI projects.
Multiple inference replicas behind a load balancer in Kubernetes scale in response to the volume of requests.
vLLM and Co. support the OpenAI API format – existing applications can switch to your own inference system via an endpoint URL. Models are stored in S3, deployments are managed by Kubernetes, and Advanced Monitoring keeps an eye on latency and utilisation.
Customers typically implement this use case using these building blocks – which can be combined and expanded at any time.
LLM inference on your own GPU infrastructure: Run language models with high performance and in compliance with the GDPR – NVIDIA A100 & RTX 6000 Ada from €92.59 per month. The cornerstone of the service is Cloud GPU from €92.59 per month – billed by the hour, with no minimum contract term. This is supplemented, as required, by Kubernetes and Advanced Monitoring. All data remains in Germany: our own data centres in Hallstadt near Bamberg, certified to ISO 27001 on the basis of IT-Grundschutz and BSI C5:2020 Type 1. New accounts receive a €200 starting credit.
| Building block | Price |
|---|---|
| Cloud GPU | from €92.59 per month |
| Kubernetes | from €29.99 per month |
| Advanced Monitoring | from €4.99 per month |
Sign up and receive €200 credit at centron within your first 60 days.
This promotional offer applies to new accounts only. Available exclusively to businesses.