AI & Compute – Serving language models in production · cloud GPUs for LLM inference

LLM inference: tokens served, costs under control

Self-hosted language model endpoints with predictable latency and predictable costs – as an alternative to the token prices of third-party APIs. AI inference is the decisive step in which a trained model responds to new inputs – lightning-fast, precise and scalable. With large language models in particular it needs powerful, specialised hardware and high parallel processing: without the right infrastructure, even the best model becomes a bottleneck.

ccloud³ Console · inference endpoint
eu-de · Hallstadt Data Centre
GPU
RTX 6000 Ada
Latency p95
180 ms
Tokens/s
2,400
LLM vLLM · Llama 3 70B (Q4)
● Serving
API OpenAI-compatible endpoint
● Active
K8S Replicas · 2 → 4
● Scaled
MON Advanced Monitoring · latency
● Green
Load peak absorbedreplica added · 40 s
  • Scalable, secure and usable across industries – GPUs flexibly scalable, locally hosted and GDPR-compliant, for healthcare, finance or technology.
  • Massive computing power for complex AI models – parallel data processing for large language models, fast and reliable even under high load.
  • Reliable infrastructure – made in Germany – operated in the ISO-certified data centre near Bamberg: low latency, full GDPR compliance.
  • Guaranteed performance instead of oversubscription – dedicated GPUs, clear SLAs and locally connected infrastructure without shared performance.
Why centron

Inference as a separate infrastructure

Once a certain volume is reached, in-house inference outperforms any token API – in terms of cost, latency and data protection.

Dedicated inference GPUs

vLLM or TGI runs your models on RTX A4000 or RTX 6000 Ada – without neighbours, with constant latency.

Cost per hour instead of tokens

From €92.59 per month: With high volumes, this quickly pays for itself compared to token prices.

No prompt leaves the house

Customer data in prompts remains on your German infrastructure – eliminating the GDPR concern that plagues many AI projects.

Scales with usage

Multiple inference replicas behind a load balancer in Kubernetes scale in response to the volume of requests.

Use cases

AI that thinks along: practical inference for real challenges

The strengths of LLM inference show where people need support: analysing complex data, recognising subtle differences or deciding under uncertainty. All of this works without laborious programming, through intelligent learning from data – making AI a practical aid for better decisions and safer processes.

Healthcare

LLM inference in diagnostics or patient communication – models recognise subtle differences in complex data and support decisions under uncertainty; privacy-compliant on infrastructure for healthcare.

Financial sector

Analysis, advice or fraud detection: AI models give financial experts well-founded recommendations and warn of anomalies – on sovereign infrastructure for the financial industry.

Technology & industry

Process automation, quality control or system monitoring – models warn of system failures and help farmers detect diseased plants, for example; examples from manufacturing and predictive maintenance.

Knowledge management with RAG

Across industries within companies: knowledge management with RAG makes internal documents searchable via language model – implemented with centron cloud GPUs with high performance and privacy compliance; training and context data sit in the AI data lake.

API-compatible

Your own OpenAI-compatible endpoint

vLLM and Co. support the OpenAI API format – existing applications can switch to your own inference system via an endpoint URL. Models are stored in S3, deployments are managed by Kubernetes, and Advanced Monitoring keeps an eye on latency and utilisation.

from €92.59 per month (billed by the hour)In-house inference rather than token-based billing
Calculate costs using the price calculator
  • OpenAI-compatible – Drop-in via endpoint change
  • Constant latency – dedicated GPUs
  • Model Registry – Weights versioned in S3
  • Observable – Metrics & alerts included
Recommended modules

The right centron products

Customers typically implement this use case using these building blocks – which can be combined and expanded at any time.

Cloud GPU
Ab
92,59 € / Monat
NVIDIA performance
  • RTX A4000 from €92.59 per month
  • Quadro RTX 6000 from €170.83 · A100 from €489.47 · RTX 6000 Ada from €858.19 per month
  • Dedicated, not shared
  • No minimum term
Kubernetes
Ab
29,99 € / Monat
Container orchestration
  • AutoScaler included
  • Traffic at a fixed price
  • CI/CD-ready
Advanced Monitoring
Ab
4,99 € / Monat
Full visibility
  • Real-time metrics
  • Customised alarms
  • External checks
In a nutshell

How much does LLM inference cost at centron?

LLM inference on your own GPU infrastructure: Run language models with high performance and in compliance with the GDPR – NVIDIA A100 & RTX 6000 Ada from €92.59 per month. The cornerstone of the service is Cloud GPU from €92.59 per month – billed by the hour, with no minimum contract term. This is supplemented, as required, by Kubernetes and Advanced Monitoring. All data remains in Germany: our own data centres in Hallstadt near Bamberg, certified to ISO 27001 on the basis of IT-Grundschutz and BSI C5:2020 Type 1. New accounts receive a €200 starting credit.

Packages and prices
Building blockPrice
Cloud GPUfrom €92.59 per month
Kubernetesfrom €29.99 per month
Advanced Monitoringfrom €4.99 per month
AI & Compute FAQ

Frequently asked questions

What is LLM inference?

LLM inference refers to the process in which a trained large language model processes new inputs and then makes predictions, answers or decisions – for example generating a text, analysing content or classifying data.

Why is inference decisive for the practical use of AI?

Inference is the step in which AI shows its knowledge: it analyses new data and responds to it. Without fast, reliable inference a trained model remains theoretical – only high-performance inference makes it usable in everyday life, e.g. in chatbots, diagnostic systems or automated processes.

What role do GPUs play in LLM inference?

GPUs enable the parallel processing of large data volumes, which is essential for compute-intensive tasks such as LLM inference. GPUs therefore make the real-world use of LLMs feasible in the first place; inference on CPUs is theoretical only. They also improve response times and ensure efficient use of resources – especially with large language models.

When does running your own inference pay off compared with API providers?

As a rule of thumb: as soon as your token bill permanently exceeds the cost of a suitable GPU instance, or data protection rules out external APIs. An RTX A4000 from €92.59 per month already serves quantised models in production – check your volume in the price calculator.

What are the differences between GPU dedicated servers and cloud GPUs for LLM inference?

GPU dedicated servers offer maximum predictability and constant performance – but are less flexible when load peaks occur or additional resources are needed at short notice. Cloud GPUs, on the other hand, enable fast provisioning, flexible scalability and usage-based costs. For LLM inference this is often the more efficient solution, as additional GPUs can be added as needed and no long-term hardware investment is required.

How much does it cost to get started?

The starting prices are deliberately low: Cloud GPU from €92.59 per month, Kubernetes from €29.99 per month and Advanced Monitoring from €4.99 per month. New accounts receive a €200 starting credit valid for 60 days – you can calculate the cost of your specific configuration transparently using the price calculator.

What should companies look out for with cloud GPUs in Germany (data sovereignty, latency, oversubscription)?

With cloud GPUs in Germany, companies should above all pay attention to data sovereignty, transparent resource provisioning and a stable network connection. Computing resources should be operated in German data centres to ensure GDPR compliance and avoid risks from extraterritorial laws such as the US CLOUD Act. The centron cloud is additionally attested to BSI C5:2020 Type 1, creating traceability for security- and compliance-critical workloads. Low latency and guaranteed performance are equally important: many international providers work with oversubscription, where several users share a GPU. That can lead to fluctuating performance. centron offers scalable resources, clear SLAs and locally connected infrastructure so that LLM inference runs reliably, quickly and without performance drops.

Why should companies use centron’s cloud GPUs for LLM inference?

centron offers powerful cloud GPU instances from its ISO-certified data centre in Hallstadt. You benefit from high availability, fast scalability, GDPR compliance and a transparent cost structure – ideal for production AI applications in sensitive industries.

What advantages does cloud-based LLM inference offer over your own hardware?

Cloud GPUs enable flexible use as needed, avoid high investment costs and offer immediate access to scalable computing power. Especially with changing requirements or short-term projects, these are clear advantages over rigid on-premises infrastructure.

In which industries is LLM inference used most?

LLM inference is used in numerous areas: healthcare, e.g. for diagnostics or patient communication; the financial sector for analysis, advice or fraud detection; technology & industry for process automation, quality control or system monitoring; across industries within companies for knowledge management with RAG. With centron cloud GPUs these scenarios can be implemented with high performance and privacy compliance.

Can this be implemented in a way that complies with the GDPR?

Yes – that is the central advantage of your own inference: with API services from third countries, prompts and thus frequently personal content leave the EU. If the model runs on centron GPUs, inputs, context and outputs remain entirely in Germany and are not used for third-party training purposes. Operations take place in data centres certified to ISO 27001 based on IT-Grundschutz; for ccloud³ / Managed Cloud there is an unqualified BSI C5:2020 Type 1 attestation. Details can be found in the Trust Center.

Can companies rent servers from centron?

centron offers flexible server solutions for different requirements: virtual machines for scalable cloud infrastructures and managed servers for the supported operation of business-critical systems. Companies thus get exactly the infrastructure that fits their needs: from maximum technical flexibility to comprehensive all-round service.

Get started for free

Sign up and receive €200 credit at centron within your first 60 days.

This promotional offer applies to new accounts only. Available exclusively to businesses.