LLM inference: tokens served, costs under control
Self-hosted language model endpoints with predictable latency and predictable costs – as an alternative to the token prices of third-party APIs. AI inference is the decisive step in which a trained model responds to new inputs – lightning-fast, precise and scalable. With large language models in particular it needs powerful, specialised hardware and high parallel processing: without the right infrastructure, even the best model becomes a bottleneck.
- Scalable, secure and usable across industries – GPUs flexibly scalable, locally hosted and GDPR-compliant, for healthcare, finance or technology.
- Massive computing power for complex AI models – parallel data processing for large language models, fast and reliable even under high load.
- Reliable infrastructure – made in Germany – operated in the ISO-certified data centre near Bamberg: low latency, full GDPR compliance.
- Guaranteed performance instead of oversubscription – dedicated GPUs, clear SLAs and locally connected infrastructure without shared performance.
Inference as a separate infrastructure
Once a certain volume is reached, in-house inference outperforms any token API – in terms of cost, latency and data protection.
Dedicated inference GPUs
vLLM or TGI runs your models on RTX A4000 or RTX 6000 Ada – without neighbours, with constant latency.
Cost per hour instead of tokens
From €92.59 per month: With high volumes, this quickly pays for itself compared to token prices.
No prompt leaves the house
Customer data in prompts remains on your German infrastructure – eliminating the GDPR concern that plagues many AI projects.
Scales with usage
Multiple inference replicas behind a load balancer in Kubernetes scale in response to the volume of requests.
AI that thinks along: practical inference for real challenges
The strengths of LLM inference show where people need support: analysing complex data, recognising subtle differences or deciding under uncertainty. All of this works without laborious programming, through intelligent learning from data – making AI a practical aid for better decisions and safer processes.
Healthcare
LLM inference in diagnostics or patient communication – models recognise subtle differences in complex data and support decisions under uncertainty; privacy-compliant on infrastructure for healthcare.
Financial sector
Analysis, advice or fraud detection: AI models give financial experts well-founded recommendations and warn of anomalies – on sovereign infrastructure for the financial industry.
Technology & industry
Process automation, quality control or system monitoring – models warn of system failures and help farmers detect diseased plants, for example; examples from manufacturing and predictive maintenance.
Knowledge management with RAG
Across industries within companies: knowledge management with RAG makes internal documents searchable via language model – implemented with centron cloud GPUs with high performance and privacy compliance; training and context data sit in the AI data lake.
Your own OpenAI-compatible endpoint
vLLM and Co. support the OpenAI API format – existing applications can switch to your own inference system via an endpoint URL. Models are stored in S3, deployments are managed by Kubernetes, and Advanced Monitoring keeps an eye on latency and utilisation.
- OpenAI-compatible – Drop-in via endpoint change
- Constant latency – dedicated GPUs
- Model Registry – Weights versioned in S3
- Observable – Metrics & alerts included
The right centron products
Customers typically implement this use case using these building blocks – which can be combined and expanded at any time.
- RTX A4000 from €92.59 per month
- Quadro RTX 6000 from €170.83 · A100 from €489.47 · RTX 6000 Ada from €858.19 per month
- Dedicated, not shared
- No minimum term
- AutoScaler included
- Traffic at a fixed price
- CI/CD-ready
- Real-time metrics
- Customised alarms
- External checks
How much does LLM inference cost at centron?
LLM inference on your own GPU infrastructure: Run language models with high performance and in compliance with the GDPR – NVIDIA A100 & RTX 6000 Ada from €92.59 per month. The cornerstone of the service is Cloud GPU from €92.59 per month – billed by the hour, with no minimum contract term. This is supplemented, as required, by Kubernetes and Advanced Monitoring. All data remains in Germany: our own data centres in Hallstadt near Bamberg, certified to ISO 27001 on the basis of IT-Grundschutz and BSI C5:2020 Type 1. New accounts receive a €200 starting credit.
| Building block | Price |
|---|---|
| Cloud GPU | from €92.59 per month |
| Kubernetes | from €29.99 per month |
| Advanced Monitoring | from €4.99 per month |
Frequently asked questions
What is LLM inference?
Why is inference decisive for the practical use of AI?
What role do GPUs play in LLM inference?
When does running your own inference pay off compared with API providers?
What are the differences between GPU dedicated servers and cloud GPUs for LLM inference?
How much does it cost to get started?
What should companies look out for with cloud GPUs in Germany (data sovereignty, latency, oversubscription)?
Why should companies use centron’s cloud GPUs for LLM inference?
What advantages does cloud-based LLM inference offer over your own hardware?
In which industries is LLM inference used most?
Can this be implemented in a way that complies with the GDPR?
Can companies rent servers from centron?
More on AI & Machine Learning
Get started for free
Sign up and receive €200 credit at centron within your first 60 days.
This promotional offer applies to new accounts only. Available exclusively to businesses.