High Availability – Availability as a principle · disaster recovery for SMEs

High availability, thought through

From redundant data centres to clusters and replication to failover in seconds: this is how centron keeps your services running. Most businesses have a backup – far fewer can say how long it would take to be operational again. Anyone who has to rebuild servers, network, access and applications by hand after a total failure needs days, even if all the data is there.

ccloud³ Console · HA cluster
eu-de · Hallstadt Data Centre
Pool
High Availability
SLA
99.8%
Failover
seconds
LB Load balancer · 2 nodes
● Active
REP cProtect · fire compartment B
● 15 min
BAK cBacks · second location
● Backed up
TST Restart test · Q3
● Documented
Node failedfailover · users notice nothing
  • Four levels, one decision – from a simple backup to a cluster that never makes the outage visible in the first place.
  • Distributed operation since 2016 – backups and systems across several locations in Germany, by default in Hallstadt near Bamberg.
  • A tested restart – test environments for recovery, billed by the minute.
  • Certified to the BSI standard – ISO 27001 based on IT-Grundschutz and BSI C5 attestation.
The HA modules

Availability is built up in layers

True high availability is not a single product, but rather the interplay between infrastructure, architecture and processes.

Redundant infrastructure

UPS, emergency power systems, separate fire compartments and multiple redundant power connections – the physical infrastructure at our data centre in Hallstadt near Bamberg.

Highly available architecture

Managed clusters with load balancing, real-time replication and automatic failover keep applications online – even if individual components fail.

Contingency planning

cProtect replicates data every 15 minutes to a separate fire compartment within the Hallstadt data centre – with immutable data states and failover within seconds, even in the event of a ransomware attack.

Vigilant processes

24/7 monitoring with alerts and – in the case of managed services – proactive fault resolution by centron experts before users even notice anything is wrong.

Use cases

Four levels of resilience

The difference between backup and restart only becomes visible in an emergency – and then it is too late to clarify it. The permissible downtime and permissible data loss per system determine which of the four levels is required.

1. Backup

The data is there, the restart is done by hand. Downtime: hours to days. Suitable for systems whose failure is bearable, such as test environments or archives. Implemented via cBacks.

2. Backup at a second location

The same backup, additionally at another location. Protects against events that affect the entire site: fire, flooding, prolonged power failure. Implemented via managed backup or S3 Object Storage.

3. Failover

A second system stands ready and takes over in the event of failure. Downtime: minutes. The switch-over can happen automatically or on instruction. Suitable for systems whose standstill noticeably disrupts operations – implemented with cProtect.

4. High availability

Several systems run in parallel; the failure of a single one goes unnoticed by users. Downtime: close to zero. The most elaborate and expensive level, sensible for systems without which no work is possible – as a managed cluster in the High Availability Pool.

From Need to Architecture

How much uptime does your system require?

Not every application requires geographical redundancy – but every one requires a well-thought-out plan. During the high-availability consultation, we analyse the costs of downtime, define RPO and RTO targets, and design the appropriate architecture: from a pair of load balancers to cross-site replication.

99.8 %High Availability Pool · guaranteed per month in the service specification
Find out more about cProtect
  • Needs analysis – Clarify downtime costs and objectives
  • Definition of RPO/RTO as a basis for architecture
  • A tailor-made solution From clusters to geographical redundancy
  • Contractual SLAs to give you certainty in your planning
Disaster recovery plan

What belongs in a plan – and who is responsible for what

Test recovery without spare hardware: you restore a backup into a new ccloud³ instance, check that the application starts and the data is correct, measure the time actually needed and delete the instance again afterwards – only the time it exists is billed. The measured time almost always deviates from the planned time, usually upwards; that is precisely the value of the test. What is documented: date, restored system, time needed, problems encountered and the resulting changes to the plan – exactly what auditors and insurers want to see. The evidence is provided by the Trust Center.

What belongs in a disaster recovery plan
ItemContent
Prioritisation of systemsAn order by importance for business operations. Without it, the restart follows technical effort rather than need.
Recovery objectivesFor each system, the permissible downtime and permissible data loss. Both values determine which of the four levels is required and at what interval backups are made.
ResponsibilitiesNamed persons responsible with deputies, plus availability outside working hours, including that of the service provider.
Restart procedureThe individual steps, documented in a form that can also be carried out by people who do not look after the system daily.
Test cycleA fixed schedule for rehearsal. A plan without a documented recovery is no evidence in an audit; at least one recovery per year per critical system is customary.
Who is responsible for what
TaskYoucentron
Classification of systems and recovery objectivesentirelyadvice on technical feasibility
Disaster recovery plancreation and maintenanceprovision of operational details
Applications and operating systemsconfiguration, application restartmanaged service on request
Backup and recoverydefinition of interval and retentionexecution and provision
Platform and data centresredundancy, monitoring, emergency management of operations
Recovery testplanning, execution, documentationprovision of the test environment
Evidence for auditors and insurersinclusion in your own documentationcertificates with scope, BSI C5 attestation and TOM in the Trust Center
Recommended modules

The right centron products

High availability emerges from the interplay of these modules – combinable and extensible at any time. We work out the right configuration together with you.

cProtect
From
€0.03 / month
Recovery in seconds
  • Replication every 15 minutes
  • Failover in seconds
  • Separate fire compartment
Managed Backup
Price
on request
Supported backup
  • Design and operation by centron
  • Retention as defined
  • Restores tested regularly
Managed Cluster
Price
on request
Customised SLAs
  • Load balancing and real-time replication
  • Automatic failover
  • Operation and updates by centron
Contractual undertakings

Availability, in black and white

These values are laid down in the service specifications and apply per calendar month. cProtect has no availability commitment of its own; the service follows the SLA of the infrastructure on which the failover system runs. For a highly available system, the High Availability Pool is therefore the foundation.

SLAs & key figures
ServiceValueScope
ccloud³ VMs in the High Availability Pool99.8 %per month · ccloud³ service specification
ccloud³ VMs in the Worker and Power Pool99.5 %per month · ccloud³ service specification
Managed Server99.5 % / 99.8 %Worker/Power Pool or HA Pool
Volumes and S3 Object Storage99.9 %per month · service specification
Colocation network99.9 %annual average · Colocation service specification
cProtect replication intervalapprox. 15 minutescProtect service specification
Customised SLAsby arrangementManaged Cluster and Full Managing
High Availability FAQ

Frequently asked questions on high availability and disaster recovery

What do 99.5%, 99.8% and 99.9% availability mean in concrete terms?

The SLAs apply per month. With around 730 hours in a month, 99.5% corresponds to up to about 3.7 hours of unavailability, 99.8% to up to about 1.5 hours and 99.9% to up to about 44 minutes. For highly available systems, the High Availability Pool with 99.8% is therefore the right basis. Volumes, object storage and the colocation network are at 99.9%.

Isn’t a backup enough?

Backups protect data, not availability. Restoring large systems takes hours. High availability starts before that, with redundant systems and failover, so that operations never come to a standstill in the first place. Backups remain mandatory nonetheless.

What belongs in a disaster recovery plan?

An order of systems by importance, for each system the permissible downtime and permissible data loss, named responsibilities with deputies and availability, the restart steps in executable form and a fixed schedule for rehearsal.

What is disaster recovery as a service?

Disaster recovery as a service is a model in which the fallback environment is not kept permanently but is provided by the provider and activated when needed. The advantage is that no second set of hardware has to be bought and operated; billing is by usage.

What is the difference between high availability and failover?

With failover, a second system standing by takes over when the first fails; the switch-over takes seconds to minutes and is usually noticeable. With high availability, several systems run in parallel so that the failure of a single one remains invisible to users. High availability is more elaborate and more expensive.

How do I achieve high availability for my application?

Typically with a managed cluster: several nodes in the High Availability Pool, load balancing, replication and automatic failover. centron plans and operates this and backs it with individual SLAs.

How much resilience does a medium-sized company need?

For small and medium-sized businesses that cannot be answered across the board, only per system. It makes sense to assess each system by what a day of standstill would cost. Systems without which no work is possible justify failover or high availability; for everything else, a backup with a second location is usually sufficient.

What happens if the production system fails?

With cProtect, a state of your system no more than around 15 minutes old is ready in a separate fire compartment of the Hallstadt data centre. In the event of a fault, the failover system is activated and takes over operations. cProtect has no availability SLA of its own; the SLA of the infrastructure on which the failover system runs applies.

How does cProtect work?

cProtect replicates the VMs in its own environment and creates a consistent backup of the VM every 15 minutes. 15 + 1 states are generated, i.e. the 16th state overwrites the first. In the event of a failover, the replication VM is used so that operation of the application running on the server is ensured again within a very short time – details on cProtect.

How far back in time can I go with cProtect?

By default, cProtect creates a backup every 15 minutes. This lets you jump back up to four hours in time. On request, up to 24 hours are also possible. This option is particularly effective against ransomware, crypto-trojans and other malware.

Does a backup protect against ransomware?

Only if the backup itself is not reachable. Attackers look for the backups first and encrypt or delete them. Effective backups are those that cannot be changed within the retention period. How this is implemented is described by cProtect.

Does cProtect protect me against cyber attacks?

cProtect is an effective way to undo infections by ransomware, crypto-trojans and other malware – simply turn back time. The prerequisite is that the infection is noticed promptly.

How can I protect myself even better against cyber attacks?

cProtect reaches its limits if the infection is not noticed quickly enough. If you want to be on the safe side here, we recommend our managed backup: based on a previously defined backup plan, we regularly create backups of your systems and retain them for up to four weeks.

What alternative to cProtect is there?

As an alternative to cProtect we offer two solutions: managed backup – with our backup landscape we create data backups based on a predefined backup plan – and the failover cluster: on request we build a dedicated or virtual failover cluster for you.

What does a geo-backup offer?

A geo-backup stores the backup at another location. It therefore remains available even if an event completely hits the primary site, such as fire, flooding or a prolonged power failure. centron has run distributed operations across several locations in Germany since 2016.

How often should a recovery be tested?

At least once a year per critical system, plus after every major change to the application or environment. What matters is less the frequency than the documentation: a test without a recorded result helps neither with improvement nor with auditors.

Where are our backups located?

By default in the centron data centre in Hallstadt near Bamberg. The scope of our ISO 27001 certification based on IT-Grundschutz covers Hallstadt, Nuremberg and Frankfurt am Main. Further locations can be included for distributed operations.

Get started for free

Sign up and receive €200 credit at centron within your first 60 days.

This promotional offer applies to new accounts only. Available exclusively to businesses.