Database Failover Interview Questions
Master Database Failover with interview-focused questions covering failover architecture, automatic failover, manual failover, switchover, leader election, quorum, split brain, heartbeat, replica promotion, high availability, and enterprise production best practices.
Introduction
Modern enterprise applications must remain available even when database servers fail.
Imagine
- Database Server Crash
- Network Failure
- Data Center Outage
- Hardware Failure
- Cloud Instance Failure
Without failover
Database Down
↓
Application Down
↓
Business Loss
Database Failover automatically transfers database responsibilities from a failed server to a healthy server, ensuring minimal downtime.
Failover is one of the core concepts behind
- High Availability (HA)
- Disaster Recovery (DR)
- Cloud Native Applications
- Distributed Databases
Database Failover Architecture
flowchart LR
Application --> LoadBalancer["Load Balancer"]
LoadBalancer["Load Balancer"] --> PrimaryDatabase["Primary Database"]
PrimaryDatabase["Primary Database"] --> ReplicaDatabase["Replica Database"]
PrimaryDatabase["Primary Database"]
-.Heartbeat.->
FailoverManager["Failover Manager"]
ReplicaDatabase["Replica Database"]
-.Heartbeat.->
FailoverManager["Failover Manager"]
1. What is Database Failover?
Answer
Database Failover is the process of automatically or manually switching database operations from a failed Primary database to a healthy Replica.
The goal is
- High Availability
- Minimal Downtime
- Continuous Service
2. Why is Failover required?
Without Failover
Primary Failure
↓
Application Stops
With Failover
Primary Failure
↓
Replica Promotion
↓
Application Continues
3. What is High Availability (HA)?
High Availability means
the database remains available
even if hardware or software failures occur.
Typical target
99.99%
or
99.999% availability
High Availability
flowchart TD
HighAvailability["High Availability"] --> Replication
HighAvailability["High Availability"] --> Failover
HighAvailability["High Availability"] --> Monitoring
4. What is Automatic Failover?
Automatic Failover means
the system automatically detects
Primary failure
and promotes
a Replica
without administrator intervention.
Automatic Failover Workflow
flowchart LR
PrimaryFailure["Primary Failure"] --> FailureDetectionReplicaPromotion["Failure Detection --> Replica Promotion --> ApplicationRedirected["Application Redirected"]"]
5. What is Manual Failover?
Manual Failover requires
an administrator
to initiate
Replica promotion.
Useful for
- Planned Maintenance
- Controlled Recovery
- Testing
6. What is Switchover?
Switchover is a planned role change
between
Primary
and
Replica
without any failure.
Usually performed during
- Maintenance
- Upgrades
- Hardware Replacement
Switchover Workflow
Primary
↓
Switchover
↓
Replica becomes Primary
↓
Old Primary becomes Replica
7. Failover vs Switchover
| Failover | Switchover |
|---|---|
| Unplanned | Planned |
| Failure Driven | Maintenance Driven |
| Emergency | Controlled |
| Possible Data Loss | No Data Loss (Ideally) |
8. What is Replica Promotion?
Replica Promotion means
a Replica
becomes
the new Primary
after failover.
Replica Promotion
flowchart LR
PrimaryDown["Primary Down"] --> ReplicaNewprimarynewPrimary["Replica --> NewPrimary["New Primary"]"]
9. What is Leader Election?
Leader Election is the process of selecting
one node
to become
the new Primary.
Commonly used in
- PostgreSQL
- Kafka
- ZooKeeper
- etcd
- Kubernetes
Leader Election
flowchart LR
Replica1["Replica 1"] --> Election
Replica2["Replica 2"] --> Election
Replica3["Replica 3"] --> Election
Election --> Leader
10. Why is Leader Election required?
Without Leader Election
multiple databases
may believe
they are Primary.
This creates
data inconsistency.
11. What is Split Brain?
Split Brain occurs when
multiple servers
believe
they are
the Primary
at the same time.
This results in
conflicting writes.
Split Brain
Network Partition
↓
Primary A
Primary B
↓
Both Accept Writes
↓
Conflict
12. Why is Split Brain dangerous?
Problems include
- Data Corruption
- Lost Updates
- Duplicate Records
- Inconsistent Data
13. How can Split Brain be prevented?
Common solutions
- Quorum
- Fencing Tokens
- Consensus Algorithms
- Leader Election
- Witness Node
14. What is Quorum?
Quorum is
the minimum number of nodes
required
to make cluster decisions.
Example
5 Nodes
↓
Majority = 3
Quorum Example
flowchart LR
Node1 --> Quorum
Node2 --> Quorum
Node3 --> Quorum
Node4 --> Quorum
Node5 --> Quorum
15. What is Heartbeat?
Heartbeat is
a periodic health check
between
cluster nodes.
If heartbeat stops,
the node is suspected
to have failed.
Heartbeat
flowchart LR
Primary
-.Heartbeat.->
Replica
Replica
-.Heartbeat.->
Primary
16. What is Failure Detection?
Failure Detection monitors
- Heartbeats
- Network Connectivity
- Node Health
If the Primary becomes unreachable,
failover begins.
17. What is Failback?
Failback means
restoring
the original Primary
after it has recovered.
Typical steps
Old Primary Recovered
↓
Synchronize Data
↓
Optional Switchover Back
18. What happens during Failover?
Typical sequence
Primary Failure
↓
Failure Detection
↓
Leader Election
↓
Replica Promotion
↓
Application Redirect
↓
Normal Operation
Complete Failover Flow
flowchart LR
PrimaryFailure["Primary Failure"] --> HeartbeatLostLeaderElection["Heartbeat Lost --> Leader Election --> Replica Promotion --> Client Redirect --> ApplicationRunning["Application Running"]"]
19. What causes Failover?
Common causes
- Hardware Failure
- Power Failure
- Network Partition
- Database Crash
- Storage Failure
- Cloud VM Failure
20. What is Recovery Time Objective (RTO)?
RTO defines
the maximum acceptable downtime
after a failure.
Example
Target
< 5 Minutes
21. What is Recovery Point Objective (RPO)?
RPO defines
how much data
the business is willing to lose.
Example
Maximum
30 Seconds
RTO vs RPO
| RTO | RPO |
|---|---|
| Downtime | Data Loss |
| Time to Recover | Point to Recover |
| Availability | Durability |
22. Banking Example
Primary Database
↓
Failure
↓
Replica Promoted
↓
ATM Transactions Continue
23. E-Commerce Example
Order Database
↓
Primary Crash
↓
Automatic Failover
↓
Customers Continue Ordering
24. Airline Booking Example
Booking Server
↓
Replica Promotion
↓
No Booking Downtime
25. SaaS Example
Cloud Region Failure
↓
Replica Region Activated
↓
Application Available
26. Production Example
100 Million Users
↓
Primary Database Crash
↓
30 Seconds
↓
Replica Promoted
↓
Traffic Restored
27. Common Failover Challenges
- Split Brain
- Slow Detection
- Replication Lag
- Long Recovery Time
- Network Partition
- DNS Delay
28. Common Failover Tools
Examples
- PostgreSQL Patroni
- MySQL Group Replication
- Oracle Data Guard
- SQL Server Always On
- Redis Sentinel
- Kubernetes Operators
29. Best Practices for Failover
- Automate failover where possible.
- Continuously monitor node health.
- Test failover regularly.
- Minimize replication lag.
- Deploy replicas in different Availability Zones.
- Prevent split brain using quorum.
- Document recovery procedures.
- Keep backups independent of failover.
- Monitor RTO and RPO.
- Validate applications after failover.
Database Failover Workflow
flowchart LR
Primary
-.Heartbeat.->
FailoverManager["Failover Manager"]
FailoverManager["Failover Manager"] --> Replica
Replica --> NewPrimary["New Primary"]
NewPrimary["New Primary"] --> Application
Enterprise Best Practices
- Deploy at least one replica for every production database.
- Use automatic failover for critical systems.
- Monitor heartbeat continuously.
- Test disaster recovery every quarter.
- Minimize replication lag.
- Keep replicas in different regions or availability zones.
- Use quorum-based leader election.
- Avoid split-brain scenarios.
- Monitor failover duration.
- Practice regular failover drills.
Quick Revision
| Topic | Key Point |
|---|---|
| Failover | Automatic Recovery |
| Switchover | Planned Role Change |
| Replica Promotion | Replica Becomes Primary |
| Leader Election | Select New Primary |
| Split Brain | Multiple Primaries |
| Quorum | Majority Decision |
| Heartbeat | Health Check |
| Failback | Restore Original Primary |
| RTO | Recovery Time |
| RPO | Data Loss Objective |
Interview Tips
Interviewers frequently ask
- What is Database Failover?
- Failover vs Switchover.
- What is Replica Promotion?
- Explain Leader Election.
- What is Split Brain?
- How is Split Brain prevented?
- What is Quorum?
- What is Heartbeat?
- Explain RTO vs RPO.
- Describe the failover process in production.
A strong interview explanation is:
"Database failover is the automatic or manual process of promoting a healthy replica to become the new primary when the original primary fails. High availability solutions continuously monitor heartbeats, perform leader election, promote a replica, and redirect application traffic. To avoid data corruption, enterprise systems use quorum-based decisions and split-brain prevention mechanisms. Successful failover minimizes downtime (RTO) and data loss (RPO), ensuring business continuity."
Summary
Database Failover is a critical capability for building highly available, fault-tolerant, and resilient enterprise systems. Through automatic failover, leader election, heartbeat monitoring, replica promotion, and quorum-based decision making, organizations can maintain service availability even during hardware failures, software crashes, or network outages.
Understanding failover concepts such as Switchover, Failback, Split Brain, RTO, RPO, and production best practices is essential for Backend Developers, Database Engineers, DevOps Engineers, Cloud Engineers, and Solution Architects responsible for designing mission-critical distributed systems.