Two sentinels
Losing one leaves no majority, so failover can never be authorised
- master-name
- mymaster
- sentinels
- 2
Generate sentinel.conf for a monitored primary. Two things decide whether failover works, and only one of them is the quorum: performing a failover also needs a majority of all sentinels to be reachable, which is why two sentinels can never fail over.
Worked setups you can load into the form above. Each one is a decision the generator makes differently, and the reason it makes it.
Losing one leaves no majority, so failover can never be authorised
Four buys nothing over three and makes a split vote possible
These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.
A majority of two is two, so losing either one makes failover impossible. Two sentinels detect an outage and can do nothing about it.
Instead:Run three, on three separate hosts.
The quorum only controls agreement that the primary is down. The majority requirement for performing the failover is separate and cannot be lowered.
Instead:Add sentinels rather than lowering the quorum.
After a promotion the client keeps writing to the old primary, which now rejects writes with READONLY.
Instead:Use the client library's Sentinel support.
Sentinel rewrites the file to record the current primary and its peers. Overwriting it discards that state.
Instead:Template it once, then exclude it.
Sentinel has two separate thresholds and configuring only the visible one is the usual reason a failover does not happen when it should.
Losing one leaves a single sentinel, which is not a majority of two, so no failover can be authorised. Three is the minimum, on three separate hosts, and an even number above that buys nothing over the odd number below it.
Sentinel does not proxy traffic; it tells clients where the primary is. A client configured with the primary's host and port keeps pointing at the old primary after a promotion, and the symptom is writes failing with READONLY.
Too low and an ordinary garbage collection pause triggers a failover nobody wanted. Too high and a real outage lasts longer than it needs to. The right value depends on how long your primary can legitimately stall.
It records the current primary and the other sentinels it has discovered. Deploying the same file to every host from configuration management will overwrite that state, so exclude it after the first start.