Windows Failover Cluster: Quorum Loss, Witness Failures and Nodes That Will Not Join
Most failover cluster incidents are quorum incidents wearing a different name. This guide explains how votes are counted at the moment of failure, how to read the cluster log, and how to recover a cluster that has shut itself down to protect data.
What quorum actually counts
A failover cluster stays online while it holds a majority of votes. Each node normally has one vote, and a witness may hold one more. If a majority cannot be assembled, the remaining nodes deliberately stop — this is correct behaviour that prevents a split brain writing to shared storage from two directions.
Get-ClusterQuorum Get-ClusterNode | Select-Object Name, State, NodeWeight, DynamicWeight Get-Cluster | Select-Object Name, DynamicQuorum, WitnessDynamicWeight
Two columns matter and are easily confused:
- NodeWeight — the configured vote, 0 or 1. Set by the administrator.
- DynamicWeight — the vote the cluster is currently granting. Adjusted automatically by dynamic quorum.
Since Windows Server 2012 R2, dynamic quorum removes votes from nodes as they leave, so a cluster can survive sequential failures that static arithmetic says should have killed it. It cannot survive simultaneous loss of a majority, which is why a single network event that isolates several nodes at once is far more dangerous than the same nodes failing one at a time.
Dynamic witness complements this: the witness vote is granted only when the node count is even, and withdrawn when it is odd, keeping the total odd at all times.
Get-ClusterResource -Name "File Share Witness" | Get-ClusterParameter (Get-Cluster).WitnessDynamicWeight # 1 = witness currently voting, 0 = not
A witness reporting WitnessDynamicWeight = 0 on an even-node cluster is a real problem and usually means the witness resource is failed or unreachable.
Event IDs that tell you what happened
| Event ID | Source | Meaning |
|---|---|---|
| 1135 | FailoverClustering | A node was removed from active membership. Network or heartbeat loss, not necessarily a crash. |
| 1177 | FailoverClustering | The node stopped because quorum was lost. The cluster service shut down on purpose. |
| 1069 | FailoverClustering | A clustered role or resource failed. |
| 1146 | FailoverClustering | The resource hosting subsystem (RHS) crashed — usually a misbehaving resource DLL. |
| 1230 | FailoverClustering | A resource deadlocked and was terminated. |
| 5120 | FailoverClustering | Cluster Shared Volume paused — storage path loss. Common and important. |
| 5142 | FailoverClustering | CSV no longer accessible from this node. |
Get-WinEvent -FilterHashtable @{LogName='System'; ProviderName='Microsoft-Windows-FailoverClustering'} -MaxEvents 100 |
Where-Object Id -in 1135,1177,1069,1146,5120 |
Format-Table TimeCreated, Id, Message -Wrap
The sequence matters more than any single event. A 1135 followed seconds later by 1177 on the surviving nodes describes a network partition. A 1177 with no preceding 1135 points at storage or at the witness.
Generate and read the cluster log
The cluster log is far more detailed than the event log and is not written as a text file until you ask for it. Timestamps are UTC by default, so convert them or you will misalign events by hours:
# last 24 hours, local time, to C:\Logs Get-ClusterLog -Destination C:\Logs -TimeSpan 1440 -UseLocalTime
Search for the moment of the decision rather than reading linearly:
Select-String -Path C:\Logs\*_cluster.log -Pattern 'quorum|lost quorum|NetftIsConnected|IsolatedNode|Shutting down' | Select-Object -First 60
Useful markers inside the log:
[QUORUM]— vote arithmetic at the time of the event.[NETFTAPI]— the cluster virtual adapter; heartbeat loss appears here first.[RCM]— Resource Control Manager, showing resource state changes and failover decisions.[DCM]— Distributed Cluster Manager, for CSV events.
Heartbeat and network isolation
Clustering health checks are sensitive by design: the defaults assume a low-latency LAN. On stretched clusters, oversubscribed virtual switches or networks with aggressive QoS, a node can be evicted for a delay that no user would ever notice.
Get-Cluster | Select-Object SameSubnetDelay, SameSubnetThreshold,
CrossSubnetDelay, CrossSubnetThreshold,
RouteHistoryLength
| Setting | Default | Meaning |
|---|---|---|
| SameSubnetDelay | 1000 ms | Interval between heartbeats |
| SameSubnetThreshold | 5 (10 on 2019+) | Missed heartbeats before the node is considered down |
| CrossSubnetDelay | 1000 ms | Interval for nodes on different subnets |
| CrossSubnetThreshold | 5 (20 on 2019+) | Missed heartbeats across subnets |
Raising thresholds on a stretched cluster is legitimate; raising them to mask a genuine network fault is not — it extends the time before a real failure is detected.
(Get-Cluster).SameSubnetThreshold = 10 (Get-Cluster).CrossSubnetThreshold = 20
Confirm the cluster networks are classified correctly. A storage or backup network accidentally marked as cluster-capable will carry heartbeats and introduce jitter:
Get-ClusterNetwork | Select-Object Name, Role, Address, State # Role: 0 = none, 1 = cluster only, 3 = cluster and client
Witness problems
The witness is the most commonly neglected component in a cluster and the one that decides the outcome of a tie.
Get-ClusterQuorum | Select-Object Cluster, QuorumResource, QuorumType
- File share witness — must be on a server outside the cluster, with the cluster name object granted both share and NTFS write permission. Granting rights to the node computer accounts rather than the CNO is a frequent misconfiguration.
- Disk witness — small shared LUN; fails with the storage path, so it is a poor choice when storage is the likely failure domain.
- Cloud witness — an Azure storage account; requires outbound HTTPS to
*.core.windows.netfrom every node, which proxies and firewalls often block.
# reconfigure witness types Set-ClusterQuorum -NodeAndFileShareMajority "\\fs01\ClusterWitness$" Set-ClusterQuorum -NodeAndCloudMajority -AccountName "stwitness01" -AccessKey "<key>" Set-ClusterQuorum -NodeMajority # odd node count, no witness
Microsoft's current guidance is to configure a witness in all cases, including odd-node clusters, because dynamic quorum can reduce an odd cluster to an even one during a failure.
Recovering a cluster that will not start
When too many nodes are lost, the cluster service will not start — correctly, because it cannot prove it is the authoritative copy. Forcing it is a deliberate data-risk decision and should be taken on exactly one node.
# start this node as authoritative, ignoring quorum Start-ClusterNode -Name NODE1 -FixQuorum
Immediately afterwards, correct the vote configuration before starting anything else, or the next transient fault repeats the outage:
(Get-ClusterNode NODE2).NodeWeight = 0 (Get-ClusterNode NODE3).NodeWeight = 0 Get-ClusterNode | Select-Object Name, NodeWeight, DynamicWeight
Bring the remaining nodes back one at a time, restoring their weight as each rejoins. Never force quorum on two nodes concurrently — that is precisely the split brain the shutdown was protecting you from.
Validate before declaring the incident closed. The validation report is also what Microsoft support will ask for first:
Test-Cluster -Node NODE1,NODE2,NODE3 -ReportName C:\Logs\ClusterValidation Test-Cluster -Include "Storage","Network","Inventory"
Run storage tests only when the clustered roles are offline — they take disks offline as part of the test and will cause an outage if run against a live cluster.
Checklist
Get-ClusterNodefirst: compare NodeWeight with DynamicWeight.- Establish whether the witness is voting —
WitnessDynamicWeight. - Build a timeline from Event IDs 1135, 1177, 5120 across all nodes, not just the one that complained.
- 1135 before 1177 means network. 1177 alone points at witness or storage.
- Generate the cluster log with
-UseLocalTimeand search[QUORUM]and[NETFTAPI]. - Check cluster network roles — a storage network carrying heartbeats causes phantom evictions.
- Only force quorum on one node, then immediately fix node weights.
- Finish with
Test-Clusterand keep the report.
Get-ClusterNode votes, not with the application.