Chapter 16 · Always On Availability Groups, Failover Cluster Instances, and HA Design
Failover Cluster Instances: Shared Storage, Instance Failover, and Operational Tradeoffs
Understand SQL Server Failover Cluster Instances as instance-level HA with shared storage, virtual identity, dependencies and explicit operational tradeoffs.
Learning outcomes
A ServiceHub administrator sees “Always On” in documentation and assumes FCI and AG are interchangeable. They are not. A Failover Cluster Instance is one SQL Server instance whose service, network identity and storage dependencies are managed as a clustered role across WSFC nodes. Only one node owns/runs that instance at a time. The databases normally remain on shared storage; failover moves compute ownership to the same database files.
Describe FCI as instance-level HA with one active instance identity and cluster-managed dependencies.
Explain shared-storage implications, virtual network identity, system-database coverage and storage as a separate failure domain.
Distinguish FCI failover from AG replica role movement and from disaster recovery.
Inspect FCI evidence through SERVERPROPERTY/DMVs and reason about patch/failover workflows.
Apply SQL Server 2025 edition and Windows infrastructure constraints without making FCI a mandatory paid lab.
1. One instance identity, multiple possible owners
Clients connect to the FCI’s virtual network name rather than to a physical node. WSFC owns a group containing SQL Server resources and dependencies such as network name/IP and storage. If NODE-A fails and the cluster moves the group to NODE-B, the same SQL instance identity starts there and opens the same database files. This is different from an AG, where each replica is a separate SQL Server instance with its own database copy.
SELECT SERVERPROPERTY('ProductVersion') AS product_version, SERVERPROPERTY('ProductUpdateLevel') AS update_level, SERVERPROPERTY('Edition') AS edition, SERVERPROPERTY('EngineEdition') AS engine_edition, SERVERPROPERTY('IsClustered') AS is_fci, SERVERPROPERTY('IsHadrEnabled') AS hadr_enabled, SERVERPROPERTY('MachineName') AS machine_name, SERVERPROPERTY('ServerName') AS server_name;GOSELECT name, compatibility_level, recovery_model_descFROM sys.databasesWHERE name = N'ServiceHubLab';GO
| Property | FCI | Availability Group |
|---|---|---|
| Protection unit | SQL Server instance | Selected user database(s) |
| Data copies | Usually one shared storage copy | One database copy per replica |
| System databases/jobs/logins | Instance moves, so instance-scoped state moves with it | Need separate synchronization/management unless using features such as contained AG where supported |
| Storage failure | Shared storage can remain a common dependency | Independent replica storage can survive a primary storage loss |
| Client endpoint | FCI virtual network name | AG listener for database role routing |
| Express | Not supported | Not supported |
| SQL Server 2025 Standard | FCI up to two nodes | Basic AG only |
| SQL Server 2025 Enterprise | FCI up to 16 nodes | Advanced AG feature set |
2. Shared storage simplifies consistency but concentrates dependency
Because an FCI does not continuously copy database pages between nodes, there is no log-send/redo lag between FCI owners. But the data storage must remain accessible to the new owner. Shared SAN, Storage Spaces Direct or other supported cluster storage must be designed, validated and monitored independently. RAID or storage-array redundancy is not the same as an offsite database replica, and FCI does not replace backup.
SELECT SERVERPROPERTY('IsClustered') AS is_failover_cluster_instance, SERVERPROPERTY('ComputerNamePhysicalNetBIOS') AS current_physical_node, SERVERPROPERTY('ServerName') AS sql_instance_network_name, SERVERPROPERTY('Edition') AS edition;GOSELECT * FROM sys.dm_os_cluster_nodes;GOSELECT name,physical_name,state_desc,type_descFROM sys.master_filesWHERE database_id IN (1,2,3,4,DB_ID(N'ServiceHubLab'))ORDER BY database_id,file_id;GO
On the local standalone lab, IsClustered=0 and the
cluster-node DMV may contain no meaningful cluster topology.
That negative result is correct. Do not fabricate FCI output.
3. Dependency and ownership drive failover timing
An FCI failover includes failure detection, cluster ownership arbitration, resource transitions, SQL Server service startup on the target node, database crash recovery, listener/network convergence and client reconnection. A fast node transition does not guarantee instant database availability if recovery has substantial work to do. The transaction log and checkpoint behavior from Chapters 8 and 15 therefore matter directly to FCI RTO.
USE ServiceHubLab;GOIF OBJECT_ID(N'lab16.FciDependency',N'U') IS NULLCREATE TABLE lab16.FciDependency( dependency_name varchar(40) PRIMARY KEY, owner_layer varchar(20) NOT NULL, shared_across_nodes bit NOT NULL, failure_effect nvarchar(200) NOT NULL);TRUNCATE TABLE lab16.FciDependency;INSERT lab16.FciDependency VALUES ('SQL Server service','WSFC',0,N'Must start on new owner node'), ('Virtual network name','WSFC',1,N'Client endpoint must come online/resolvable'), ('Database storage','Storage/WSFC',1,N'New owner must have valid storage access'), ('SQL Agent jobs','SQL instance',1,N'Move with the FCI instance'), ('External backup target','External',1,N'Must remain reachable from every possible owner');SELECT * FROM lab16.FciDependency ORDER BY dependency_name;GO
4. Patch by moving ownership deliberately, not by hoping
A rolling maintenance workflow normally validates cluster health, confirms a viable target owner, drains or coordinates application work, moves the clustered SQL role, verifies database/application health on the new owner, patches the passive node, then repeats according to the organization’s change design. SQL Server CU level and Windows cluster-node patch compatibility must follow Microsoft support guidance. A successful role move is not a substitute for an unplanned-failure drill.
If the same storage array, fabric, credential, corruption event or administrative mistake can make the database files unavailable to every node, adding more FCI nodes does not remove that failure domain. Pair FCI with independent backups and, when required, database-level remote DR such as an AG or other supported technology.
# Optional real WSFC/FCI topology only.Get-ClusterGroup | Sort-Object NameGet-ClusterResource | Sort-Object OwnerGroup,Name# Planned ownership changes require an approved runbook and application coordination.# Move-ClusterGroup -Name '<SQL Server FCI group>' -Node '<validated target>'
5. What FCI protects—and what stays outside the boundary
FCI gives strong instance continuity across a defined set of cluster nodes, but the protection boundary must be documented precisely. The SQL Server service, system databases, SQL Server Agent metadata, server-level objects stored in the instance, and databases on cluster-managed storage move as part of the instance. External dependencies do not magically move: Active Directory and DNS, Kerberos service principal names, external certificates or secrets, file shares, backup destinations, linked servers, remote endpoints, application configuration and monitoring systems all need to work from every possible owner node.
This is why a failover test should include more than
SELECT @@SERVERNAME. Verify that the
listener/virtual name resolves, TLS certificates validate for
that name, the SQL Agent service starts, critical jobs are
appropriately enabled, backup paths are reachable from the new
node, database mail or alerting still works where used, and an
application transaction succeeds. If an external dependency is
node-local, failover can reveal a latent single-node
configuration bug even though WSFC and SQL Server are
functioning correctly.
USE ServiceHubLab;GOIF OBJECT_ID(N'lab16.FciReadiness',N'U') IS NULLCREATE TABLE lab16.FciReadiness( check_name varchar(60) PRIMARY KEY, expected_on_all_nodes bit NOT NULL, verified bit NOT NULL, evidence nvarchar(300) NULL);TRUNCATE TABLE lab16.FciReadiness;INSERT lab16.FciReadiness VALUES ('Cluster validation current',1,0,N'Attach Test-Cluster evidence'), ('SQL service account rights',1,0,N'Verify service logon and permissions'), ('Backup destination reachable',1,0,N'Test from each possible owner'), ('TLS certificate/name valid',1,0,N'Validate virtual SQL name'), ('Monitoring/alerting follows owner',1,0,N'Generate a test signal');SELECT * FROM lab16.FciReadiness ORDER BY check_name;GO
Do not “fix” a failed readiness check by widening service-account rights or making backup shares world-writable. Preserve least privilege from Chapter 14 and prove access using the actual SQL Server/Agent service identities. HA that requires excessive privilege to survive failover creates a different class of production risk.
6. Production judgment
SQL Server 2025 FCI is available in Standard and Enterprise families, not Express; Standard supports up to two FCI nodes while Enterprise supports more. A real FCI also requires supported Windows Server failover-clustering infrastructure and validated storage/network design. SQL Server 2025 FCI setup has current TLS prerequisites; use current setup documentation rather than an old build sheet.
Choose FCI when instance-level failover and a single storage copy fit the failure model. Choose AG when independent database copies, remote DR/read scale or database-level role movement are required. Combining FCI and AG is possible, but it layers two different failover systems and therefore increases operational testing requirements.
Check your understanding
- Why does FCI normally avoid replica redo lag?
- What important failure domain can remain common across all FCI nodes?
- Why do SQL Agent jobs usually move naturally with FCI but not with a conventional AG failover?
- Why can a fast cluster role move still produce a longer database RTO?
- What SQL Server 2025 editions support FCI?
Review the answers
1. Because nodes normally mount the same database files rather than maintain independent continuously replayed copies.
2. The shared storage and its fabric/control plane.
3. FCI moves the SQL instance; an AG moves database roles between separate instances whose instance-scoped objects are independent.
4. The new owner still has to start SQL Server and recover databases, then clients must reconnect.
5. Standard and Enterprise families support FCI; Express does not.
Authoritative references
- Always On failover cluster instances — FCI architecture and SQL Server 2025 notes
- Install a SQL Server failover cluster — setup and infrastructure requirements
- SQL Server 2025 editions and supported features — FCI edition limits
- FCI with Availability Groups — layering FCI and AG