Chapter 16 · Always On Availability Groups, Failover Cluster Instances, and HA Design

Failover Cluster Instances: Shared Storage, Instance Failover, and Operational Tradeoffs

Understand SQL Server Failover Cluster Instances as instance-level HA with shared storage, virtual identity, dependencies and explicit operational tradeoffs.

Advanced165–205 minutesFCI architecture labSQL Server 2025 CU7 · 17.0.4065.4Compatibility 170 · Developer/Express simulationSSMS 22.8.2 · Last reviewed August 2026

Learning outcomes

A ServiceHub administrator sees “Always On” in documentation and assumes FCI and AG are interchangeable. They are not. A Failover Cluster Instance is one SQL Server instance whose service, network identity and storage dependencies are managed as a clustered role across WSFC nodes. Only one node owns/runs that instance at a time. The databases normally remain on shared storage; failover moves compute ownership to the same database files.

01

Describe FCI as instance-level HA with one active instance identity and cluster-managed dependencies.

02

Explain shared-storage implications, virtual network identity, system-database coverage and storage as a separate failure domain.

03

Distinguish FCI failover from AG replica role movement and from disaster recovery.

04

Inspect FCI evidence through SERVERPROPERTY/DMVs and reason about patch/failover workflows.

05

Apply SQL Server 2025 edition and Windows infrastructure constraints without making FCI a mandatory paid lab.

1. One instance identity, multiple possible owners

Clients connect to the FCI’s virtual network name rather than to a physical node. WSFC owns a group containing SQL Server resources and dependencies such as network name/IP and storage. If NODE-A fails and the cluster moves the group to NODE-B, the same SQL instance identity starts there and opens the same database files. This is different from an AG, where each replica is a separate SQL Server instance with its own database copy.

sql · record the local engine and HA feature context
SELECT SERVERPROPERTY('ProductVersion') AS product_version,       SERVERPROPERTY('ProductUpdateLevel') AS update_level,       SERVERPROPERTY('Edition') AS edition,       SERVERPROPERTY('EngineEdition') AS engine_edition,       SERVERPROPERTY('IsClustered') AS is_fci,       SERVERPROPERTY('IsHadrEnabled') AS hadr_enabled,       SERVERPROPERTY('MachineName') AS machine_name,       SERVERPROPERTY('ServerName') AS server_name;GOSELECT name, compatibility_level, recovery_model_descFROM sys.databasesWHERE name = N'ServiceHubLab';GO
Property FCI Availability Group
Protection unit SQL Server instance Selected user database(s)
Data copies Usually one shared storage copy One database copy per replica
System databases/jobs/logins Instance moves, so instance-scoped state moves with it Need separate synchronization/management unless using features such as contained AG where supported
Storage failure Shared storage can remain a common dependency Independent replica storage can survive a primary storage loss
Client endpoint FCI virtual network name AG listener for database role routing
Express Not supported Not supported
SQL Server 2025 Standard FCI up to two nodes Basic AG only
SQL Server 2025 Enterprise FCI up to 16 nodes Advanced AG feature set

2. Shared storage simplifies consistency but concentrates dependency

Because an FCI does not continuously copy database pages between nodes, there is no log-send/redo lag between FCI owners. But the data storage must remain accessible to the new owner. Shared SAN, Storage Spaces Direct or other supported cluster storage must be designed, validated and monitored independently. RAID or storage-array redundancy is not the same as an offsite database replica, and FCI does not replace backup.

sql · inspect cluster/instance evidence without changing the server
SELECT SERVERPROPERTY('IsClustered') AS is_failover_cluster_instance,       SERVERPROPERTY('ComputerNamePhysicalNetBIOS') AS current_physical_node,       SERVERPROPERTY('ServerName') AS sql_instance_network_name,       SERVERPROPERTY('Edition') AS edition;GOSELECT * FROM sys.dm_os_cluster_nodes;GOSELECT name,physical_name,state_desc,type_descFROM sys.master_filesWHERE database_id IN (1,2,3,4,DB_ID(N'ServiceHubLab'))ORDER BY database_id,file_id;GO

On the local standalone lab, IsClustered=0 and the cluster-node DMV may contain no meaningful cluster topology. That negative result is correct. Do not fabricate FCI output.

3. Dependency and ownership drive failover timing

An FCI failover includes failure detection, cluster ownership arbitration, resource transitions, SQL Server service startup on the target node, database crash recovery, listener/network convergence and client reconnection. A fast node transition does not guarantee instant database availability if recovery has substantial work to do. The transaction log and checkpoint behavior from Chapters 8 and 15 therefore matter directly to FCI RTO.

sql · build an FCI dependency review in the local design lab
USE ServiceHubLab;GOIF OBJECT_ID(N'lab16.FciDependency',N'U') IS NULLCREATE TABLE lab16.FciDependency(  dependency_name varchar(40) PRIMARY KEY,  owner_layer varchar(20) NOT NULL,  shared_across_nodes bit NOT NULL,  failure_effect nvarchar(200) NOT NULL);TRUNCATE TABLE lab16.FciDependency;INSERT lab16.FciDependency VALUES ('SQL Server service','WSFC',0,N'Must start on new owner node'), ('Virtual network name','WSFC',1,N'Client endpoint must come online/resolvable'), ('Database storage','Storage/WSFC',1,N'New owner must have valid storage access'), ('SQL Agent jobs','SQL instance',1,N'Move with the FCI instance'), ('External backup target','External',1,N'Must remain reachable from every possible owner');SELECT * FROM lab16.FciDependency ORDER BY dependency_name;GO

4. Patch by moving ownership deliberately, not by hoping

A rolling maintenance workflow normally validates cluster health, confirms a viable target owner, drains or coordinates application work, moves the clustered SQL role, verifies database/application health on the new owner, patches the passive node, then repeats according to the organization’s change design. SQL Server CU level and Windows cluster-node patch compatibility must follow Microsoft support guidance. A successful role move is not a substitute for an unplanned-failure drill.

Wrong approach: treat shared storage as “HA already solved.”

If the same storage array, fabric, credential, corruption event or administrative mistake can make the database files unavailable to every node, adding more FCI nodes does not remove that failure domain. Pair FCI with independent backups and, when required, database-level remote DR such as an AG or other supported technology.

powershell · Windows PowerShell — real FCI evidence/change commands to study, not run blindly
# Optional real WSFC/FCI topology only.Get-ClusterGroup | Sort-Object NameGet-ClusterResource | Sort-Object OwnerGroup,Name# Planned ownership changes require an approved runbook and application coordination.# Move-ClusterGroup -Name '<SQL Server FCI group>' -Node '<validated target>'

5. What FCI protects—and what stays outside the boundary

FCI gives strong instance continuity across a defined set of cluster nodes, but the protection boundary must be documented precisely. The SQL Server service, system databases, SQL Server Agent metadata, server-level objects stored in the instance, and databases on cluster-managed storage move as part of the instance. External dependencies do not magically move: Active Directory and DNS, Kerberos service principal names, external certificates or secrets, file shares, backup destinations, linked servers, remote endpoints, application configuration and monitoring systems all need to work from every possible owner node.

This is why a failover test should include more than SELECT @@SERVERNAME. Verify that the listener/virtual name resolves, TLS certificates validate for that name, the SQL Agent service starts, critical jobs are appropriately enabled, backup paths are reachable from the new node, database mail or alerting still works where used, and an application transaction succeeds. If an external dependency is node-local, failover can reveal a latent single-node configuration bug even though WSFC and SQL Server are functioning correctly.

sql · capture an FCI readiness checklist as data before a maintenance window
USE ServiceHubLab;GOIF OBJECT_ID(N'lab16.FciReadiness',N'U') IS NULLCREATE TABLE lab16.FciReadiness(  check_name varchar(60) PRIMARY KEY, expected_on_all_nodes bit NOT NULL,  verified bit NOT NULL, evidence nvarchar(300) NULL);TRUNCATE TABLE lab16.FciReadiness;INSERT lab16.FciReadiness VALUES ('Cluster validation current',1,0,N'Attach Test-Cluster evidence'), ('SQL service account rights',1,0,N'Verify service logon and permissions'), ('Backup destination reachable',1,0,N'Test from each possible owner'), ('TLS certificate/name valid',1,0,N'Validate virtual SQL name'), ('Monitoring/alerting follows owner',1,0,N'Generate a test signal');SELECT * FROM lab16.FciReadiness ORDER BY check_name;GO

Do not “fix” a failed readiness check by widening service-account rights or making backup shares world-writable. Preserve least privilege from Chapter 14 and prove access using the actual SQL Server/Agent service identities. HA that requires excessive privilege to survive failover creates a different class of production risk.

6. Production judgment

SQL Server 2025 FCI is available in Standard and Enterprise families, not Express; Standard supports up to two FCI nodes while Enterprise supports more. A real FCI also requires supported Windows Server failover-clustering infrastructure and validated storage/network design. SQL Server 2025 FCI setup has current TLS prerequisites; use current setup documentation rather than an old build sheet.

Choose FCI when instance-level failover and a single storage copy fit the failure model. Choose AG when independent database copies, remote DR/read scale or database-level role movement are required. Combining FCI and AG is possible, but it layers two different failover systems and therefore increases operational testing requirements.

Check your understanding

  1. Why does FCI normally avoid replica redo lag?
  2. What important failure domain can remain common across all FCI nodes?
  3. Why do SQL Agent jobs usually move naturally with FCI but not with a conventional AG failover?
  4. Why can a fast cluster role move still produce a longer database RTO?
  5. What SQL Server 2025 editions support FCI?
Review the answers

1. Because nodes normally mount the same database files rather than maintain independent continuously replayed copies.

2. The shared storage and its fabric/control plane.

3. FCI moves the SQL instance; an AG moves database roles between separate instances whose instance-scoped objects are independent.

4. The new owner still has to start SQL Server and recover databases, then clients must reconnect.

5. Standard and Enterprise families support FCI; Express does not.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.