Monitor Automatic Page Repair in a Distributed Availability Group

A distributed availability group connects two separate underlying availability groups. Automatic page repair operates through an availability group’s own primary and secondary replicas; the distributed layer should not be treated as one large pool in which any remote replica can directly supply a damaged page.

Last updated: October 11, 2026.

SELECT
    DB_NAME(database_id) AS DatabaseName,
    file_id,
    page_id,
    error_type,
    page_status,
    modification_time
FROM sys.dm_hadr_auto_page_repair
ORDER BY modification_time DESC;
GO

Run this query on every instance hosting a replica in both underlying availability groups. SQL Server 2022 and later requires VIEW SERVER PERFORMANCE STATE; earlier supported versions use VIEW SERVER STATE. Preserve results because the DMV retains only a limited number of attempts per database. Status 4 or 5 indicates success or failure according to the role handling the request, so interpret it with the replica role and error log.

Apply the local availability-group repair model

Microsoft’s automatic page repair documentation says a primary broadcasts an eligible page request to the secondaries in its availability group and uses the first valid response. A secondary with a redo-time page error requests the page from its primary. Repairs cover eligible data-page errors such as 823, 824, and 829, not every control or allocation page.

Microsoft defines a distributed availability group as two separate availability groups linked through their primaries. Therefore, treating repair partners as members of the damaged replica’s underlying local group is the conservative interpretation of the documented architecture.

Repair success does not close the incident

A successful row in sys.dm_hadr_auto_page_repair means SQL Server replaced the unreadable page; it does not prove that storage is healthy. Correlate the event with error 823 or 824 entries, msdb.dbo.suspect_pages, operating-system logs, storage telemetry, and recent integrity checks. Run an appropriate DBCC CHECKDB plan and investigate the hardware path.

Alert on new repair attempts across every replica instead of checking only the global primary. A secondary may encounter corruption while replaying log records, suspend its database, and record the evidence locally before an administrator notices application symptoms.

If no healthy local repair partner exists, use the established restore and corruption response process. Continue with availability-group timeout behavior, backup verification, and restoring a backup elsewhere.

Related Web Cheat Sheet guides

Sergey Kornilov

Sergey Kornilov