AZ-802 monitoring and troubleshooting

Updated September 29, 2026

Monitor and troubleshoot Windows Server environments is worth 15–20% of AZ-802. It covers three areas: monitoring with Windows Server tools and Azure services, troubleshooting Windows Server itself, and troubleshooting Active Directory. It is the domain that ties the others together, because a troubleshooting question can be about DNS, Hyper-V, BitLocker or Kerberos. What it tests is method: which tool shows the problem, and which fix addresses the cause rather than the symptom.

Monitoring

Windows Server tools

ToolUse
Performance MonitorLive counters: processor, memory, disk, network
Data collector setsScheduled or triggered collection of counters, traces and configuration, with reports
Windows Admin CenterDashboards and alerts across servers
System InsightsLocal machine learning that forecasts CPU, storage and network capacity
Event ViewerEvent logs, custom views and subscriptions

Event subscriptions forward events from many servers to a collector. In collector-initiated mode the collector pulls from a listed set of sources; in source-initiated mode sources push to the collector, configured through Group Policy, which scales better. Both rely on WinRM and the Windows Event Collector service.

Useful counters and thresholds to recognise: sustained % Processor Time above about 85%, Available MBytes near zero with high Pages/sec, and Avg. Disk sec/Read well above 20 ms.

Azure services

  • The Azure Monitor Agent collects data from Azure VMs and Arc-enabled servers. What it collects and where it sends it is defined in data collection rules (DCRs), which you associate with machines. One DCR can cover many servers.
  • Alerts fire on metrics, log queries or the activity log, and notify through action groups.
  • VM insights shows performance charts and, optionally, a dependency map of processes and connections. It uses the Azure Monitor Agent and a DCR; the map also needs the Dependency agent.

Troubleshooting Windows Server

ProblemFirst tools
ConnectivityTest-NetConnection, ping, tracert, firewall rules
Name resolutionResolve-DnsName, nslookup, ipconfig /flushdns, Clear-DnsClientCache
Windows UpdateGet-WindowsUpdateLog, Update Manager history, WSUS or proxy settings
Time servicew32tm /query /status, w32tm /resync, w32tm /config
PerformancePerformance Monitor, Resource Monitor, data collector sets
VM and Arc extensionsExtension status in the portal, azcmagent check, extension logs on the machine
Disk encryptionBitLocker status, recovery key location, manage-bde -status
StorageDisk and volume health, Storage Spaces health, event logs

Time is worth extra attention because Kerberos allows a clock skew of only five minutes. In a domain, members sync from DCs, DCs from the PDC emulator of their domain, and the PDC emulator of the forest root domain from a reliable external source. Configure that one with w32tm /config /manualpeerlist:... /syncfromflags:manual /reliable:yes. If the PDC emulator role moves, reconfigure the new holder.

Troubleshooting Active Directory

Recovering deleted objects

SituationMethod
AD Recycle Bin enabledRestore-ADObject or AD Administrative Center; attributes and memberships intact
No Recycle Bin, recent backupAuthoritative restore of the object from DSRM with ntdsutil
No Recycle Bin, no backupTombstone reanimation; most attributes lost

The Recycle Bin is enabled with Enable-ADOptionalFeature and cannot be turned off again. It only helps with objects deleted after it was enabled.

DSRM and the database

Directory Services Restore Mode starts a DC with AD DS offline, using the DSRM administrator password set at promotion. From there you restore from backup, run an authoritative restore with ntdsutil to mark objects as newer than their replicated copies, or check and compact the database with ntdsutil.

SYSVOL

SYSVOL replicates with DFS Replication. If SYSVOL is damaged on one DC, a non-authoritative sync makes that DC take SYSVOL from its partners. If it is broken everywhere, an authoritative sync on one DC, with the others set to non-authoritative, rebuilds it from a single good copy. Both are set through the msDFSR-Enabled and msDFSR-Options attributes on the DC’s SYSVOL subscription object.

Replication

repadmin /replsummary gives an overview of failures, repadmin /showrepl shows per-partner status, and dcdiag tests a DC broadly. Common causes are DNS problems, blocked ports between sites, time skew and lingering objects after a DC was offline longer than the tombstone lifetime.

Kerberos and secure channels

  • Kerberos: check time skew, duplicate or missing SPNs (setspn -X, setspn -L), and tickets with klist. Errors such as KRB_AP_ERR_MODIFIED often point at an SPN on the wrong account.
  • Secure channel: when a computer reports that the trust relationship with the domain failed, test with Test-ComputerSecureChannel or nltest /sc_verify and repair with Test-ComputerSecureChannel -Repair or Reset-ComputerMachinePassword. Reverting a VM checkpoint to before a machine password change is a classic cause.

Sample questions

Question 1. Users intermittently fail to sign in, and the security log on a DC shows Kerberos errors about clock skew. The forest root PDC emulator role was moved to a new DC last week. What should you do?

  • A. Enable Hyper-V time synchronisation on every VM
  • B. Reset the computer account passwords of the affected servers
  • C. Configure the new PDC emulator to sync from an external time source with w32tm
  • D. Increase the Kerberos maximum clock skew to 30 minutes
Show answer

Answer: C

Only the forest root PDC emulator should sync with an external time source. When the role moved, the new holder was not configured, so the domain hierarchy drifts. Configuring the new PDC emulator with w32tm /config and a manual peer list fixes the source. Syncing members with the Hyper-V integration service or disabling Kerberos is wrong, and resetting computer passwords addresses secure channels, not time.

Want more questions like this? Full AZ-802 practice tests →

Question 2. After a VM was reverted to a checkpoint from two months ago, users cannot sign in on it with the message that the trust relationship between this workstation and the primary domain failed. What is the quickest fix that does not require removing the server from the domain?

  • A. Restart the Netlogon service on the server
  • B. Flush the DNS cache on the server
  • C. Perform an authoritative restore of the computer object
  • D. Run Test-ComputerSecureChannel -Repair
Show answer

Answer: D

Reverting the checkpoint restored an old machine account password, so the secure channel breaks. Test-ComputerSecureChannel -Repair, run with domain credentials, resets the password and restores the channel without a rejoin. Restarting Netlogon, a DNS flush or an authoritative restore of the computer object do not fix the mismatched password.

Want more questions like this? Full AZ-802 practice tests →

Question 3. You must collect the Security event log from 150 servers to one server, and new servers should start forwarding automatically when they join the Servers OU. What should you configure?

  • A. A source-initiated event subscription with a GPO on the Servers OU
  • B. A collector-initiated event subscription listing all 150 servers
  • C. A data collector set on each server
  • D. System Insights on each server
Show answer

Answer: A

A source-initiated subscription, with the target subscription manager set through a GPO linked to the Servers OU, makes every server in the OU forward events to the collector automatically. A collector-initiated subscription requires listing each source. Data collector sets gather performance data, and System Insights forecasts capacity.

Want more questions like this? Full AZ-802 practice tests →

What to practise

Break things in your lab on purpose and fix them. Delete a user with and without the Recycle Bin. Restart a DC into DSRM. Stop replication between sites by blocking a port and read the repadmin output. Revert a member server to an old checkpoint and repair its secure channel. Move the PDC emulator and reconfigure time. In Azure, create a DCR that collects the System log from an Arc-enabled server and an alert on a log query, and enable VM insights on one VM.