Replication between primary and additional domain controllers has failed: how should AD, DNS, time, and SYSVOL be checked?
Domain-controller replication failures affect accounts, passwords, GPOs, and sign-ins. Start with replication summaries and error codes, then verify DNS, sites, RPC, time, and DFSR.
Domain-controller replication failures affect accounts, passwords, GPOs, and sign-ins. For this case, first verify repadmin replication error codes and AD DNS and site subnets, then use RPC, dynamic ports and firewall rules to decide whether remediation is needed.
Define the failure boundary first
For this directory and identity case, establish the failure boundary with repadmin and AD DNS, then continue to RPC, dynamic ports and firewall rules. Capture the current state, incident time and one known-good comparison before changing production configuration.
Work through the dependency chain
| Check | Why it matters | Recommended action |
|---|---|---|
| 01 · repadmin replication error codes | Verify repadmin replication error codes on the affected path using logs, counters or state information rather than relying only on the configured rule. | Record the current value, evidence source and timestamp for repadmin replication error codes. If adjustment is required, change one condition only and retain the original setting for rollback. |
| 02 · AD DNS and site subnets | Verify AD DNS and site subnets on the affected path using logs, counters or state information rather than relying only on the configured rule. | Check AD DNS and site subnets read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation. |
| 03 · RPC, dynamic ports and firewall rules | Verify RPC, dynamic ports and firewall rules on the affected path using logs, counters or state information rather than relying only on the configured rule. | Compare RPC, dynamic ports and firewall rules with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production. |
| 04 · domain-controller time synchronization | Review the current state, related logs and recent changes for domain-controller time synchronization, then align them with the incident timeline before deciding whether a change is required. | Record the current value, evidence source and timestamp for domain-controller time synchronization. If adjustment is required, change one condition only and retain the original setting for rollback. |
| 05 · DFSR/SYSVOL state | Review the current state, related logs and recent changes for DFSR/SYSVOL state, then align them with the incident timeline before deciding whether a change is required. | Check DFSR/SYSVOL state read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation. |
| 06 · directory objects, USN and database health | Review the current state, related logs and recent changes for directory objects, USN and database health, then align them with the incident timeline before deciding whether a change is required. | Compare directory objects, USN and database health with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production. |
repadmin /replsummary
repadmin /showrepl
dcdiag /test:dnsChange only after the evidence is clear
- Start with read-only evidence. Check repadmin replication error codes and AD DNS and site subnets before changing configuration.
- If the first checks are normal, continue with RPC, dynamic ports and firewall rules and domain-controller time synchronization, keeping evidence tied to the incident time.
- Change configuration only when the evidence explains the symptom. For DFSR/SYSVOL state, preserve the original value and define the rollback trigger before adjustment.
- Validate directory objects, USN and database health in a controlled scope before expanding to production users or traffic.
Validation and rollback
- Validate the complete user or application workflow; do not stop at the single status of repadmin replication error codes.
- Recheck DFSR/SYSVOL state and directory objects, USN and database health after the change and confirm that no new bypass, permission expansion or secondary error has appeared.
- Archive evidence from repadmin replication error codes through directory objects, USN and database health, together with before/after configuration, business validation and the rollback point.
Common wrong turns
- Changing repadmin and AD DNS at the same time, which makes the original cause impossible to prove.
- Treating a normal result for RPC, dynamic ports and firewall rules as proof that domain-controller time synchronization and the rest of the business path are healthy.
- Leaving a temporary exception related to DFSR/SYSVOL or directory objects, USN and database health in production without an owner, expiry time and rollback note.
Related questions
Where should I start with “Replication between primary and additional domain controllers has failed: how should AD, DNS, time, and SYSVOL be checked?”?
Start with repadmin replication error codes and AD DNS and site subnets; they establish the first useful troubleshooting boundary without changing production state.
What should be checked after the first layer looks normal?
Continue with RPC, dynamic ports and firewall rules and domain-controller time synchronization, then correlate the result with the incident time and the actual user or application path.
What should be retained after the change?
Keep evidence for DFSR/SYSVOL state and directory objects, USN and database health, plus the original configuration, validation result, observation notes and rollback point.
