DNS returns Server Failure or SERVFAIL: checking forwarders, firewall port 53 and cache
SERVFAIL means the resolver could not complete the query. Check authoritative data, forwarders, TCP/UDP 53, DNSSEC, timeouts and cache in sequence.
SERVFAIL means the resolver could not complete the query. For this case, first verify authoritative-zone state and forwarder reachability and timeout, then use TCP/UDP 53 firewall policy to decide whether remediation is needed.
Define the failure boundary first
For this network and security-boundary case, establish the failure boundary with authoritative-zone state and forwarder reachability and timeout, then continue to TCP/UDP 53. Capture the current state, incident time and one known-good comparison before changing production configuration.
Work through the dependency chain
| Check | Why it matters | Recommended action |
|---|---|---|
| 01 · authoritative-zone state | Verify authoritative-zone state on the affected path using logs, counters or state information rather than relying only on the configured rule. | Record the current value, evidence source and timestamp for authoritative-zone state. If adjustment is required, change one condition only and retain the original setting for rollback. |
| 02 · forwarder reachability and timeout | Verify forwarder reachability and timeout on the affected path using logs, counters or state information rather than relying only on the configured rule. | Check forwarder reachability and timeout read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation. |
| 03 · TCP/UDP 53 firewall policy | Verify TCP/UDP 53 firewall policy on the affected path using logs, counters or state information rather than relying only on the configured rule. | Compare TCP/UDP 53 firewall policy with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production. |
| 04 · recursion and root-hint configuration | Review the current state, related logs and recent changes for recursion and root-hint configuration, then align them with the incident timeline before deciding whether a change is required. | Record the current value, evidence source and timestamp for recursion and root-hint configuration. If adjustment is required, change one condition only and retain the original setting for rollback. |
| 05 · DNSSEC/EDNS compatibility | Review the current state, related logs and recent changes for DNSSEC/EDNS compatibility, then align them with the incident timeline before deciding whether a change is required. | Check DNSSEC/EDNS compatibility read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation. |
| 06 · negative cache and server cache | Review the current state, related logs and recent changes for negative cache and server cache, then align them with the incident timeline before deciding whether a change is required. | Compare negative cache and server cache with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production. |
nslookup example.com 192.0.2.53
Resolve-DnsName example.com -Server 192.0.2.53Change only after the evidence is clear
- Start with read-only evidence. Check authoritative-zone state and forwarder reachability and timeout before changing configuration.
- If the first checks are normal, continue with TCP/UDP 53 firewall policy and recursion and root-hint configuration, keeping evidence tied to the incident time.
- Change configuration only when the evidence explains the symptom. For DNSSEC/EDNS compatibility, preserve the original value and define the rollback trigger before adjustment.
- Validate negative cache and server cache in a controlled scope before expanding to production users or traffic.
Validation and rollback
- Validate the complete user or application workflow; do not stop at the single status of authoritative-zone state.
- Recheck DNSSEC/EDNS compatibility and negative cache and server cache after the change and confirm that no new bypass, permission expansion or secondary error has appeared.
- Archive evidence from authoritative-zone state through negative cache and server cache, together with before/after configuration, business validation and the rollback point.
Common wrong turns
- Changing authoritative-zone state and forwarder reachability and timeout at the same time, which makes the original cause impossible to prove.
- Treating a normal result for TCP/UDP 53 as proof that recursion and root-hint configuration and the rest of the business path are healthy.
- Leaving a temporary exception related to DNSSEC/EDNS or negative cache and server cache in production without an owner, expiry time and rollback note.
Related questions
Where should I start with “DNS returns Server Failure or SERVFAIL: checking forwarders, firewall port 53 and cache”?
Start with authoritative-zone state and forwarder reachability and timeout; they establish the first useful troubleshooting boundary without changing production state.
What should be checked after the first layer looks normal?
Continue with TCP/UDP 53 firewall policy and recursion and root-hint configuration, then correlate the result with the incident time and the actual user or application path.
What should be retained after the change?
Keep evidence for DNSSEC/EDNS compatibility and negative cache and server cache, plus the original configuration, validation result, observation notes and rollback point.
