Insights /Network, VPN and firewall

DNS returns Server Failure or SERVFAIL: checking forwarders, firewall port 53 and cache

SERVFAIL means the resolver could not complete the query. Check authoritative data, forwarders, TCP/UDP 53, DNSSEC, timeouts and cache in sequence.

Quick answer

SERVFAIL means the resolver could not complete the query. For this case, first verify authoritative-zone state and forwarder reachability and timeout, then use TCP/UDP 53 firewall policy to decide whether remediation is needed.

Define the failure boundary first

For this network and security-boundary case, establish the failure boundary with authoritative-zone state and forwarder reachability and timeout, then continue to TCP/UDP 53. Capture the current state, incident time and one known-good comparison before changing production configuration.

Work through the dependency chain

CheckWhy it mattersRecommended action
01 · authoritative-zone stateVerify authoritative-zone state on the affected path using logs, counters or state information rather than relying only on the configured rule.Record the current value, evidence source and timestamp for authoritative-zone state. If adjustment is required, change one condition only and retain the original setting for rollback.
02 · forwarder reachability and timeoutVerify forwarder reachability and timeout on the affected path using logs, counters or state information rather than relying only on the configured rule.Check forwarder reachability and timeout read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation.
03 · TCP/UDP 53 firewall policyVerify TCP/UDP 53 firewall policy on the affected path using logs, counters or state information rather than relying only on the configured rule.Compare TCP/UDP 53 firewall policy with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production.
04 · recursion and root-hint configurationReview the current state, related logs and recent changes for recursion and root-hint configuration, then align them with the incident timeline before deciding whether a change is required.Record the current value, evidence source and timestamp for recursion and root-hint configuration. If adjustment is required, change one condition only and retain the original setting for rollback.
05 · DNSSEC/EDNS compatibilityReview the current state, related logs and recent changes for DNSSEC/EDNS compatibility, then align them with the incident timeline before deciding whether a change is required.Check DNSSEC/EDNS compatibility read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation.
06 · negative cache and server cacheReview the current state, related logs and recent changes for negative cache and server cache, then align them with the incident timeline before deciding whether a change is required.Compare negative cache and server cache with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production.
Read-only examples
nslookup example.com 192.0.2.53
Resolve-DnsName example.com -Server 192.0.2.53

Change only after the evidence is clear

  1. Start with read-only evidence. Check authoritative-zone state and forwarder reachability and timeout before changing configuration.
  2. If the first checks are normal, continue with TCP/UDP 53 firewall policy and recursion and root-hint configuration, keeping evidence tied to the incident time.
  3. Change configuration only when the evidence explains the symptom. For DNSSEC/EDNS compatibility, preserve the original value and define the rollback trigger before adjustment.
  4. Validate negative cache and server cache in a controlled scope before expanding to production users or traffic.

Validation and rollback

  • Validate the complete user or application workflow; do not stop at the single status of authoritative-zone state.
  • Recheck DNSSEC/EDNS compatibility and negative cache and server cache after the change and confirm that no new bypass, permission expansion or secondary error has appeared.
  • Archive evidence from authoritative-zone state through negative cache and server cache, together with before/after configuration, business validation and the rollback point.

Common wrong turns

  • Changing authoritative-zone state and forwarder reachability and timeout at the same time, which makes the original cause impossible to prove.
  • Treating a normal result for TCP/UDP 53 as proof that recursion and root-hint configuration and the rest of the business path are healthy.
  • Leaving a temporary exception related to DNSSEC/EDNS or negative cache and server cache in production without an owner, expiry time and rollback note.

Related questions

Where should I start with “DNS returns Server Failure or SERVFAIL: checking forwarders, firewall port 53 and cache”?

Start with authoritative-zone state and forwarder reachability and timeout; they establish the first useful troubleshooting boundary without changing production state.

What should be checked after the first layer looks normal?

Continue with TCP/UDP 53 firewall policy and recursion and root-hint configuration, then correlate the result with the incident time and the actual user or application path.

What should be retained after the change?

Keep evidence for DNSSEC/EDNS compatibility and negative cache and server cache, plus the original configuration, validation result, observation notes and rollback point.

PreviousDesigning DNS conditional forwarders for an isolated network that must resolve only vendor domainsNextWith two enterprise DNS servers, must forwarders, cache and root hints be identical?

Need an assessment based on the actual environment?