Insights /Backup, NAS and business continuity

Should TrueNAS use hardware RAID, or should ZFS manage the disks directly?

ZFS needs direct visibility of disks, SMART data, and error states. An HBA or JBOD mode is usually preferred, followed by vdev design based on performance, capacity, and rebuild windows.

Quick answer

ZFS needs direct visibility of disks, SMART data, and error states. For this case, first verify HBA/JBOD passthrough and direct disk visibility to ZFS, then use SMART and error visibility to decide whether remediation is needed.

Start with the business decision

For this backup and storage case, establish the failure boundary with HBA/JBOD passthrough and ZFS, then continue to SMART and error visibility. Capture the current state, incident time and one known-good comparison before changing production configuration.

Where the options actually differ

CheckWhy it mattersRecommended action
01 · HBA/JBOD passthroughVerify HBA/JBOD passthrough on the affected path using logs, counters or state information rather than relying only on the configured rule.Record the current value, evidence source and timestamp for HBA/JBOD passthrough. If adjustment is required, change one condition only and retain the original setting for rollback.
02 · direct disk visibility to ZFSVerify direct disk visibility to ZFS on the affected path using logs, counters or state information rather than relying only on the configured rule.Check direct disk visibility to ZFS read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation.
03 · SMART and error visibilityVerify SMART and error visibility on the affected path using logs, counters or state information rather than relying only on the configured rule.Compare SMART and error visibility with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production.
04 · vdev fault-tolerance designReview the current state, related logs and recent changes for vdev fault-tolerance design, then align them with the incident timeline before deciding whether a change is required.Record the current value, evidence source and timestamp for vdev fault-tolerance design. If adjustment is required, change one condition only and retain the original setting for rollback.
05 · rebuild duration and URE riskReview the current state, related logs and recent changes for rebuild duration and URE risk, then align them with the incident timeline before deciding whether a change is required.Check rebuild duration and URE risk read-only and save the result. If it differs from the baseline, correlate it with the incident time and recent changes before remediation.
06 · cache and power-loss protectionReview the current state, related logs and recent changes for cache and power-loss protection, then align them with the incident timeline before deciding whether a change is required.Compare cache and power-loss protection with a known-good peer, the log timeline and the real application path; confirm whether it is causal before changing production.

Questions to answer before selection

  1. Start with read-only evidence. Check HBA/JBOD passthrough and direct disk visibility to ZFS before changing configuration.
  2. If the first checks are normal, continue with SMART and error visibility and vdev fault-tolerance design, keeping evidence tied to the incident time.
  3. Change configuration only when the evidence explains the symptom. For rebuild duration and URE risk, preserve the original value and define the rollback trigger before adjustment.
  4. Validate cache and power-loss protection in a controlled scope before expanding to production users or traffic.

Implementation and operating impact

  • Validate the complete user or application workflow; do not stop at the single status of HBA/JBOD passthrough.
  • Recheck rebuild duration and URE risk and cache and power-loss protection after the change and confirm that no new bypass, permission expansion or secondary error has appeared.
  • Archive evidence from HBA/JBOD passthrough through cache and power-loss protection, together with before/after configuration, business validation and the rollback point.

Acceptance criteria

  • Changing HBA/JBOD passthrough and ZFS at the same time, which makes the original cause impossible to prove.
  • Treating a normal result for SMART and error visibility as proof that vdev and the rest of the business path are healthy.
  • Leaving a temporary exception related to URE or cache and power-loss protection in production without an owner, expiry time and rollback note.

Related questions

Where should I start with “Should TrueNAS use hardware RAID, or should ZFS manage the disks directly?”?

Start with HBA/JBOD passthrough and direct disk visibility to ZFS; they establish the first useful troubleshooting boundary without changing production state.

What should be checked after the first layer looks normal?

Continue with SMART and error visibility and vdev fault-tolerance design, then correlate the result with the incident time and the actual user or application path.

What should be retained after the change?

Keep evidence for rebuild duration and URE risk and cache and power-loss protection, plus the original configuration, validation result, observation notes and rollback point.

PreviousCan duplicate hostnames across virtual desktops cause domain trust, DNS, Group Policy, and sign-in problems?NextIf ransomware encrypts a shared file server, how can backups be protected from deletion or encryption at the same time?

Need an assessment based on your actual environment?