Bug 2527805 - kernel NULL dereference in listening_get_first during /proc/net/tcp read leaves listener spinlock held and causes system soft-lockup
Summary: kernel NULL dereference in listening_get_first during /proc/net/tcp read leav...
Keywords:
Status: NEW
Alias: None
Product: Fedora
Classification: Fedora
Component: kernel
Version: 44
Hardware: x86_64
OS: Linux
unspecified
urgent
Target Milestone: ---
Assignee: Justin M. Forbes
QA Contact: Fedora Extras Quality Assurance
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2026-09-02 23:31 UTC by villeaj.poutiainen
Modified: 2026-09-02 23:36 UTC (History)
14 users (show)

Fixed In Version:
Clone Of:
Environment:
Last Closed:
Type: ---
Embargoed:


Attachments (Terms of Use)
important Oops + subsequent lockup evidence (71.44 KB, text/plain)
2026-09-02 23:35 UTC, villeaj.poutiainen
no flags Details
contains GDB/disassembly/type information (22.79 KB, text/plain)
2026-09-02 23:35 UTC, villeaj.poutiainen
no flags Details
Fedora kernel configuration (292.79 KB, text/plain)
2026-09-02 23:36 UTC, villeaj.poutiainen
no flags Details

Description villeaj.poutiainen 2026-09-02 23:31:59 UTC
1. Please describe the problem:

Fedora 44 experienced a kernel NULL-pointer dereference while a long-running
pasta.avx2 process was reading /proc/net/tcp as part of RootlessKit/pasta
automatic TCP/UDP port discovery.

The initial kernel failure was:

BUG: kernel NULL pointer dereference, address: 0000000000000230
#PF: supervisor read access in kernel mode
Oops: 0000 [#1] SMP NOPTI
CPU: 1 UID: 1000 PID: 1634 Comm: pasta.avx2
RIP: listening_get_first+0xc0/0x120

Call path:

listening_get_first
tcp_get_idx
tcp_seq_start
seq_read_iter
seq_read
proc_reg_read
vfs_read
ksys_read

The userspace syscall was read(fd=16, ...).

I obtained the exact Fedora debuginfo/vmlinux matching the crashed kernel,
7.1.10-200.fc44.x86_64.

vmlinux Build ID:

4c9470ef910bac97b4554b740c84f337cada5950

GDB analysis of that exact binary gives these structure offsets:

struct seq_file.file     = 0x60
struct file.f_inode      = 0x20
struct inode.i_private   = 0x230

The relevant compiled instructions in listening_get_first are:

+0x8f: call _raw_spin_lock
...
+0xb8: mov 0x60(%rbp),%rdx
+0xbc: mov 0x20(%rdx),%rdx
+0xc0: mov 0x230(%rdx),%rdx
...
+0xff: call _raw_spin_unlock

At the fault:

RDX = 0
CR2 = 0x230

The exact binary/disassembly therefore shows that
seq->file->f_inode was NULL while the active /proc/net/tcp read was in
progress. The fault at +0xc0 occurred when the code then attempted to access
inode->i_private at offset 0x230.

The function had already acquired the TCP listener hash-bucket spinlock at
+0x8f. The normal _raw_spin_unlock() path is later at +0xff, so the Oops
occurred while that spinlock was held.

The kernel continued running after the Oops. It also reported:

note: pasta.avx2[1634] exited with irqs disabled
note: pasta.avx2[1634] exited with preempt_count 1

A few minutes later, process "ss", PID 934233, became stuck at 100% system
CPU in:

native_queued_spin_lock_slowpath
_raw_spin_lock
listening_get_first+0x94
tcp_get_idx
tcp_seq_next
seq_read_iter

The listener-bucket lock address present in the original pasta Oops was:

ffff8f0d05894ec0

The later ss soft-lockup was attempting to acquire that same address:

ffff8f0d05894ec0

The watchdog first reported ss stuck for 26 seconds, then continued reporting
the same stack at 52, 85, 112, 138, 164, 190, 216 and at least 253 seconds.
RCU subsequently reported both normal and expedited stalls on the affected
CPU.

The machine became effectively unusable and required a forced reboot.

Crash-time system and userspace configuration:

Fedora 44 x86_64
AMD Ryzen 9 9950X
Sapphire NITRO+ B850M WIFI
64 GiB RAM
NVIDIA RTX 3090

passt/pasta: 0^20260728.gf8df3f1-2.fc44
RootlessKit: 3.1.0-1.fc44
moby-engine: 29.7.2-1.fc44

The passt package had been installed on 2026-08-07, before the affected
RootlessKit/pasta process started on 2026-08-26, so the exact pasta source
version involved is known.

At the time of the crash, rootless Docker launched RootlessKit with:

--net=pasta
--port-driver=implicit

This caused pasta to use automatic TCP/UDP port discovery. That mode keeps
/proc/net TCP/UDP descriptors open and periodically rescans them; the
crashing read(fd=16, ...) is consistent with that scan path.


The first Oops was tainted:

G OE

The O/E flags correspond to out-of-tree/unsigned NVIDIA kernel modules.

I found no MCE or other CPU/RAM hardware-error report associated with this
kernel Oops. However, I have not reproduced the failure on different
hardware or on an untainted kernel, so I am not claiming that hardware or
unrelated kernel memory corruption has been conclusively ruled out.

What is established by the captured Oops and exact binary analysis is:

1. seq->file->f_inode was NULL during the /proc/net/tcp iteration.
2. this NULL dereference occurred while a TCP listener hash-bucket spinlock
   was held;
3. the Oops prevented the normal unlock path from executing; and
4. a later /proc/net/tcp reader ("ss") became permanently stuck attempting
   to acquire the same lock address.

What caused the live seq_file/file state to have f_inode == NULL is not yet
known. A kernel file-lifetime/reuse bug or other memory corruption are
possible explanations, but neither has been proven.

I have not assessed exploitability or security impact. The demonstrated impact is a local system-wide denial of service, but it is not known whether the underlying condition can be intentionally triggered by an unprivileged user.


2. What is the Version-Release number of the kernel:

7.1.10-200.fc44.x86_64


3. Did it work previously in Fedora? If so, what kernel version did the issue
   *first* appear? Old kernels are available for download at
   https://koji.fedoraproject.org/koji/packageinfo?packageID=8 :

Unknown.

This failure has been observed once. I have not established the first
affected kernel version and therefore do not claim that this is a regression.

The machine is currently running 7.1.12-200.fc44.x86_64, but the
RootlessKit configuration has also been changed from the crash-time
"implicit" port driver to "pesto". In the current configuration pasta uses:

--tcp-ports=none
--udp-ports=none

Therefore the absence of another crash on 7.1.12 is not a valid test of
whether the kernel bug is still present.


4. Can you reproduce this issue? If so, please provide the steps to reproduce
   the issue below:

Not currently reproduced deliberately.

The observed triggering environment was a long-running RootlessKit/pasta
instance configured with:

--net=pasta
--port-driver=implicit

which caused pasta to perform periodic automatic TCP/UDP port discovery by
reading the proc networking socket tables, including /proc/net/tcp.

I have not deliberately restored that configuration and attempted to force
the crash because the observed failure can leave a kernel spinlock held,
produce RCU/soft-lockup stalls, and require a hard reset.

I still have a Timeshift snapshot from 2026-08-31 21:00, approximately nine hours before the crash, which may be useful later for a controlled reproduction
environment.

I currently have significant workloads running on this machine and do not
have enough spare CPU resources to construct and run a dedicated
reproduction VM.


5. Does this problem occur with the latest Rawhide kernel? To install the
   Rawhide kernel, run `sudo dnf install fedora-repos-rawhide` followed by
   `sudo dnf update --enablerepo=rawhide kernel`:

Not tested.

If maintainers identify a specific reproduction or diagnostic test that
would materially help isolate the cause, I can preserve the available
pre-crash system snapshot and attempt such testing when resources permit.


6. Are you running any modules that not shipped with directly Fedora's kernel?:

Yes.

The NVIDIA kernel modules were loaded:

nvidia
nvidia_modeset
nvidia_drm
nvidia_uvm

The initial Oops was tainted G OE (out-of-tree and unsigned modules).

There is currently no evidence tying the NVIDIA modules to the
/proc/net/tcp/listening_get_first failure, but I am reporting the taint
explicitly. I have not reproduced the problem on an untainted kernel.


7. Please attach the kernel logs. You can get the complete kernel log
   for a boot with `journalctl --no-hostname -k > dmesg.txt`. If the
   issue occurred on a previous boot, use the journalctl `-b` flag.

Attached:

kernel-crash-public.txt
kernel-config.txt
vmlinux-analysis-public.txt

kernel-crash-public.txt is a sanitized extract of the affected boot's kernel
journal beginning at the initial pasta.avx2 NULL-pointer Oops and continuing
through all subsequent ss watchdog/RCU soft-lockup reports. The hostname was
replaced with "HOST"; unrelated earlier kernel-log content was omitted.

vmlinux-analysis-public.txt contains GDB output from the exact Fedora
7.1.10-200.fc44.x86_64 debuginfo/vmlinux, including the complete
listening_get_first disassembly, source-line mappings for +0x94 and +0xc0,
and the exact struct seq_file, struct file and struct inode layouts.

kernel-config.txt is the exact configuration of the affected
7.1.10-200.fc44.x86_64 kernel.

I intentionally did not attach the complete system journal because it
contains substantial unrelated application information. If
additional surrounding kernel or service-log evidence is needed, I can
provide targeted sanitized excerpts.

Reproducible: Always

Comment 1 villeaj.poutiainen 2026-09-02 23:35:05 UTC
Created attachment 2156511 [details]
important Oops + subsequent lockup evidence

Comment 2 villeaj.poutiainen 2026-09-02 23:35:43 UTC
Created attachment 2156512 [details]
contains GDB/disassembly/type information

Comment 3 villeaj.poutiainen 2026-09-02 23:36:16 UTC
Created attachment 2156513 [details]
Fedora kernel configuration


Note You need to log in before you can comment on or make changes to this bug.