Bug 2521546 - kernel 7.1.8: Oops "Split lock detected" in skb_clone via icmp_unreach/tcp_v4_err, repeated panics
Summary: kernel 7.1.8: Oops "Split lock detected" in skb_clone via icmp_unreach/tcp_v4...
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Fedora
Classification: Fedora
Component: kernel
Version: 44
Hardware: x86_64
OS: Linux
unspecified
urgent
Target Milestone: ---
Assignee: Justin M. Forbes
QA Contact: Fedora Extras Quality Assurance
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2026-08-22 22:19 UTC by don
Modified: 2026-09-04 01:28 UTC (History)
14 users (show)

Fixed In Version: kernel-7.1.13-200.fc44 kernel-7.1.13-100.fc43
Clone Of:
Environment:
Last Closed: 2026-09-04 01:11:40 UTC
Type: ---
Embargoed:


Attachments (Terms of Use)
vmcore-dmesg.txt from the kdump capture (107.17 KB, text/plain)
2026-08-22 22:22 UTC, don
no flags Details
vmcore-dmesg.txt from the kdump capture (106.81 KB, text/plain)
2026-08-22 22:29 UTC, don
no flags Details

Description don 2026-08-22 22:19:16 UTC
1. Please describe the problem:

The machine panics with "Oops: Split lock detected" in skb_clone, reached from the ICMP error path (icmp_rcv -> icmp_unreach -> tcp_v4_err -> ip_icmp_error -> skb_clone) in softirq context while a userland process (NetworkManager in the captured case) was in a TCP connect. The faulting atomic uses a misaligned pointer (RDX = ffff8a3b5e29881f), which suggests a corrupted or mis-built skb rather than a split-lock policy issue. The kernel is not tainted.

Three panics on 2026-08-22 (19:57, 22:00, 22:56 local) on a Lenovo ThinkPad 21QE004QUK (BIOS N4JET22W 1.12, Intel Core Ultra, iwlwifi, e1000e, LUKS root). Only the third was captured by kdump. In each case the minutes before the crash show tailscaled repeatedly failing to reach its DERP relay ("could not connect to the 'London' relay server", 49 times in the captured boot), so the ICMP error path was being exercised by a stream of failed outbound connects. The host also runs NetBird (wireguard, tun) and nftables rules.

Oops header:
  Oops: Split lock detected: 0000 [#1] SMP NOPTI
  CPU: 10 UID: 0 PID: 1215 Comm: NetworkManager Kdump: loaded Not tainted 7.1.8-200.fc44.x86_64 #1 PREEMPT(lazy)
  RIP: 0010:skb_clone+0x159/0x1e0
Call trace (IRQ part): ip_icmp_error, tcp_v4_err, icmp_unreach, icmp_rcv, ip_protocol_deliver_rcu, ip_local_deliver_finish, __netif_receive_skb_one_core, process_backlog, __napi_poll, net_rx_action, handle_softirqs, do_softirq
Task part: __dev_queue_xmit, ip_finish_output2, ip_output, __ip_queue_xmit, __tcp_transmit_skb, tcp_connect, tcp_v4_connect, __inet_stream_connect, __sys_connect
Full log attached (vmcore-dmesg.txt); the 394 MB vmcore is available on request.

2. What is the Version-Release number of the kernel:

kernel-7.1.8-200.fc44.x86_64

3. Did it work previously in Fedora? If so, what kernel version did the issue *first* appear?

Unknown. kernel-6.19.10-300.fc44 is still installed and no panic was observed while it was the running kernel, but it has not yet been run long enough under the same network conditions to call this a regression. I am now testing 7.1.9-200.fc44 and will report back.

4. Can you reproduce this issue? If so, please provide the steps to reproduce the issue below:

Not deterministically. Three occurrences in about three hours, all during periods of repeated failed outbound TCP connects (tailscaled probing an unreachable relay) while NetBird and nftables were active. Uptime at the captured crash was 2932 s.

5. Does this problem occur with the latest Rawhide kernel?

Not tested.

6. Are you running any modules that not shipped with directly Fedora's kernel?:

No. The kernel is untainted; loaded modules are in-tree (wireguard, tun, iwlwifi, e1000e, nf_conntrack, nft_*, ip_set, xt_*). The host runs SentinelOne and Tailscale/NetBird daemons, which load eBPF programs but no kernel modules.

7. Please attach the kernel logs:

Attached: vmcore-dmesg.txt from /var/crash/127.0.0.1-2026-08-22-22:56:47 (kdump capture of the third panic). The earlier two boots' journals end abruptly with no kernel message.

Reproducible: Always

Comment 1 don 2026-08-22 22:22:13 UTC
Created attachment 2155317 [details]
vmcore-dmesg.txt from the kdump capture

Comment 2 don 2026-08-22 22:28:13 UTC
Comment on attachment 2155317 [details]
vmcore-dmesg.txt from the kdump capture

Could a kernel triager please make this private - I have uploaded a scrubbed version.

Comment 3 don 2026-08-22 22:29:23 UTC
Created attachment 2155318 [details]
vmcore-dmesg.txt from the kdump capture

Comment 4 junjie.cao 2026-09-01 00:59:14 UTC
Decoded the Oops: at skb_clone+0x159 the Code bytes are

  8b 93 c0 00 00 00    mov  0xc0(%rbx),%edx   ; skb->end
  48 03 93 c8 00 00 00 add  0xc8(%rbx),%rdx   ; + skb->head = skb_shinfo(skb)
  f0 ff 42 20          lock incl 0x20(%rdx)   ; atomic_inc(&shinfo->dataref)

RDX = ffff8a3b5e29881f, so dataref sits at ...883f: a 4-byte atomic
crossing a 64-byte line. skb_shared_info is misaligned by 0x1f because
skb->head came from a page_pool fragment at an odd offset.

Root cause is upstream: page_pool_alloc_frag_netmem() aligns fragment
sizes to dma_get_cache_alignment(), which is 1 on x86, so one odd-sized
request (skb_pp_cow_data() on the generic-XDP path -- NetBird/tailscale
attach an XDP program to lo) leaves the per-cpu pool's frag_offset
misaligned for every buffer carved out after it. Fixed in net.git,
Cc: stable:
https://git.kernel.org/netdev/net/c/dc0df5a0c62c

Fedora 44 backport so 7.1.y users don't wait for the stable round trip:
https://gitlab.com/cki-project/kernel-ark/-/merge_requests/4730

Why only Intel machines crash: a kernel-mode #AC dies outright
(arch/x86/kernel/traps.c exc_alignment_check -> die("Split lock
detected")); AMD has no split-lock detection, so the same misaligned
atomic just runs slowly there. Workaround until the fix lands:
boot with split_lock_detect=off.

Comment 5 Fedora Update System 2026-09-02 17:55:15 UTC
FEDORA-2026-a7b1ccd14c (kernel-7.1.13-100.fc43) has been submitted as an update to Fedora 43.
https://bodhi.fedoraproject.org/updates/FEDORA-2026-a7b1ccd14c

Comment 6 Fedora Update System 2026-09-02 17:55:57 UTC
FEDORA-2026-0d885c0533 (kernel-7.1.13-200.fc44) has been submitted as an update to Fedora 44.
https://bodhi.fedoraproject.org/updates/FEDORA-2026-0d885c0533

Comment 7 Fedora Update System 2026-09-03 01:39:02 UTC
FEDORA-2026-0d885c0533 has been pushed to the Fedora 44 testing repository.
Soon you'll be able to install the update with the following command:
`sudo dnf upgrade --enablerepo=updates-testing --refresh --advisory=FEDORA-2026-0d885c0533`
You can provide feedback for this update here: https://bodhi.fedoraproject.org/updates/FEDORA-2026-0d885c0533

See also https://fedoraproject.org/wiki/QA:Updates_Testing for more information on how to test updates.

Comment 8 Fedora Update System 2026-09-03 01:54:41 UTC
FEDORA-2026-a7b1ccd14c has been pushed to the Fedora 43 testing repository.
Soon you'll be able to install the update with the following command:
`sudo dnf upgrade --enablerepo=updates-testing --refresh --advisory=FEDORA-2026-a7b1ccd14c`
You can provide feedback for this update here: https://bodhi.fedoraproject.org/updates/FEDORA-2026-a7b1ccd14c

See also https://fedoraproject.org/wiki/QA:Updates_Testing for more information on how to test updates.

Comment 9 Fedora Update System 2026-09-04 01:11:40 UTC
FEDORA-2026-0d885c0533 (kernel-7.1.13-200.fc44) has been pushed to the Fedora 44 stable repository.
If problem still persists, please make note of it in this bug report.

Comment 10 Fedora Update System 2026-09-04 01:28:03 UTC
FEDORA-2026-a7b1ccd14c (kernel-7.1.13-100.fc43) has been pushed to the Fedora 43 stable repository.
If problem still persists, please make note of it in this bug report.


Note You need to log in before you can comment on or make changes to this bug.