Bug 2508154 - NULL function pointer call in mid8250_probe() panics boot on Intel Atom C3000 (Denverton) HSUART [8086:19d8] since kernel 7.1.4; 7.1.3-201 works
Summary: NULL function pointer call in mid8250_probe() panics boot on Intel Atom C3000...
Keywords:
Status: CLOSED COMPLETED
Alias: None
Product: Fedora
Classification: Fedora
Component: kernel
Version: 44
Hardware: x86_64
OS: Linux
unspecified
urgent
Target Milestone: ---
Assignee: Justin M. Forbes
QA Contact: Fedora Extras Quality Assurance
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2026-07-28 18:28 UTC by Arcadiy Ivanov
Modified: 2026-08-05 08:47 UTC (History)
15 users (show)

Fixed In Version:
Clone Of:
Environment:
Last Closed: 2026-08-05 08:47:48 UTC
Type: ---
Embargoed:


Attachments (Terms of Use)
Screen 1 (11.60 MB, image/jpeg)
2026-07-28 18:30 UTC, Arcadiy Ivanov
no flags Details
Screen 2 (10.58 MB, image/jpeg)
2026-07-28 18:30 UTC, Arcadiy Ivanov
no flags Details
Screen 3 (8.87 MB, image/jpeg)
2026-07-28 18:31 UTC, Arcadiy Ivanov
no flags Details

Description Arcadiy Ivanov 2026-07-28 18:28:36 UTC
1. Please describe the problem:

On a headless Intel Atom C3758R (Denverton) appliance, kernels 7.1.4-204.fc44
and 7.1.5-200.fc44 panic during boot. Kernel 7.1.3-201.fc44 works.

The 8250_mid driver calls a NULL function pointer while probing the onboard
Atom C3000 HSUART controller (PCI 8086:19d8). Because this happens in
do_initcalls() on PID 1, init dies and the kernel panics.

Note the panic is INVISIBLE on a default boot. It occurs after console_init()
has registered tty0 against a dummy console device, but before simpledrm binds
the EFI framebuffer, so no console is attached when the oops prints. The only
symptom without instrumentation is a machine that goes silent right after GRUB
and never comes up. It was only made visible by booting with:

    earlycon=efifb keep_bootcon ignore_loglevel

The panic also cannot be recovered from pstore on this machine, because the
ERST backend is full:

    pstore: backend (erst) writing error (-28)

Oops and call trace (transcribed from console, kernel 7.1.5-200.fc44.x86_64):

  RIP: 0010:0x0
  Code: Unable to access opcode bytes at 0xffffffffffffffd6.
  RSP: 0000:ffffce37c00278f8 EFLAGS: 00010286
  RAX: 0000000000000000 RBX: 0000000000000040 RCX: 0000000000000000
  RDX: ffffce37c00b1000 RSI: ffffce37c0027908 RDI: ffff8b7301e61038
  RBP: ffff8b7302a19000 R08: 0000000000000000 R09: 0000000000000000
  R10: ffff8b7301158a80 R11: ffffce37c00b1fff R12: ffff8b7301e61038
  R13: ffff8b7302a190d0 R14: ffffffffb3bba2e8 R15: 0000000000000006
  FS:  0000000000000000(0000) GS:ffff8b8289c96000(0000) knlGS:0000000000000000
  CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
  CR2: ffffffffffffffd6 CR3: 000000020342c000 CR4: 00000000003506f0
  Call Trace:
   <TASK>
   mid8250_probe+0x197/0x310
   ? __pfx_mid8250_set_termios+0x10/0x10
   local_pci_probe+0x4a/0x90
   ? cpu_hotplug_disable+0x15/0x30
   pci_call_probe+0x5a/0x1a0
   ? pci_assign_irq+0x30/0x1a0
   ? pci_match_device+0x13e/0x180
   pci_device_probe+0xa8/0x1b0
   ? __pfx_pci_device_probe+0x10/0x10
   call_driver_probe+0x23/0x100
   ? driver_sysfs_add+0x55/0xc0
   really_probe+0xd7/0x2d0
   ? __pm_runtime_resume+0x5f/0x90
   __driver_probe_device+0x8f/0x190
   driver_probe_device+0x1f/0xa0
   ? __pfx___driver_attach+0x10/0x10
   __driver_attach+0xca/0x1f0
   ? __pfx___driver_attach+0x10/0x10
   bus_for_each_dev+0x97/0x100
   bus_add_driver+0x15a/0x260
   ? __pfx_mid8250_pci_driver_init+0x10/0x10
   driver_register+0x75/0xe0
   do_one_initcall+0x8e/0x3f0
   do_initcalls+0x150/0x180
   kernel_init_freeable+0x107/0x160
   ? __pfx_kernel_init+0x10/0x10
   kernel_init+0x1a/0x140
   ret_from_fork+0x1a1/0x270
   ? __pfx_kernel_init+0x10/0x10
   ret_from_fork_asm+0x1a/0x30
   </TASK>
  Modules linked in:
  CR2: 0000000000000000
  ---[ end trace 0000000000000000 ]---
  pstore: backend (erst) writing error (-28)
  note: swapper/0[1] exited with irqs disabled
  Kernel panic - not syncing: Attempted to kill init! exitcode=0x00000009
  Kernel Offset: 0x31200000 from 0xffffffff81000000 (relocation range:
    0xffffffff80000000-0xffffffffbfffffff)
  ---[ end Kernel panic - not syncing: Attempted to kill init!
    exitcode=0x00000009 ]---

RIP is 0x0, i.e. control was transferred through a NULL function pointer
rather than a NULL data dereference.

For comparison, the same device probes successfully on 7.1.3-201:

  Serial: 8250/16550 driver, 32 ports, IRQ sharing enabled
  8250_mid 0000:00:1a.0: enabling device (0141 -> 0143)
  8250_mid 0000:00:1a.0: Found HSU DMA, 2 channels
  0000:00:1a.0: ttyS4 at MMIO 0xdf419000 (irq = 37, base_baud = 115200) is a TI16750

2. What is the Version-Release number of the kernel:

Broken:  kernel-core-7.1.4-204.fc44.x86_64
         kernel-core-7.1.5-200.fc44.x86_64
Working: kernel-core-7.1.3-201.fc44.x86_64

3. Did it work previously in Fedora? If so, what kernel version did the issue
   *first* appear?  Old kernels are available for download at
   https://koji.fedoraproject.org/koji/packageinfo?packageID=8 :

Yes. Regression window is 7.1.3-201 -> 7.1.4-204. Every earlier kernel on this
machine worked: 7.0.4-200, 7.0.6-200, 7.0.8-200, 7.0.9-204, 7.0.10-200/201,
7.0.12-201, 7.0.13-200, 7.0.14-201, 7.1.3-200, 7.1.3-201.

4. Can you reproduce this issue? If so, please provide the steps to reproduce
   the issue below:

Always reproducible on this hardware.

  - Step 1. Boot kernel 7.1.4-204 or 7.1.5-200 on an Atom C3000 board with the
            onboard HSUART (PCI 8086:19d8) enabled.
  - Step 2. The machine goes silent after GRUB and never comes up.
  - Step 3. Add "earlycon=efifb keep_bootcon ignore_loglevel" to see the panic.
  - Step 4. Boot 7.1.3-201; the system comes up normally.

5. Does this problem occur with the latest Rawhide kernel? To install the
   Rawhide kernel, run ``sudo dnf install fedora-repos-rawhide`` followed by
   ``sudo dnf update --enablerepo=rawhide kernel``:

Not tested.


6. Are you running any modules that not shipped with directly Fedora's kernel?:

akmod-karellen-intel-qat-1x (Intel QAT, locally built). Not a factor: the
panic occurs in a built-in initcall before any out-of-tree module loads, and
the same module is present on the working 7.1.3-201 boot.

7. Please attach the kernel logs. You can get the complete kernel log
   for a boot with ``journalctl --no-hostname -k > dmesg.txt``. If the
   issue occurred on a previous boot, use the journalctl ``-b`` flag.


CPU:       Intel(R) Atom(TM) CPU C3758R @ 2.40GHz (Denverton, Goldmont)
           family 6, model 95, stepping 1, 8 cores, microcode 0x3e
Board:     QDNV01 / QDNV01, American Megatrends BIOS 5.13, 02/21/2024
Memory:    64 GB
Firmware:  UEFI, EFI v2.6, Secure Boot disabled
Device:    00:1a.0 Serial controller: Intel Corporation Atom Processor C3000
           Series HSUART Controller [8086:19d8] (rev 11), driver 8250_mid

Kernel command line (working and failing entries are identical apart from the
debug options added in step 3):

  root=UUID=... ro intel_iommu=on cpufreq.default_governor=performance

Reproducible: Always

Comment 1 Arcadiy Ivanov 2026-07-28 18:30:02 UTC
Created attachment 2152882 [details]
Screen 1

Comment 2 Arcadiy Ivanov 2026-07-28 18:30:54 UTC
Created attachment 2152883 [details]
Screen 2

Comment 3 Arcadiy Ivanov 2026-07-28 18:31:21 UTC
Created attachment 2152884 [details]
Screen 3

Comment 4 Justin M. Forbes 2026-07-28 19:09:38 UTC
Possible to give the rawhide kernel a try? That at least tells us if it already fixed upstream, or if it still needs a fix.  As can occasionally happen in stable, one fix gets backported and something that it is dependent upon does not.

Comment 5 Arcadiy Ivanov 2026-07-28 19:49:35 UTC
I'll try.

Comment 6 Arcadiy Ivanov 2026-07-28 19:51:19 UTC
@Justin

========================================================================
sha    : 7fb13fd7e9a59a37cd911efff83abe19e3ee029d
date   : 2026-07-15T07:35:46Z
author : Jiangshan Yi

serial: 8250_mid: Fix NULL function pointer dereference on DNV/ICX-D/SNR platforms

Commit b1b4efea05a5 ("serial: 8250_mid: Disable DMA for selected
platforms") replaced the dnv_board setup and exit callbacks with
PTR_IF(false, ...), which evaluates to NULL. However, the three call
sites in mid8250_probe() and mid8250_remove() unconditionally
dereference these function pointers without NULL checks, causing a NULL
pointer dereference (kernel oops) on any Denverton (DNV), Ice Lake Xeon
D (ICX-D/CDF), or Snowridge (SNR) platform.

Fix this by adding the missing NULL checks before calling the setup and
exit callbacks.

Fixes: b1b4efea05a5 ("serial: 8250_mid: Disable DMA for selected platforms")
Cc: stable <stable>
Reviewed-by: Andy Shevchenko <andriy.shevchenko.com>
Signed-off-by: Jiangshan Yi <yijiangshan>
Link: https://patch.msgid.link/20260715073546.1875083-1-yijiangshan@kylinos.cn
Signed-off-by: Greg Kroah-Hartman <gregkh>

files  : ['drivers/tty/serial/8250/8250_mid.c']
========================================================================
sha    : b1b4efea05a56c0995e4702a86d6624b4fdff32f
date   : 2026-06-26T09:49:37Z
author : Andy Shevchenko

serial: 8250_mid: Disable DMA for selected platforms

In accordance with Errata (specification updates)
HSUART May Stop Functioning when DMA is Active.

- Denverton document #572409, rev 3.4, DNV60
- Ice Lake Xeon D document #714070, ICXD65
- Snowridge document #731931, SNR44

For a quick fix just disable the respective callbacks during the device probe.
Depending on the future development we might remove them completely.

Reported-by: micas-opensource <zjianan156>
Closes: https://lore.kernel.org/linux-serial/20250625031409.2404219-1-opensource@ruijie.com.cn/
Fixes: 6ede6dcd87aa ("serial: 8250_mid: add support for DMA engine handling from UART MMIO")
Cc: stable <stable>
Signed-off-by: Andy Shevchenko <andriy.shevchenko.com>
Link: https://patch.msgid.link/20260626094937.561776-1-andriy.shevchenko@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh>

files  : ['drivers/tty/serial/8250/8250_mid.c']

Comment 7 Justin M. Forbes 2026-07-28 20:01:13 UTC
Nice find, I have queued it for the next kernel builds

Comment 8 Arcadiy Ivanov 2026-07-28 20:15:08 UTC
Thank you!

Comment 9 Arcadiy Ivanov 2026-08-05 08:47:48 UTC
Confirm 7.1.6-201.f44 is back to working condition on Danverton.


Note You need to log in before you can comment on or make changes to this bug.