Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
For bugs related to Red Hat Enterprise Linux 5 product line. The current stable release is 5.10. For Red Hat Enterprise Linux 6 and above, please visit Red Hat JIRA https://issues.redhat.com/secure/CreateIssue!default.jspa?pid=12332745 to report new issues.

Bug 500446

Summary: [RHEL5.4] igb: debug kernel reveals incorrect call used to free multiqueue netdev
Product: Red Hat Enterprise Linux 5 Reporter: Andy Gospodarek <agospoda>
Component: kernelAssignee: Andy Gospodarek <agospoda>
Status: CLOSED ERRATA QA Contact: Red Hat Kernel QE team <kernel-qe>
Severity: medium Docs Contact:
Priority: low    
Version: 5.4CC: dzickus, jburke, jtluka, peterm
Target Milestone: rc   
Target Release: ---   
Hardware: All   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: Bug Fix
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2009-09-02 08:53:16 UTC Type: ---
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Andy Gospodarek 2009-05-12 18:00:26 UTC
When running with a debug kernel and shutting down the system, we get the following:

INIT: Switching to runINIT: Sending processes the TERM signal
Shutting down Avahi daemon: [  OK  ]
Stopping HAL daemon: [  OK  ]
Stopping yum-updatesd: [  OK  ]
Stopping anacron: [  OK  ]
Stopping atd: [  OK  ]
Stopping cups: [  OK  ]
Shutting down xfs: [  OK  ]
Shutting down console mouse services: [  OK  ]
Stopping sshd: [  OK  ]
Shutting down sm-client: [  OK  ]
Shutting down sendmail: [  OK  ]
Stopping xinetd: [  OK  ]
Stopping acpi daemon: [  OK  ]
Stopping crond: [  OK  ]
Stopping autofs:  Stopping automount: [  OK  ]
[  OK  ]
Shutting down ntpd: [  OK  ]
Stopping ipsec:  could not open include filename: '/etc/ipsec.d/*.conf' (tried  and )
ipsec_setup: Stopping Openswan IPsec...
[  OK  ]
Stopping system message bus: [  OK  ]
Stopping RPC idmapd: [  OK  ]
Stopping NFS statd: [  OK  ]
Stopping mcstransd: [  OK  ]
Stopping portmap: [  OK  ]
Shutting down restorecond: [  OK  ]
Stopping auditd: audit(1241895760.072:19): audit_pid=0 old=4011 by auid=4294967295 subj=system_u:system_r:auditd_t:s0
[  OK  ]
Stopping PC/SC smart card daemon (pcscd): [  OK  ]
Shutting down kernel logger: [  OK  ]
Shutting down system logger: [  OK  ]
Shutting down hidd: [  OK  ]
[  OK  ][  OK  ]Stopping Bluetooth services:[  OK  ][  OK  ]
Disabling ondemand cpu frequency scaling: [  OK  ]
Starting killall:  [  OK  ]
Sending all processes the TERM signal... 
type=1701 audit(1241895763.356:20): auid=4294967295 uid=0 gid=0 ses=4294967295 subj=system_u:system_r:iscsid_t:s0 pid=3566 comm="iscsid" sig=11
Sending all processes the KILL signal... 
Saving random seed:  
Syncing hardware clock to system time type=1111 audit(1241895768.998:21): user pid=9258 uid=0 auid=4294967295 subj=system_u:system_r:hwclock_t:s0 msg='changing system time: exe="/sbin/hwclock" (hostname=?, addr=?, terminal=console res=success)'

Turning off swap:  
Turning off quotas:  
Unmounting pipe file systems:  
Unmounting file systems:  
Please stand by while rebooting the system...
md: stopping all md devices.
Synchronizing SCSI cache for disk sda: 
slab error in verify_redzone_free(): cache `size-2048': memory outside object was overwritten
 [<c0477005>] cache_free_debugcheck+0xaa/0x195
 [<f898a982>] igb_free_queues+0x2b/0x60 [igb]
 [<c0477342>] kfree+0x7d/0xc7
 [<f898a982>] igb_free_queues+0x2b/0x60 [igb]
 [<f898c416>] igb_suspend+0x4c/0x160 [igb]
 [<c050061d>] pci_device_shutdown+0x13/0x14
 [<c0566bab>] device_shutdown+0x47/0x6c
 [<c0430666>] kernel_restart+0x8/0x39
 [<c04307e1>] sys_reboot+0x143/0x1b4
 [<c06222e3>] schedule+0xa1f/0xaab
 [<c040adf2>] sched_clock+0x8b/0x9c
 [<c0438e76>] lock_release_holdtime+0x25/0x43
 [<c0624514>] _spin_unlock_irqrestore+0x34/0x39
 [<c043a331>] trace_hardirqs_on+0xf8/0x118
 [<c0437d89>] hrtimer_try_to_cancel+0x3c/0x42
 [<c0437d99>] hrtimer_cancel+0xa/0x14
 [<c0623354>] do_nanosleep+0x43/0x6a
 [<c0437ed2>] hrtimer_nanosleep+0x50/0x106
 [<c0438e76>] lock_release_holdtime+0x25/0x43
 [<c0450009>] audit_syscall_entry+0x14b/0x17d
 [<c04080e3>] do_syscall_trace+0xab/0xb1
 [<c0404f93>] syscall_call+0x7/0xb
 =======================
f5a8adfc: redzone 1:0x0, redzone 2:0x170fc2a5.
------------[ cut here ]------------
kernel BUG at mm/slab.c:2796!
invalid opcode: 0000 [#1]
SMP 
last sysfs file: /devices/system/cpu/cpu1/cpufreq/cpuinfo_max_freq
Modules linked in: deflate zlib_deflate ccm serpent blowfish twofish ecb xcbc crypto_hash cbc md5 sha256 sha512 des aes_generic testmgr_cipher testmgr crypto_blkcipher aes_i586 xfrm6_esp xfrm4_esp aead crypto_algapi tunnel4 xfrm6_tunnel tunnel6 autofs4 hidp rfcomm l2cap bluetooth sunrpc ipv6 xfrm_nalgo crypto_api ib_iser rdma_cm ib_cm iw_cm ib_sa ib_mad ib_core ib_addr iscsi_tcp libiscsi_tcp libiscsi2 scsi_transport_iscsi2 scsi_transport_iscsi acpi_cpufreq dm_multipath scsi_dh video hwmon backlight sbs i2c_ec button battery asus_acpi ac lp i2c_i801 sr_mod i2c_core cdrom igb parport_pc parport sg pcspkr dm_raid45 dm_message dm_region_hash dm_mem_cache dm_snapshot dm_zero dm_mirror dm_log dm_mod ahci libata sd_mod scsi_mod ext3 jbd uhci_hcd ohci_hcd ehci_hcd
CPU:    10
EIP:    0060:[<c047707d>]    Not tainted VLI
EFLAGS: 00010002   (2.6.18-145.el5dz_testdebug #1) 
EIP is at cache_free_debugcheck+0x122/0x195
eax: f5a8adf4   ebx: f7ffc0c0   ecx: 0000080c   edx: 00000008
esi: f5a8adfc   edi: f5a8a5e8   ebp: 00000001   esp: c26fae3c
ds: 007b   es: 007b   ss: 0068
Process reboot (pid: 8259, ti=c26fa000 task=f735f120 task.ti=c26fa000)
Stack: f898a982 f5a8a5c0 f7ffc0c0 c27c05d4 f5a8ae00 00000286 c0477342 f5cc2400 
       00000094 00000001 00000002 f898a982 f5cc25b8 f5cc2000 f7e82ef8 f898c416 
       00000002 f5cc2400 28121969 f7e8374c 28121969 00e2aff4 c26fa000 c050061d 
Call Trace:
 [<f898a982>] igb_free_queues+0x2b/0x60 [igb]
 [<c0477342>] kfree+0x7d/0xc7
 [<f898a982>] igb_free_queues+0x2b/0x60 [igb]
 [<f898c416>] igb_suspend+0x4c/0x160 [igb]
 [<c050061d>] pci_device_shutdown+0x13/0x14
 [<c0566bab>] device_shutdown+0x47/0x6c
 [<c0430666>] kernel_restart+0x8/0x39
 [<c04307e1>] sys_reboot+0x143/0x1b4
 [<c06222e3>] schedule+0xa1f/0xaab
 [<c040adf2>] sched_clock+0x8b/0x9c
 [<c0438e76>] lock_release_holdtime+0x25/0x43
 [<c0624514>] _spin_unlock_irqrestore+0x34/0x39
 [<c043a331>] trace_hardirqs_on+0xf8/0x118
 [<c0437d89>] hrtimer_try_to_cancel+0x3c/0x42
 [<c0437d99>] hrtimer_cancel+0xa/0x14
 [<c0623354>] do_nanosleep+0x43/0x6a
 [<c0437ed2>] hrtimer_nanosleep+0x50/0x106
 [<c0438e76>] lock_release_holdtime+0x25/0x43
 [<c0450009>] audit_syscall_entry+0x14b/0x17d
 [<c04080e3>] do_syscall_trace+0xab/0xb1
 [<c0404f93>] syscall_call+0x7/0xb
 =======================
Code: 8c 00 00 00 8b 78 0c 89 f0 29 f8 f7 f1 3b 83 98 00 00 00 89 c5 72 08 0f 0b eb 0a 67 b4 64 c0 89 e8 0f af c1 8d 04 07 39 c6 74 08 <0f> 0b ec 0a 67 b4 64 c0 f6 83 95 00 00 00 02 74 15 89 f0 b9 05 
EIP: [<c047707d>] cache_free_debugcheck+0x122/0x195 SS:ESP 0068:c26fae3c
 <0>Kernel panic - not syncing: Fatal exception
 BUG: warning at arch/i386/kernel/smp.c:550/smp_call_function() (Not tainted)
 [<c0415b7a>] stop_this_cpu+0x0/0x38
 [<c041595b>] smp_call_function+0x57/0xc5
 [<c04253e7>] printk+0x18/0x8e
 [<c04159dc>] smp_send_stop+0x13/0x26
 [<c0424903>] panic+0x4c/0x171
 [<c040660f>] die+0x25d/0x291
 [<c0406cce>] do_invalid_op+0x0/0x9d
 [<c0406d5f>] do_invalid_op+0x91/0x9d
 [<c047707d>] cache_free_debugcheck+0x122/0x195
 [<c0424d05>] release_console_sem+0x19c/0x1df
 [<c04253c3>] vprintk+0x30f/0x31b
 [<c0404f93>] syscall_call+0x7/0xb
 [<c04253e7>] printk+0x18/0x8e
 [<c0404f93>] syscall_call+0x7/0xb
 [<c0405bc5>] error_code+0x39/0x40
 [<c047707d>] cache_free_debugcheck+0x122/0x195
 [<f898a982>] igb_free_queues+0x2b/0x60 [igb]
 [<c0477342>] kfree+0x7d/0xc7
 [<f898a982>] igb_free_queues+0x2b/0x60 [igb]
 [<f898c416>] igb_suspend+0x4c/0x160 [igb]
 [<c050061d>] pci_device_shutdown+0x13/0x14
 [<c0566bab>] device_shutdown+0x47/0x6c
 [<c0430666>] kernel_restart+0x8/0x39
 [<c04307e1>] sys_reboot+0x143/0x1b4
 [<c06222e3>] schedule+0xa1f/0xaab
 [<c040adf2>] sched_clock+0x8b/0x9c
 [<c0438e76>] lock_release_holdtime+0x25/0x43
 [<c0624514>] _spin_unlock_irqrestore+0x34/0x39
 [<c043a331>] trace_hardirqs_on+0xf8/0x118
 [<c0437d89>] hrtimer_try_to_cancel+0x3c/0x42
 [<c0437d99>] hrtimer_cancel+0xa/0x14
 [<c0623354>] do_nanosleep+0x43/0x6a
 [<c0437ed2>] hrtimer_nanosleep+0x50/0x106
 [<c0438e76>] lock_release_holdtime+0x25/0x43
 [<c0450009>] audit_syscall_entry+0x14b/0x17d
 [<c04080e3>] do_syscall_trace+0xab/0xb1
 [<c0404f93>] syscall_call+0x7/0xb


This is a result of the multiqueue netdev being allocated with the call alloc_netdev, but being freed with kfree rather than free_netdev.  This patch fixes the issue:

--- a/drivers/net/igb/igb_main.c
+++ b/drivers/net/igb/igb_main.c
@@ -285,7 +285,7 @@ static void igb_free_queues(struct igb_adapter *adapter)
 	for (i = 0; i < adapter->num_rx_queues; i++) {
 		struct igb_ring *ring = &(adapter->rx_ring[i]);
 		ring->count = adapter->rx_ring_count;
-		kfree(ring->netdev);
+		free_netdev(ring->netdev);
 	}
 
 	adapter->num_rx_queues = 0;

Comment 2 RHEL Program Management 2009-05-12 18:19:20 UTC
This request was evaluated by Red Hat Product Management for inclusion in a Red
Hat Enterprise Linux maintenance release.  Product Management has requested
further review of this request by Red Hat Engineering, for potential
inclusion in a Red Hat Enterprise Linux Update release for currently deployed
products.  This request is not yet committed for inclusion in an Update
release.

Comment 3 Don Zickus 2009-05-14 19:35:57 UTC
in kernel-2.6.18-148.el5
You can download this test kernel from http://people.redhat.com/dzickus/el5

Please do NOT transition this bugzilla state to VERIFIED until our QE team
has sent specific instructions indicating when to do so.  However feel free
to provide a comment indicating that this fix has been verified.

Comment 5 Jan Tluka 2009-07-20 15:12:52 UTC
Patch is in -158.el5. Adding SanityOnly.

Comment 7 errata-xmlrpc 2009-09-02 08:53:16 UTC
An advisory has been issued which should help the problem
described in this bug report. This report is therefore being
closed with a resolution of ERRATA. For more information
on therefore solution and/or where to find the updated files,
please follow the link below. You may reopen this bug report
if the solution does not work for you.

http://rhn.redhat.com/errata/RHSA-2009-1243.html