Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
For bugs related to Red Hat Enterprise Linux 5 product line. The current stable release is 5.10. For Red Hat Enterprise Linux 6 and above, please visit Red Hat JIRA https://issues.redhat.com/secure/CreateIssue!default.jspa?pid=12332745 to report new issues.

Bug 461658

Summary: kdump: incomplete header in vmcore
Product: Red Hat Enterprise Linux 5 Reporter: Prarit Bhargava <prarit>
Component: kernelAssignee: Neil Horman <nhorman>
Status: CLOSED DUPLICATE QA Contact: Martin Jenner <mjenner>
Severity: medium Docs Contact:
Priority: medium    
Version: 5.3CC: anderson, clalance, jturner, vgoyal
Target Milestone: rc   
Target Release: ---   
Hardware: All   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: Bug Fix
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2008-09-10 21:15:00 UTC Type: ---
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Prarit Bhargava 2008-09-09 19:03:28 UTC
Description of problem:

After kdumping a system the resulting vmcore has an incomplete header.

Version-Release number of selected component (if applicable): latest RHEL-5 kernel, -109.el5


How reproducible: 100%


Steps to Reproduce:
1. Build kernel with prarit's followup patch from BZ 450244.
2. Install, reboot, and crash using echo c > /proc/sysrq-trigger
  
Actual results:

vmcore has incomplete header


Expected results:  vmcore should have a complete header

[root@xeratul 2008-09-09-13:01]# readelf -a vmcore
ELF Header:
 Magic:   7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
 Class:                             ELF64
 Data:                              2's complement, little endian
 Version:                           1 (current)
 OS/ABI:                            UNIX - System V
 ABI Version:                       0
 Type:                              CORE (Core file)
 Machine:                           Advanced Micro Devices X86-64
 Version:                           0x1
 Entry point address:               0x0
 Start of program headers:          64 (bytes into file)
 Start of section headers:          0 (bytes into file)
 Flags:                             0x0
 Size of this header:               64 (bytes)
 Size of program headers:           56 (bytes)
 Number of program headers:         4
 Size of section headers:           0 (bytes)
 Number of section headers:         0
 Section header string table index: 0

There are no sections in this file.

There are no sections in this file.

Program Headers:
 Type           Offset             VirtAddr           PhysAddr
                FileSiz            MemSiz              Flags  Align
 NOTE           0x0000000000000120 0x0000000000000000 0x0000000000000000
                0x00000000000003f8 0x00000000000003f8         0
 LOAD           0x0000000000000518 0xffff810000000000 0x0000000000000000
                0x00000000000a0000 0x00000000000a0000  RWE    0
 LOAD           0x00000000000a0518 0xffff810000100000 0x0000000000100000
                0x0000000000f00000 0x0000000000f00000  RWE    0
 LOAD           0x0000000000fa0518 0xffff810009000000 0x0000000009000000
                0x0000000036e8ac00 0x0000000036e8ac00  RWE    0

There is no dynamic section in this file.

There are no relocations in this file.

There are no unwind sections in this file.

No version information found in this file.

Notes at offset 0x00000120 with length 0x000003f8:
 Owner         Data size       Description
 VMCOREINFO            0x000003de      Unknown note type: (0x00000000)

Additional info:  This kernel can be easily reproduced by checking out the latest RHEL5 git tree (version -109.el5), adding my followup patch from 450244, and compiling.

Comment 1 Prarit Bhargava 2008-09-09 19:08:34 UTC
nhorman, FYI -- this was the issue that anderson was talking about on RHKL after I posted my one-line followup patch for 450244.

P.

Comment 2 Neil Horman 2008-09-09 19:28:16 UTC
Ok, thanks.  Just out of curiousity, have you checked to see if the -97.el5 kernel has this problem as well?  I figure I'll start this with a bisection, just as before to see where it broke, since I have known good cores from the -92 kernel

Comment 3 Neil Horman 2008-09-09 20:55:34 UTC
Prarit, Vivek, I just tested 2.6.18-97.el5 on nec-em16.rhts.bos.redhat.com, and crashing it produces a working vmcore file, so something else between -98 and -109 is causing this to happen.

I've got a kernel with your memmap parsing patch removed prarit, and I'll try that next, although for the life of me, I cant see how that is going to cause a malformed vmcore header.

Comment 4 Dave Anderson 2008-09-09 21:05:52 UTC
/proc/xen on bare-metal kernels?

> Hi Dave,
>
> Following code in kexec-tools, prepares the elf header for kernel text
> region. (kexec-tools-2.0.0/kexec/crashdump-elf.c). You might want to
> just do gdb and see if we are preparing this header or not at the time of
> loading crashdump kernel.
>
>         /* Setup an PT_LOAD type program header for the region where
>          * Kernel is mapped if info->kern_size is non-zero.
>          */
>
>         if (info->kern_size && !xen_present()) {
>                 phdr = (PHDR *) bufp;
>                 bufp += sizeof(PHDR);
>                 phdr->p_type    = PT_LOAD;
>                 phdr->p_flags   = PF_R|PF_W|PF_X;
>                 phdr->p_offset  = phdr->p_paddr = info->kern_paddr_start;
>                 phdr->p_vaddr   = info->kern_vaddr_start;
>                 phdr->p_filesz  = phdr->p_memsz = info->kern_size;
>                 phdr->p_align   = 0;
>                 (elf->e_phnum)++;
>                 dbgprintf_phdr("Kernel text Elf header", phdr);
>         }
>
> Thanks
> Vivek
>

Well, given kexec-tool's xen_present() is this:

int xen_present(void)
{
        struct stat buf;

        return stat("/proc/xen", &buf) == 0;
}

Then xen_present() would return TRUE on my -109 bare-metal kernel:

  # uname -r
  2.6.18-109.el5.prarit
  # file /proc/xen
  /proc/xen: directory
  #

So it skips the segment...

Is this /proc/xen addition fairly new?

Dave

Comment 5 Dave Anderson 2008-09-09 21:16:16 UTC
On an unrelated note, I keep bumping into this error on my test box
with a freshly-installed kexec-tools kexec-tools-1.102pre-36.el5

/etc/init.d/kdump: line 134: [: too many arguments

here:
        if [ -n "$FORCE_REBUILD" -a $modified_files != " " ]
        then
                modified_files="force_rebuild"
        fi

It's because $modified_files is empty.  I'm not sure if it's
symptomatic of some other issue such that $modified_file should
either be a list of files or a space, but I keep getting an
empty string.  So it seems like it should be:

        if [ -n "$FORCE_REBUILD" -a "$modified_files" != " " ]

Comment 6 Neil Horman 2008-09-09 22:21:45 UTC
Dang it!  yes, its rather recent.  The xen people updated the layout of their /proc tree so that all kernels now have a /proc/xen.  I've got another bug on this in the kdump initscript.    Thats probably it

/me is so tired of people changing things in the kernel without testing kdump.

Thats likely it, I'll try write a patch tomorrow.

The modified files thing is queued to be fixed, I've just not checked it in yet.  Sorry about that

Comment 7 Chris Lalancette 2008-09-10 15:44:24 UTC
*** Bug 461784 has been marked as a duplicate of this bug. ***

Comment 8 Chris Lalancette 2008-09-10 15:45:45 UTC
As discussed on rhkernel-list and IRC, this is actually a Xen kernel bug, so it
will be fixed there.  Don Dutile is working on it (this is a public announcement for others who happen by here).

Comment 9 Chris Lalancette 2008-09-10 20:27:19 UTC
*** Bug 461785 has been marked as a duplicate of this bug. ***

Comment 10 Chris Lalancette 2008-09-10 21:15:00 UTC
Actually, since I'm fairly certain this is because the /proc/xen badness, I'm going to close this one as a dup of 461532.  That just explains the issue in a little more detail, and this should naturally be fixed once that is fixed.

Chris Lalancette

*** This bug has been marked as a duplicate of bug 461532 ***