Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1829332

Summary: node-exporter crashes due to bad /proc/cpuinfo parsing
Product: OpenShift Container Platform Reporter: Pablo Alonso Rodriguez <palonsor>
Component: MonitoringAssignee: Paul Gier <pgier>
Status: CLOSED DUPLICATE QA Contact: Junqi Zhao <juzhao>
Severity: low Docs Contact:
Priority: unspecified    
Version: 4.2.zCC: alegrand, anpicker, erooth, kakkoyun, lcosic, mloibl, pkrupa, surbania
Target Milestone: ---   
Target Release: 4.5.0   
Hardware: s390x   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-04-30 12:32:59 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Pablo Alonso Rodriguez 2020-04-29 12:03:54 UTC
Description of problem:

On a freshly installed s390x cluster, all the node-exporter pods crash-loop due to constant panics due to what looks like a bad processing of /proc/cpuinfo output.

Version-Release number of selected component (if applicable):

4.2.13

How reproducible:

Always

Steps to Reproduce:
1. Deploy a fresh OpenShift cluster on IBM Z on z/VM hypervisor

Actual results:

Clusteroperator monitoring degraded and node-exporter pods crash looping

Expected results:

Clusteroperator monitoring not degraded and node-exporter pods working.

Additional info:

Relevant stack trace is:

panic: runtime error: index out of range

goroutine 30 [running]:
github.com/prometheus/procfs.parseCPUInfo(0xc00017b000, 0x4ee, 0xe00, 0x4ee, 0xe00, 0x0, 0x0, 0x1)
        /go/src/github.com/prometheus/node_exporter/vendor/github.com/prometheus/procfs/cpuinfo.go:85 +0x1d32
github.com/prometheus/procfs.FS.CPUInfo(0x3ffcb4ff571, 0xa, 0x24000000000f4684, 0x1d394, 0xc0001c3d50, 0x70, 0x70)
        /go/src/github.com/prometheus/node_exporter/vendor/github.com/prometheus/procfs/cpuinfo.go:61 +0x180
github.com/prometheus/node_exporter/collector.(*cpuCollector).updateInfo(0xc00005c940, 0xc000084cc0, 0x35d4c8c600000000, 0x1e47f087846)
        /go/src/github.com/prometheus/node_exporter/collector/cpu_linux.go:96 +0x3c
github.com/prometheus/node_exporter/collector.(*cpuCollector).Update(0xc00005c940, 0xc000084cc0, 0xc311a0, 0x0)
        /go/src/github.com/prometheus/node_exporter/collector/cpu_linux.go:81 +0xdc
github.com/prometheus/node_exporter/collector.execute(0x6e60b2, 0x3, 0x7cb140, 0xc00005c940, 0xc000084cc0)
        /go/src/github.com/prometheus/node_exporter/collector/collector.go:127 +0x68
github.com/prometheus/node_exporter/collector.NodeCollector.Collect.func1(0xc000084cc0, 0xc00007ef40, 0x6e60b2, 0x3, 0x7cb140, 0xc00005c940)
        /go/src/github.com/prometheus/node_exporter/collector/collector.go:118 +0x48
created by github.com/prometheus/node_exporter/collector.NodeCollector.Collect
        /go/src/github.com/prometheus/node_exporter/collector/collector.go:117 +0xe2

(more information in comments)