Bug 1578720
| Summary: | master is restart again and again due to etcd dns resolve become unavailable or timeout | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | Johnny Liu <jialiu> |
| Component: | Installer | Assignee: | Scott Dodson <sdodson> |
| Status: | CLOSED ERRATA | QA Contact: | Johnny Liu <jialiu> |
| Severity: | high | Docs Contact: | |
| Priority: | urgent | ||
| Version: | 3.10.0 | CC: | aos-bugs, hongli, jialiu, jokerman, lxia, mifiedle, mmccomas, sdodson, wmeng, wsun |
| Target Milestone: | --- | Keywords: | TestBlocker |
| Target Release: | 3.10.0 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | No Doc Update | |
| Doc Text: |
undefined
|
Story Points: | --- |
| Clone Of: | Environment: | ||
| Last Closed: | 2018-07-30 19:15:30 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Johnny Liu
2018-05-16 09:22:33 UTC
This is blocking all the cluster which is using etcd dns url resolved by internal DNS server. Reverted the change that's suspected to have broken this. https://github.com/openshift/openshift-ansible/pull/8409 @Scott, according to my testing in my initial report, do you know why cluster internal DNS still could be resolved even without "server=/cluster.local/127.0.0.1" setting? The node is going to send a message to dnsmasq via dbus when it starts it stops dynamically changing the config. When the node is running it will forward requests to the dnsRecursiveResolvConf defined in node-config.yaml Is the node not resolving reverse lookups properly? # dig @127.0.0.1 -x 8.8.8.8 ;; ANSWER SECTION: 8.8.8.8.in-addr.arpa. 53 IN PTR google-public-dns-a.google.com. @QA Contact,the PR 8409 has been merged to 3.10.0-0.50.0,please check the bug. Verified this bug with openshift-ansible-3.10.0-0.50.0.git.0.bd68ade.el7.noarch, and PASS. Run a system container install with AH on GCE, installation is completed successfully. root@qe-jialiu310-master-etcd-1 ~]# cat /etc/resolv.conf # nameserver updated by /etc/NetworkManager/dispatcher.d/99-origin-dns.sh # Generated by NetworkManager search cluster.local c.openshift-gce-devel.internal google.internal nameserver 10.240.0.48 # oc get nodes NAME STATUS ROLES AGE VERSION qe-jialiu310-master-etcd-1 Ready master 3h v1.10.0+b81c8f8 qe-jialiu310-node-registry-router-1 Ready compute 3h v1.10.0+b81c8f8 [root@qe-jialiu310-master-etcd-1 ~]# time ping -c 1 qe-jialiu310-node-registry-router-1 PING qe-jialiu310-node-registry-router-1.c.openshift-gce-devel.internal (10.240.0.49) 56(84) bytes of data. 64 bytes from qe-jialiu310-node-registry-router-1.c.openshift-gce-devel.internal (10.240.0.49): icmp_seq=1 ttl=64 time=1.26 ms --- qe-jialiu310-node-registry-router-1.c.openshift-gce-devel.internal ping statistics --- 1 packets transmitted, 1 received, 0% packet loss, time 0ms rtt min/avg/max/mdev = 1.263/1.263/1.263/0.000 ms real 0m0.083s user 0m0.002s sys 0m0.005s [root@qe-jialiu310-master-etcd-1 ~]# ls /etc/dnsmasq.d/ origin-dns.conf origin-upstream-dns.conf [root@qe-jialiu310-master-etcd-1 ~]# cat /etc/dnsmasq.d/* no-resolv domain-needed no-negcache max-cache-ttl=1 enable-dbus dns-forward-max=10000 cache-size=10000 bind-dynamic except-interface=lo # End of config server=169.254.169.254 [root@qe-jialiu310-master-etcd-1 ~]# oc logs nodejs-mongodb-example-1-build -n install-test <--snip--> Pushing image docker-registry.default.svc:5000/install-test/nodejs-mongodb-example:latest ... Pushed 5/6 layers, 91% complete Pushed 6/6 layers, 100% complete Push successful [root@qe-jialiu310-master-etcd-1 ~]# cat /etc/redhat-release Red Hat Enterprise Linux Atomic Host release 7.5 Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2018:1816 |