Bug 1922943
| Summary: | Elasticsearch cluster is not able to initialize | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | Devendra Kulkarni <dkulkarn> |
| Component: | Logging | Assignee: | ewolinet |
| Status: | CLOSED ERRATA | QA Contact: | Anping Li <anli> |
| Severity: | high | Docs Contact: | |
| Priority: | medium | ||
| Version: | 4.6 | CC: | aelganzo, andcosta, aos-bugs, bjarolim, broose, ewolinet, kiyyappa, mmohan, naygupta, openshift-bugs-escalate, qitang, rsandu, ssonigra, ykarajag |
| Target Milestone: | --- | ||
| Target Release: | 4.6.z | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | logging-exploration | ||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2021-04-20 19:20:20 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Devendra Kulkarni
2021-02-01 07:27:11 UTC
@dev, do you see any relevant alert in the monitoring dashboard? Or could you post more logs from the ES pod? You may be experiencing a cert gen issue that we are correcting and is partially resolved by https://github.com/openshift/cluster-logging-operator/pull/849 and further corrected by https://github.com/openshift/elasticsearch-operator/pull/636. Workaround: * Setting both clusterlogging/instance and the elasticsearch/elasticsearch to "Unmanaged" * Scaling down the elasticsearch deployments to 0 replicas * Editing the elasticsearch deployments to set 'paused' to 'false' * Scaling up the es deployments to 1 replica * Observe if the es pods cluster and the cluster becomes at least into 'yellow' * Set both clusterlogging/instance and the elasticsearch/elasticsearch to "managed" (In reply to Sonigra Saurab from comment #11) > Workaround in comment #9 does not work for case # 02868728 Can you expand? Are you seeing the same symptoms? It looks like for multiple cases Elasticsearch fails to recover because a node detects a conflict with its data while recovering. In one case it was an issue with a Kibana index/alias and another it was the write index that was being pointed to by an alias. Looking for solutions online, the only thing that seems to come up [1] is that the node in question should have its storage wiped and then restarted and then let the cluster rebalance itself -- this is only possible if the cluster was configured with some redundancy (e.g. not ZeroRedundancy) and there is only one node with this issue. There is a PR to try to help further prevent this (as Elasticsearch typically prevents itself from getting into this state) for the case of the write index (it will remove any duplicate write indices for a particular alias). [1] https://discuss.elastic.co/t/node-wont-start-up-due-to-dupliate-alias-index-names/174548/17 Between the linked pull request and a docs guide [1] to help recover from this state in the off-chance it happens (given Elasticsearch actively tries to prevent this normally -- you cannot manually add another index as a write alias, the request is rejected) we have a path forward. The PR is to help prevent this while the cluster is still running, the docs update is for when the cluster has restarted and gotten to this point (given the Elasticsearch API will not be available due to it not recovering/starting up) manual intervention is necessary. [1] https://issues.redhat.com/browse/RHDEVDOCS-2850 Verified on elasticsearch-operator.4.6.0-202103270037.p0 Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (OpenShift Container Platform 4.6.25 extras update), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2021:1155 |