Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1922943

Summary: Elasticsearch cluster is not able to initialize
Product: OpenShift Container Platform Reporter: Devendra Kulkarni <dkulkarn>
Component: LoggingAssignee: ewolinet
Status: CLOSED ERRATA QA Contact: Anping Li <anli>
Severity: high Docs Contact:
Priority: medium    
Version: 4.6CC: aelganzo, andcosta, aos-bugs, bjarolim, broose, ewolinet, kiyyappa, mmohan, naygupta, openshift-bugs-escalate, qitang, rsandu, ssonigra, ykarajag
Target Milestone: ---   
Target Release: 4.6.z   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard: logging-exploration
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-04-20 19:20:20 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Devendra Kulkarni 2021-02-01 07:27:11 UTC
Description of problem: After upgrading RHOCP cluster to 4.6 and upgrading clusterlogging to v4.6, logging stack works fine for some days and then the Elasticsearch cluster is not able to initialize.

Pod status:

elasticsearch-cdm-agramz9f-1-59d9686859-8h2mb  1/2    Running    0         1h15m
elasticsearch-cdm-agramz9f-2-6c6dc6c754-l9wnf  1/2    Running    0         1h15m
elasticsearch-cdm-agramz9f-3-768f9d486c-gf5dc  1/2    Running    0         1h15m


Pod logs:

2021-01-27T07:10:44.461449322Z [2021-01-27T07:10:44,460][ERROR][c.a.o.s.a.BackendRegistry] [elasticsearch-cdm-agramz9f-1] Not yet initialized 
2021-01-27T07:11:05.618668540Z [2021-01-27T07:11:05,618][ERROR][c.a.o.s.a.BackendRegistry] [elasticsearch-cdm-agramz9f-1] Not yet initialized 
2021-01-27T07:11:14.412493347Z [2021-01-27T07:11:14,412][ERROR][c.a.o.s.a.BackendRegistry] [elasticsearch-cdm-agramz9f-1] Not yet initialized 
2021-01-27T07:11:35.618385546Z [2021-01-27T07:11:35,618][ERROR][c.a.o.s.a.BackendRegistry] [elasticsearch-cdm-agramz9f-1] Not yet initialized 



Version-Release number of selected component (if applicable): 4.6


How reproducible: 


Steps to Reproduce:
1.
2.
3.

Actual results:


Expected results: ES cluster should be initialized properly.


Additional info: 
[1] Restarting of ES pods does not help.
[2] Reseeding ES cluster does not help.


Workaround: Wipe the backend PVCs and then restart the ES pods.

Comment 1 Hui Kang 2021-02-01 14:37:12 UTC
@dev, do you see any relevant alert in the monitoring dashboard? Or could you post more logs from the ES pod?

Comment 9 Jeff Cantrill 2021-02-15 17:21:44 UTC
You may be experiencing a cert gen issue that we are correcting and is partially resolved by https://github.com/openshift/cluster-logging-operator/pull/849 and further corrected by https://github.com/openshift/elasticsearch-operator/pull/636.

Workaround:

* Setting both clusterlogging/instance and the elasticsearch/elasticsearch to "Unmanaged"
* Scaling down the elasticsearch deployments to 0 replicas
* Editing the elasticsearch deployments to set 'paused' to 'false'
* Scaling up the es deployments to 1 replica
* Observe if the es pods cluster and the cluster becomes at least into 'yellow'
* Set both clusterlogging/instance and the elasticsearch/elasticsearch to "managed"

Comment 12 Jeff Cantrill 2021-02-17 18:26:15 UTC
(In reply to Sonigra Saurab from comment #11)
> Workaround in comment #9 does not work for case # 02868728

Can you expand?  Are you seeing the same symptoms?

Comment 34 ewolinet 2021-03-17 18:47:23 UTC
It looks like for multiple cases Elasticsearch fails to recover because a node detects a conflict with its data while recovering.

In one case it was an issue with a Kibana index/alias and another it was the write index that was being pointed to by an alias.

Looking for solutions online, the only thing that seems to come up [1] is that the node in question should have its storage wiped and then restarted and then let the cluster rebalance itself -- this is only possible if the cluster was configured with some redundancy (e.g. not ZeroRedundancy) and there is only one node with this issue.

There is a PR to try to help further prevent this (as Elasticsearch typically prevents itself from getting into this state) for the case of the write index (it will remove any duplicate write indices for a particular alias).


[1] https://discuss.elastic.co/t/node-wont-start-up-due-to-dupliate-alias-index-names/174548/17

Comment 41 ewolinet 2021-03-25 19:00:46 UTC
Between the linked pull request and a docs guide [1] to help recover from this state in the off-chance it happens (given Elasticsearch actively tries to prevent this normally -- you cannot manually add another index as a write alias, the request is rejected) we have a path forward.

The PR is to help prevent this while the cluster is still running, the docs update is for when the cluster has restarted and gotten to this point (given the Elasticsearch API will not be available due to it not recovering/starting up) manual intervention is necessary.


[1] https://issues.redhat.com/browse/RHDEVDOCS-2850

Comment 46 Anping Li 2021-04-07 11:54:30 UTC
Verified on elasticsearch-operator.4.6.0-202103270037.p0

Comment 51 errata-xmlrpc 2021-04-20 19:20:20 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory (OpenShift Container Platform 4.6.25 extras update), and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2021:1155