Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
This project is now read‑only. Starting Monday, February 2, please use https://ibm-ceph.atlassian.net/ for all bug tracking management.

Bug 2380106

Summary: [NFS-Ganesha] Keepalived enters error state intermittently during node reboot in HA cluster
Product: [Red Hat Storage] Red Hat Ceph Storage Reporter: Manisha Saini <msaini>
Component: CephadmAssignee: Shweta Bhosale <shbhosal>
Status: CLOSED UPSTREAM QA Contact: Vinayak Papnoi <vpapnoi>
Severity: high Docs Contact:
Priority: unspecified    
Version: 8.1CC: akane, aravindr, cephqe-warriors, sabose
Target Milestone: ---   
Target Release: 9.1   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: Known Issue
Doc Text:
Cause: Sometimes keepalived daemon enters into error state. It is intermittent issue. keepalived.conf does not get generated properly. Consequence: Keepalived will not run and VIP will not come online. Workaround (if any): Redeploying keepalived daemon will populate the keepalived.conf properly and daemon will come online. Result: Keepalived will come online properly.
Story Points: ---
Clone Of: Environment:
Last Closed: 2026-03-04 09:54:13 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 2388233    

Description Manisha Saini 2025-07-15 08:47:02 UTC
Description of problem:
=======================

Ref BZ for more details - https://bugzilla.redhat.com/show_bug.cgi?id=2365590#c10 (For More details)

In a HA NFS-Ganesha deployment with ingress mode HAProxy-Protocol, the keepalived service intermittently fails to start after a node reboot. The issue occurs because the generated /etc/keepalived/keepalived.conf file is incomplete or incorrectly populated.
This prevents keepalived from starting and taking control of the VIP, which may impact NFS failover or availability.




Version-Release number of selected component (if applicable):
=========================

[ceph: root@cali013 /]# rpm -qa | grep nfs
libnfsidmap-2.5.4-34.el9.x86_64
nfs-utils-2.5.4-34.el9.x86_64
nfs-ganesha-selinux-6.5-25.el9cp.noarch
nfs-ganesha-6.5-25.el9cp.x86_64
nfs-ganesha-ceph-6.5-25.el9cp.x86_64
nfs-ganesha-rados-grace-6.5-25.el9cp.x86_64
nfs-ganesha-rados-urls-6.5-25.el9cp.x86_64
nfs-ganesha-rgw-6.5-25.el9cp.x86_64

# ceph --version
ceph version 19.2.1-234.el9cp (abe90cadf4fae6d213cb75dab705ca8e05ac1210) squid (stable)


How reproducible:
=================
Intermittent


Steps to Reproduce:
==================
1. Create NFS Ganesha cluster with haproxy-protocol

# ceph nfs cluster create nfsganesha "1 cali016 cali020" --ingress-mode haproxy-protocol  --ingress --virtual-ip 10.8.130.191/22 -i /var/lib/ceph/kimp.yaml


# ceph nfs cluster info nfsganesha
{
  "nfsganesha": {
    "backend": [
      {
        "hostname": "cali016",
        "ip": "10.8.130.16",
        "port": 12049
      }
    ],
    "ingress_mode": "haproxy-protocol",
    "monitor_port": 9049,
    "port": 2049,
    "virtual_ip": "10.8.130.191"
  }
}


2. Create 2 NFS exports

Export 1
---------
# ceph fs subvolume create cephfs ganesha1 --group_name ganeshagroup --enctag KEY-478ce1f-7373659b-f724-49a7-b936-662ce3dd1129

#  ceph fs subvolume getpath cephfs ganesha1 --group_name ganeshagroup
/volumes/ganeshagroup/ganesha1/69d4c85b-9a78-4f65-a62e-393f11689b7d

# ceph nfs export create cephfs nfsganesha /ganesha1 cephfs --path /volumes/ganeshagroup/ganesha1/69d4c85b-9a78-4f65-a62e-393f11689b7d --kmip_key_id KEY-478ce1f-7373659b-f724-49a7-b936-662ce3dd1129
{
  "bind": "/ganesha1",
  "cluster": "nfsganesha",
  "fs": "cephfs",
  "mode": "RW",
  "path": "/volumes/ganeshagroup/ganesha1/69d4c85b-9a78-4f65-a62e-393f11689b7d"
}

Export 2
----------
# ceph fs subvolume create cephfs ganesha2 --group_name ganeshagroup --enctag KEY-478ce1f-79de1c68-f163-4f9c-bb2f-697bcd1f66a4

#  ceph fs subvolume getpath cephfs ganesha2 --group_name ganeshagroup
/volumes/ganeshagroup/ganesha2/101191bd-14b2-41e0-b45e-acc773bdcba8

# ceph nfs export create cephfs nfsganesha /ganesha2 cephfs --path /volumes/ganeshagroup/ganesha2/101191bd-14b2-41e0-b45e-acc773bdcba8 --kmip_key_id KEY-478ce1f-79de1c68-f163-4f9c-bb2f-697bcd1f66a4
{
  "bind": "/ganesha2",
  "cluster": "nfsganesha",
  "fs": "cephfs",
  "mode": "RW",
  "path": "/volumes/ganeshagroup/ganesha2/101191bd-14b2-41e0-b45e-acc773bdcba8"
}

3. Mount the exports on 2 clients and run IO's

[root@ceph-manisaini-iuqoik-node7 mnt]# mount -t nfs 10.8.130.191:/ganesha1 /mnt/ganesha/
[root@ceph-manisaini-iuqoik-node7 mnt]# cd /mnt/ganesha/

[root@ceph-encrytion-nfs-viovis-node7 mnt]# mount -t nfs 10.8.130.191:/ganesha2 /mnt/ganesha/
[root@ceph-encrytion-nfs-viovis-node7 mnt]# cd /mnt/ganesha/

4. Reboot the active node (cali016). IO's will pause for sometime and then will resume. Ganesha will start on the passive node (cali020).

# ceph orch ps | grep nfsganesha
haproxy.nfs.nfsganesha.cali016.tzrifr     cali016  *:2049,9049       host is offline    92s ago  11m    41.8M        -  2.4.22-f8e3218    4aa9f9e449aa  e97e399496ea
haproxy.nfs.nfsganesha.cali020.gsjnkq     cali020  *:2049,9049       running (69s)       0s ago  74s    39.2M        -  2.4.22-f8e3218    4aa9f9e449aa  212c7da89047
keepalived.nfs.nfsganesha.cali016.kzfewa  cali016                    host is offline    92s ago  11m    1547k        -  2.2.8             38911a18f8ae  2e21ba06fc04
keepalived.nfs.nfsganesha.cali020.csriqp  cali020                    running (66s)       0s ago  73s    1551k        -  2.2.8             38911a18f8ae  b72ca70dcb19
nfs.nfsganesha.0.0.cali016.qujdlk         cali016  *:12049           host is offline    92s ago  11m     532M        -  6.5               88bbac176149  3ffb208f04d7
nfs.nfsganesha.0.1.cali020.ffyhbm         cali020  *:12049           running (74s)       0s ago  74s    98.4M        -  6.5               88bbac176149  d17fc9aa227c
[ceph: root@cali013 /]#


5. Once the cali016 node is up, it was observed that for sometime, the ganesha was running on both the nodes at same time causing I/O error. Then later, ganesha service gets restarted on Node - Cali020. But Ingress was in error state.


# ceph orch ps | grep nfsganesha
haproxy.nfs.nfsganesha.cali016.tzrifr     cali016  *:2049,9049       running (2m)      1s ago  17m    36.8M        -  2.4.22-f8e3218    4aa9f9e449aa  bafecd047c29
keepalived.nfs.nfsganesha.cali016.kzfewa  cali016                    error             1s ago  17m        -        -  <unknown>         <unknown>     <unknown>
nfs.nfsganesha.0.1.cali020.ffyhbm         cali020  *:12049           running (2m)      1s ago   6m    92.2M        -  6.5               88bbac176149  cc120ee08800

6. Linux untars resulted in I/O error

--------------------
tar: linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-everest.dts: Cannot open: Remote I/O error
linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-rainier-1s4u.dts
tar: linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-rainier-1s4u.dts: Cannot open: Remote I/O error
linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-rainier-4u.dts
xz: (stdin): Read error: Remote I/O error
tar: linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-rainier-4u.dts: Cannot open: Remote I/O error
linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-rainier.dts
tar: linux-6.4/arch/arm/boot/dts/aspeed-bmc-ibm-rainier.dts: Cannot open: Remote I/O error
linux-6.4/arch/arm/boot/dts/aspeed-bmc-inspur-fp5280g2.dts
tar: Unexpected EOF in archive
tar: linux-6.4/arch: Cannot utime: Remote I/O error
tar: linux-6.4/arch: Cannot change ownership to uid 0, gid 0: Remote I/O error
tar: linux-6.4/arch: Cannot change mode to rwxrwxr-x: Remote I/O error
tar: linux-6.4: Cannot utime: Remote I/O error
tar: linux-6.4: Cannot change ownership to uid 0, gid 0: Remote I/O error
tar: linux-6.4: Cannot change mode to rwxrwxr-x: Remote I/O error
tar: Error is not recoverable: exiting now
------------------


7. Restart the ingress service, It again goes in error state


[ceph: root@cali013 /]# ceph orch ls
NAME                       PORTS                   RUNNING  REFRESHED  AGE  PLACEMENT
alertmanager               ?:9093,9094                 1/1  0s ago     4w   count:1
ceph-exporter                                          5/5  1s ago     4w   *
crash                                                  5/5  1s ago     4w   *
grafana                    ?:3000                      1/1  0s ago     4w   count:1
ingress.nfs.nfsganesha     10.8.130.191:2049,9049      1/2  0s ago     21m  cali016;cali020;count:1
mds.cephfs                                             2/2  1s ago     4w   count:2
mgr                                                    2/2  0s ago     4w   count:2
mon                                                    5/5  1s ago     4w   count:5
nfs.nfsganesha             ?:12049                     1/1  1s ago     21m  cali016;cali020;count:1
node-exporter              ?:9100                      5/5  1s ago     4w   *
osd.all-available-devices                               27  1s ago     4w   *
prometheus                 ?:9095                      1/1  0s ago     4w   count:1

# ceph orch restart ingress.nfs.nfsganesha
Scheduled to restart haproxy.nfs.nfsganesha.cali016.tzrifr on host 'cali016'
Scheduled to restart keepalived.nfs.nfsganesha.cali016.kzfewa on host 'cali016'

# ceph orch ps | grep nfsganesha
haproxy.nfs.nfsganesha.cali016.tzrifr     cali016  *:2049,9049       running (8s)      2s ago  22m    35.7M        -  2.4.22-f8e3218    4aa9f9e449aa  87767b4b723f
keepalived.nfs.nfsganesha.cali016.kzfewa  cali016                    unknown           2s ago  22m        -        -  <unknown>         <unknown>     <unknown>
nfs.nfsganesha.0.1.cali020.ffyhbm         cali020  *:12049           running (7m)      0s ago  11m     117M        -  6.5               88bbac176149  cc120ee08800

[ceph: root@cali013 /]# ceph orch ps | grep nfsganesha
haproxy.nfs.nfsganesha.cali016.tzrifr     cali016  *:2049,9049       running (13s)     0s ago  22m    35.8M        -  2.4.22-f8e3218    4aa9f9e449aa  87767b4b723f
keepalived.nfs.nfsganesha.cali016.kzfewa  cali016                    running (0s)      0s ago  22m        -        -  <unknown>         38911a18f8ae  a94743684b94
nfs.nfsganesha.0.1.cali020.ffyhbm         cali020  *:12049           running (7m)      0s ago  11m     117M        -  6.5               88bbac176149  cc120ee08800


# ceph orch ps | grep nfsganesha
haproxy.nfs.nfsganesha.cali016.tzrifr     cali016  *:2049,9049       running (69s)     1s ago  23m    37.7M        -  2.4.22-f8e3218    4aa9f9e449aa  87767b4b723f
keepalived.nfs.nfsganesha.cali016.kzfewa  cali016                    error             1s ago  23m        -        -  <unknown>         <unknown>     <unknown>
nfs.nfsganesha.0.1.cali020.ffyhbm         cali020  *:12049           running (8m)      1s ago  12m     117M        -  6.5               88bbac176149  cc120ee08800


8. "ls" operation are stuck on NFS mount point at this stage
------------------------------------------------------
[root@ceph-manisaini-iuqoik-node7 ganesha]# ls



^C
[root@ceph-manisaini-iuqoik-node7 ganesha]# ^C
[root@ceph-manisaini-iuqoik-node7 ganesha]# ^C
[root@ceph-manisaini-iuqoik-node7 ganesha]# ls






^C
[root@ceph-manisaini-iuqoik-node7 ganesha]# ^C
[root@ceph-manisaini-iuqoik-node7 ganesha]#


Actual results:
================
keepalived fails to start on the rebooted node resulted in I/O on the client and all operations were stuck on NFS mount point


Expected results:
=================
keepalived should start cleanly with a valid keepalived.conf.The VIP should be bound to one of the available nodes.NFS IO should remain uninterrupted due to proper VIP failover.


Additional info:

Comment 9 Red Hat Bugzilla 2026-03-04 09:54:13 UTC
This product has been discontinued or is no longer tracked in Red Hat Bugzilla.