Fedora Account System
Red Hat Associate
Red Hat Customer
In openQA we have a 'support server' test which runs an NFS server (among other things). This test is always set to run on the most-recent stable Fedora, so it recently moved to Fedora 41. Since then, I've noticed it occasionally fails, because starting nfs-server.service fails. It's failing because rpc.mountd is crashing. Here's the backtrace (in-line, as it's short); Program terminated with signal SIGABRT, Aborted. Downloading source file /usr/src/debug/glibc-2.40-10.fc41.x86_64/nptl/pthread_sigmask.c #0 __GI___pthread_sigmask (how=how@entry=2, newmask=<optimized out>, newmask@entry=0x7fff0220fa60, oldmask=oldmask@entry=0x0) at pthread_sigmask.c:43 43 return (INTERNAL_SYSCALL_ERROR_P (result) Thread 1 (Thread 0x7fb1bd955880 (LWP 1005)): #0 __GI___pthread_sigmask (how=how@entry=2, newmask=<optimized out>, newmask@entry=0x7fff0220fa60, oldmask=oldmask@entry=0x0) at pthread_sigmask.c:43 local_newmask = {__val = {140401372702112, 94154476464104, 140733229103296, 140401372576782, 140733229103344, 5774367749484711168, 140733229103956, 140733229103504, 140733229103344, 140401372582914, 68, 5774367749484711168, 94154476464104, 140733229103504, 140733229103376, 140401372557136}} result = 0 #1 0x00007fb1bdf011b4 in clnt_vc_call (cl=0x55a20c1fa750, proc=2, xdr_args=0x7fb1bdf05e70 <xdr_rpcb>, args_ptr=0x7fff0220fb50, xdr_results=0x7fb1bdf02f00 <xdr_bool>, results_ptr=0x7fff0220fb4c, timeout=...) at /usr/src/debug/libtirpc-1.3.6-0.fc41.x86_64/src/clnt_vc.c:462 ct = 0x55a20c1fa780 xdrs = 0x55a20c1fa7e8 reply_msg = {rm_xid = 35715952, rm_direction = (REPLY | unknown: 0x7ffe), ru = {RM_cmb = {cb_rpcvers = 35715953, cb_prog = 32767, cb_vers = 0, cb_proc = 0, cb_cred = {oa_flavor = 0, oa_base = 0x0, oa_length = 203400544}, cb_verf = {oa_flavor = 0, oa_base = 0x7fb1bdf02cc0 <xdr_void> "\363\017\036\372\270\001", oa_length = 35715888}}, RM_rmb = {rp_stat = (MSG_DENIED | unknown: 0x220fb70), ru = {RP_ar = {ar_verf = {oa_flavor = 0, oa_base = 0x0, oa_length = 0}, ar_stat = 203400544, ru = {AR_versions = {low = 0, high = 0}, AR_results = {where = 0x0, proc = 0x7fb1bdf02cc0 <xdr_void>}}}, RP_dr = {rj_stat = RPC_MISMATCH, ru = {RJ_versions = {low = 0, high = 0}, RJ_why = AUTH_OK}}}}}} x_id = 1679306803 msg_x_id = 0x55a20c1fa7cc shipnow = 1 refreshes = 2 mask = {__val = {0, 140733229103936, 140733229103744, 5774367749484711168, 3417798300496035841, 3342918222317056114, 1801678707, 0, 0, 0, 0, 0, 0, 0, 0, 0}} newmask = {__val = {18446744067267100671, 140733229103728, 18446744067267100671, 0, 5, 7166464161511339311, 7957688117742693935, 5774367749484711168, 140401366642640, 94154476463952, 0, 5774367749484711168, 140733229103744, 94154476463952, 140733229103920, 140401371331564}} call_again = <optimized out> __PRETTY_FUNCTION__ = <optimized out> #2 0x00007fb1bdefe876 in rpcb_unset (program=program@entry=100005, version=version@entry=2, nconf=nconf@entry=0x0) at /usr/src/debug/libtirpc-1.3.6-0.fc41.x86_64/src/rpcb_clnt.c:746 client = 0x55a20c1fa750 rslt = 0 parms = {r_prog = 100005, r_vers = 2, r_netid = 0x7fb1bdf18db0 <CSWTCH.43> "\002", r_addr = 0x7fb1bdf18db0 <CSWTCH.43> "\002", r_owner = 0x7fff0220fb70 "0"} uidbuf = "0\000\233\371\241U\000\000\350՚\371\241U\000\000@\374 \002\377\177\000\000\240\350\233\371\241U\000" #3 0x000055a1f99ae09d in nfs_svc_unregister.constprop.0 (version=2, program=100005) at ../../support/nfs/svc_create.c:483 No locals. #4 0x000055a1f999ed48 in unregister_services () at /usr/src/debug/nfs-utils-2.8.1-0.fc41.x86_64/utils/mountd/mountd.c:110 No locals. #5 0x000055a1f999e242 in main (argc=1, argv=<optimized out>) at /usr/src/debug/nfs-utils-2.8.1-0.fc41.x86_64/utils/mountd/mountd.c:798 progname = 0x7fff02211ea2 "rpc.mountd" listeners = 0 foreground = <optimized out> c = <optimized out> ttl = <optimized out> sa = {__sigaction_handler = {sa_handler = 0x1, sa_sigaction = 0x1}, sa_mask = {__val = {0, 140733229104288, 140401370804290, 255, 4096, 4138, 16, 140401367441911, 140733229104872, 140733229104320, 140401370831919, 140401367441911, 0, 140733229104592, 140401367327532, 4096}}, sa_flags = 0, sa_restorer = 0x0} rlim = {rlim_cur = 1024, rlim_max = 524288} Here are the journal messages: Nov 11 23:58:16 support.test.openqa.fedoraproject.org systemd-coredump[1069]: [🡕] Process 1005 (rpc.mountd) of user 0 dumped core. Module libpcre2-8.so.0 from rpm pcre2-10.44-1.fc41.1.x86_64 Module libselinux.so.1 from rpm libselinux-3.7-5.fc41.x86_64 Module libcrypto.so.3 from rpm openssl-3.2.2-9.fc41.x86_64 Module libkeyutils.so.1 from rpm keyutils-1.6.3-4.fc41.x86_64 Module libkrb5support.so.0 from rpm krb5-1.21.3-3.fc41.x86_64 Module libcom_err.so.2 from rpm e2fsprogs-1.47.1-6.fc41.x86_64 Module libk5crypto.so.3 from rpm krb5-1.21.3-3.fc41.x86_64 Module libkrb5.so.3 from rpm krb5-1.21.3-3.fc41.x86_64 Module libgssapi_krb5.so.2 from rpm krb5-1.21.3-3.fc41.x86_64 Module liblzma.so.5 from rpm xz-5.6.2-2.fc41.x86_64 Module libz.so.1 from rpm zlib-ng-2.1.7-3.fc41.x86_64 Module libtirpc.so.3 from rpm libtirpc-1.3.6-0.rc1.fc41.x86_64 Module libuuid.so.1 from rpm util-linux-2.40.2-4.fc41.x86_64 Module libblkid.so.1 from rpm util-linux-2.40.2-4.fc41.x86_64 Module libxml2.so.2 from rpm libxml2-2.12.8-2.fc41.x86_64 Module rpc.mountd from rpm nfs-utils-2.8.1-0.fc41.x86_64 Stack trace of thread 1005: #0 0x00007fb1bdd499b8 pthread_sigmask.5 (libc.so.6 + 0x779b8) #1 0x00007fb1bdf011b4 clnt_vc_call (libtirpc.so.3 + 0xe1b4) #2 0x00007fb1bdefe876 rpcb_unset (libtirpc.so.3 + 0xb876) #3 0x000055a1f99ae09d nfs_svc_unregister.constprop.0 (rpc.mountd + 0x1309d) #4 0x000055a1f999ed48 unregister_services (rpc.mountd + 0x3d48) #5 0x000055a1f999e242 main (rpc.mountd + 0x3242) #6 0x00007fb1bdcd5248 __libc_start_call_main (libc.so.6 + 0x3248) #7 0x00007fb1bdcd530b __libc_start_main@@GLIBC_2.34 (libc.so.6 + 0x330b) #8 0x000055a1f999ebf5 _start (rpc.mountd + 0x3bf5) ELF object binary architecture: AMD x86-64 Nov 11 23:58:16 support.test.openqa.fedoraproject.org systemd[1]: nfs-mountd.service: Control process exited, code=dumped, status=6/ABRT Nov 11 23:58:16 support.test.openqa.fedoraproject.org audit[1]: SERVICE_START pid=1 uid=0 auid=4294967295 ses=4294967295 subj=system_u:system_r:init_t:s0 msg='unit=nfs-mountd comm="systemd> Nov 11 23:58:16 support.test.openqa.fedoraproject.org audit[1]: SERVICE_STOP pid=1 uid=0 auid=4294967295 ses=4294967295 subj=system_u:system_r:init_t:s0 msg='unit=systemd-coredump@0-1068-0> Nov 11 23:58:16 support.test.openqa.fedoraproject.org systemd[1]: nfs-mountd.service: Failed with result 'timeout'. Nov 11 23:58:16 support.test.openqa.fedoraproject.org systemd[1]: Failed to start nfs-mountd.service - NFS Mount Daemon. Nov 11 23:58:16 support.test.openqa.fedoraproject.org systemd[1]: Dependency failed for nfs-server.service - NFS server and services. Nov 11 23:58:16 support.test.openqa.fedoraproject.org systemd[1]: nfs-server.service: Job nfs-server.service/start failed with result 'dependency'. The test creates /export , /repo and /iso directories, puts some stuff in them, writes an /etc/exports file with this content: /export 172.16.2.0/24(ro) /repo 172.16.2.0/24(ro) /iso 172.16.2.0/24(ro) and then does `firewall-cmd --add-service=nfs` and `systemctl restart nfs-server.service`.
Huh, now I look into this carefully, it's actually failing repeatedly on a single update - https://bodhi.fedoraproject.org/updates/FEDORA-2024-eb19549211 , a libtirpc update for F41. Strangely, it's not happening on the equivalent update for F40 or Rawhide, it's *only* happening in F41. I don't know why.
With libtirpc-1.3.6-1.fc41 things should be back up and running...
Yes, that update has passed testing. Since the previous one never made it out of u-t, we can call this resolved now. Thanks!