Bug 2066778
| Summary: | RHEL9: [virtiofs] Guest hang after "getfattr -m '' /mnt/Makefile" | ||
|---|---|---|---|
| Product: | Red Hat Enterprise Linux 9 | Reporter: | bfu <bfu> |
| Component: | virtiofsd | Assignee: | Vivek Goyal <vgoyal> |
| Status: | CLOSED ERRATA | QA Contact: | xiagao |
| Severity: | medium | Docs Contact: | |
| Priority: | high | ||
| Version: | 9.0 | CC: | cohuck, coli, dgilbert, dhorak, gmaglione, hannsj_uhl, jinzhao, juzhang, kkiwi, knoel, lijin, ngu, pbonzini, qizhu, ribarry, slopezpa, smitterl, stefanha, thuth, vgoyal, virt-maint, virt-qe-z, xiagao, yfu, yiwei, zhenyzha |
| Target Milestone: | rc | Keywords: | Regression, Reopened, Triaged |
| Target Release: | 9.1 | Flags: | pm-rhel:
mirror+
|
| Hardware: | All | ||
| OS: | Linux | ||
| Whiteboard: | |||
| Fixed In Version: | virtiofsd-1.3.0-1.el9 | Doc Type: | If docs needed, set a value |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2022-11-15 10:30:35 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | 2077854 | ||
| Bug Blocks: | |||
this was reproducible on ARM, but passed on x86 (In reply to bfu from comment #1) > this was reproducible on ARM, but passed on x86 Is there anything in /var/log/audit/audit.log which indicates any kind of seccomp failure. (In reply to bfu from comment #0) > Description of problem: > after dropping cap_sys_admin, virtiofsd could not work > > Version-Release number of selected component (if applicable): > qemu version: 6.2.0-11.el9 > kernel version: 5.14.0-70.1.1.el9.s390x > virtiofs version:virtiofsd-1.1.0-3.el9.s390x > > How reproducible: > 100% > > Steps to Reproduce: > 1. Create a shared directory for testing on the host. > # mkdir -p /tmp/virtiofs_test > 2. Drop CAP_SYS_ADMIN > # capsh --drop=cap_sys_admin -- > # capsh --print (confirm cap_sys_admin dropped successfully ) > 3. Run the virtiofsd daemon on the host > (qemu virtiofsd) > # /usr/libexec/virtiofsd --socket-path=/tmp/vhostqemu -o > source=/tmp/virtiofs_test -o ache=auto > > Actual results: > [2022-03-22T13:19:29Z ERROR virtiofsd] Error entering sandbox: Unshare(Os { > code: 1, kind: PermissionDenied, message: "Operation not permitted" }) > > Expected results: > [2022-03-22T12:34:04Z INFO virtiofsd] Waiting for vhost-user socket > connection... > We need CAP_SYS_ADMIN in the user namespace virtiofs is running because unshare() and pivot_root() etc need it. And that's why it is failing. Now C virtiofsd will also fail on unsahre() call little later after the connection. I noticed that C version first waits for the connection from qemu and then it sets up the sandbox. While rust version first sets up the sandbox and then waits for the connection. That's why rust version is failing before receiving the connection. C version will fail after the connection. Give it a try. In short, this is not a bug. We are running privileged (started by root) and we need CAP_SYS_ADMIN in this mode. Also, I ran it on x86_64 and it fails for me without CAP_SYS_ADMIN. Please re-run it. I think you are considering "Waiting for connection" message as success with C virtiofsd. It is just an intermediate step. Let qemu connect and see if virtiofsd still runs after that or not. Yes, it's not a bug from https://bugzilla.redhat.com/show_bug.cgi?id=1860491#c24. yes, If sandbox=chroot is used, then CAP_SYS_ADMIN might not be required. By default sandbox=namespace is used which requires CAP_SYS_ADMIN. *** This bug has been marked as a duplicate of bug 1860491 *** Yes,x86 also has this issue after enable **selinux**. # getenforce Enforcing The steps are the same with comment 18. /usr/libexec/virtiofsd --socket-path=/tmp/avocado_02s5rvmd/avocado-vt-vm1-viofs-virtiofsd.sock -o source=/root/avocado/data/avocado-vt/virtio_fs_test/ -o cache=none -o sandbox=chroot -o xattr -o xattrmap=':map::user.virtiofsd.:' 2022-03-29 13:19:19: [root@dell-per440-06 kar]# [2022-03-29T17:19:19Z INFO virtiofsd] Waiting for vhost-user socket connection... 2022-03-29 13:19:22: [2022-03-29T17:19:22Z INFO virtiofsd] Client connected, servicing requests 2022-03-29 13:20:08: thread 'vring_worker' panicked at 'unterminated mapping: Error { cause: UnterminatedMapping, rule: None }', src/passthrough/mod.rs:2092:60 2022-03-29 13:20:08: note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace 2022-03-29 13:23:12: thread 'virtiofsd-backend' panicked at 'called `Result::unwrap()` on an `Err` value: PoisonError { .. }', /builddir/build/BUILD/virtiofsd-1.1.0/vendor/vhost-user-backend/src/vring.rs:242:27 2022-03-29 13:23:12: [2022-03-29T17:23:12Z ERROR virtiofsd] Waiting for daemon failed: WaitDaemon(Any { .. }) 2022-03-29 13:23:12: thread 'main' panicked at 'called `Result::unwrap()` on an `Err` value: PoisonError { .. }', src/main.rs:851:10 2022-03-29 13:23:12: [2022-03-29T17:23:12Z ERROR vhost_user_backend::handler] Error in vring worker: Any { .. } pkg: kernel-5.14.0-69.el9.x86_64 qemu-kvm-6.2.0-12.el9.x86_64 virtiofsd-1.1.0-4.el9_0.x86_64 (In reply to xiagao from comment #21) > Yes,x86 also has this issue after enable **selinux**. As the selinux is disabled at the beginning, so the test results always passed. > > # getenforce > Enforcing > > The steps are the same with comment 18. > /usr/libexec/virtiofsd > --socket-path=/tmp/avocado_02s5rvmd/avocado-vt-vm1-viofs-virtiofsd.sock -o > source=/root/avocado/data/avocado-vt/virtio_fs_test/ -o cache=none -o > sandbox=chroot -o xattr -o xattrmap=':map::user.virtiofsd.:' > > 2022-03-29 13:19:19: [root@dell-per440-06 kar]# [2022-03-29T17:19:19Z INFO > virtiofsd] Waiting for vhost-user socket connection... > 2022-03-29 13:19:22: [2022-03-29T17:19:22Z INFO virtiofsd] Client > connected, servicing requests > 2022-03-29 13:20:08: thread 'vring_worker' panicked at 'unterminated > mapping: Error { cause: UnterminatedMapping, rule: None }', > src/passthrough/mod.rs:2092:60 > 2022-03-29 13:20:08: note: run with `RUST_BACKTRACE=1` environment variable > to display a backtrace > 2022-03-29 13:23:12: thread 'virtiofsd-backend' panicked at 'called > `Result::unwrap()` on an `Err` value: PoisonError { .. }', > /builddir/build/BUILD/virtiofsd-1.1.0/vendor/vhost-user-backend/src/vring.rs: > 242:27 > 2022-03-29 13:23:12: [2022-03-29T17:23:12Z ERROR virtiofsd] Waiting for > daemon failed: WaitDaemon(Any { .. }) > 2022-03-29 13:23:12: thread 'main' panicked at 'called `Result::unwrap()` on > an `Err` value: PoisonError { .. }', src/main.rs:851:10 > 2022-03-29 13:23:12: [2022-03-29T17:23:12Z ERROR > vhost_user_backend::handler] Error in vring worker: Any { .. } > > pkg: > kernel-5.14.0-69.el9.x86_64 > qemu-kvm-6.2.0-12.el9.x86_64 > virtiofsd-1.1.0-4.el9_0.x86_64 And list the results. - enable host SELinux + "-o xattrmap=':map::user.virtiofsd.:" --> trigger this error. - disable host SELinux + "-o xattrmap=':map::user.virtiofsd.:" --> no error So, it's kind of related with SELinux. And without "-o xattrmap=':map::user.virtiofsd.:" did get rid of this error, but some scenario need it such as https://bugzilla.redhat.com/show_bug.cgi?id=1860491#c24 (In reply to xiagao from comment #26) > And list the results. > - enable host SELinux + "-o xattrmap=':map::user.virtiofsd.:" --> trigger > this error. > - disable host SELinux + "-o xattrmap=':map::user.virtiofsd.:" --> no error Can you look into /var/log/audit/audit.log and see if there are any SELinux denials (with SELinux enabled). I am still not sure what's the connection with SELinux here. > > So, it's kind of related with SELinux. > And without "-o xattrmap=':map::user.virtiofsd.:" did get rid of this error, > but some scenario need it such as > https://bugzilla.redhat.com/show_bug.cgi?id=1860491#c24 That's a very specific use case. If somebody is running without CAP_SYS_ADMIN and want to set "trusted." xattrs, then remapping of xattrs is required. First of all is somebody doing that? Even if somebody is doing that, I think there is workaround. For example, use following rule to prefix all xattrs with user.virtiofs. -o xattrmap=":prefix:all::user.virtiofs.::bad:all:::" Give this workaround a try. If this workaround works, I think this is not an exception for two reasons. - Dropping CAP_SYS_ADMIN and then trying to set "trusted.*" xattrs is probably not very common. - Even if it happens, we have a workaround. (In reply to Vivek Goyal from comment #27) > (In reply to xiagao from comment #26) > > And list the results. > > - enable host SELinux + "-o xattrmap=':map::user.virtiofsd.:" --> trigger > > this error. > > - disable host SELinux + "-o xattrmap=':map::user.virtiofsd.:" --> no error > > Can you look into /var/log/audit/audit.log and see if there are any SELinux > denials (with SELinux enabled). I am still not sure what's the connection > with SELinux here. No any denials logs in host audit log. > > > > > So, it's kind of related with SELinux. > > And without "-o xattrmap=':map::user.virtiofsd.:" did get rid of this error, > > but some scenario need it such as > > https://bugzilla.redhat.com/show_bug.cgi?id=1860491#c24 > > That's a very specific use case. If somebody is running without > CAP_SYS_ADMIN and want to set "trusted." xattrs, then remapping of xattrs is > required. As far as I know, it's not a common use case from x86 side. Zhenyu, bfu, from multi-arch side, what's your opinion? > > First of all is somebody doing that? Even if somebody is doing that, I think > there is workaround. For example, use following rule to prefix all xattrs > with user.virtiofs. > > > -o xattrmap=":prefix:all::user.virtiofs.::bad:all:::" Yes, this workaround works. > > Give this workaround a try. > > If this workaround works, I think this is not an exception for two reasons. > > - Dropping CAP_SYS_ADMIN and then trying to set "trusted.*" xattrs is > probably not very common. > - Even if it happens, we have a workaround. Totally agree. Hi Sergio, could you help move this bz to MODIFIED/ON_QA status? Thanks, Xiaoling Extend ITM as the original ITM is close and the bug status still is MODIFIED. Vivek hi, Could you help to change status to ON-QA, I think the pkg is ready. BR, Xiaoling So who is supposed to move it to ON_QA status. I guess once this bug is added to errata, it will automatically move to ON_QA. May be this bz was not added to errata filed for virtiofsd rebase and that's why. https://errata.devel.redhat.com/advisory/97746 Hmm..., so what are the options now? Sergio, any ideas? The test has been PASS on s390x [root@l42 ~]# rpm -qa | grep virtiofsd virtiofsd-1.3.0-1.el9.s390x Test result: (1/4) Host_RHEL.m9.u1.nographic.qcow2.virtio_scsi.up.virtio_net.Guest.RHEL.9.1.0.s390x.io-github-autotest-qemu.unattended_install.cdrom.extra_cdrom_ks.default_install.aio_threads.s390-virtio: PASS (376.10 s) (2/4) Host_RHEL.m9.u1.nographic.qcow2.virtio_scsi.up.virtio_net.Guest.RHEL.9.1.0.s390x.io-github-autotest-qemu.virtio_fs_set_capability.remove_capability.cap_dac_read_search.with_cache.auto.s390-virtio: PASS (27.71 s) (3/4) Host_RHEL.m9.u1.nographic.qcow2.virtio_scsi.up.virtio_net.Guest.RHEL.9.1.0.s390x.io-github-autotest-qemu.virtio_fs_set_capability.remove_capability.cap_sys_admin.with_xattr.with_cache.none.s390-virtio: PASS (28.30 s) (4/4) Host_RHEL.m9.u1.nographic.qcow2.virtio_scsi.up.virtio_net.Guest.RHEL.9.1.0.s390x.io-github-autotest-qemu.virtio_fs_set_capability.remove_capability.cap_sys_admin.without_xattr.with_cache.none.s390-virtio: PASS (28.21 s) As the test result on both s390x and arm, set this bz to verified Just clearing the needinfo as it seem addressed now. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (virtiofsd bug fix and enhancement update), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2022:8166 |
Description of problem: Guest hang after "getfattr -m '' /mnt/Makefile" Version-Release number of selected component (if applicable): qemu version: 6.2.0-11.el9 kernel version: 5.14.0-70.1.1.el9.s390x virtiofs version:virtiofsd-1.1.0-3.el9.s390x How reproducible: 100% Steps to Reproduce: 1. active hpage [root@l42 kar]# rmmod kvm [root@l42 kar]# modprobe kvm hpage=1 [root@l42 kar]# echo 4224 > /proc/sys/vm/nr_hugepages [root@l42 kar]# mount -t hugetlbfs -o pagesize=1024K none /mnt/kvm_hugepage [root@l42 kar]# cd /usr/local/lib/python3.9/site-packages/avocado_framework_plugin_vt-94.0-py3.9.egg/virttest; echo 3 > /proc/sys/vm/drop_caches 2. Create a shared directory for testing on the host. # mkdir -p /tmp/virtiofs_test 3. Drop CAP_SYS_ADMIN [root@ampere-hr350a-01 ~]# capsh --drop=cap_sys_admin -- [root@ampere-hr350a-01 ~]# capsh --print | grep cap_sys_admin Current: =ep cap_sys_admin-ep Current IAB: !cap_sys_admin 4. Run the virtiofsd daemon on the host [root@l42 virttest]# /usr/libexec/virtiofsd --socket-path=/tmp/vhostqemu -o source=/root/avocado/data/avocado-vt/virtio_fs_test/ -o cache=none -o sandbox=chroot -o xattr -o xattrmap=':map::user.virtiofsd.:' 5. boot up guest with virtio-fs devices /usr/libexec/qemu-kvm \ -name 'avocado-vt-vm1' \ -sandbox on \ -machine s390-ccw-virtio,memory-backend=mem-machine_mem \ -nodefaults \ -vga none \ -m 4096 \ -object memory-backend-file,mem-path=/mnt/kvm_hugepage,size=4G,id=mem-machine_mem \ -smp 6,maxcpus=6,cores=3,threads=1,sockets=2 \ -cpu 'host' \ -chardev socket,id=chardev_serial0,server=on,path=/tmp/bfu,wait=off \ -device sclpconsole,id=serial0,chardev=chardev_serial0 \ -device virtio-scsi-ccw,id=virtio_scsi_ccw0 \ -blockdev node-name=file_image1,driver=file,aio=threads,filename=/home/kar/vt_test_images/rhel900-s390x-virtio-scsi.qcow2,cache.direct=on,cache.no-flush=off \ -blockdev node-name=drive_image1,driver=qcow2,cache.direct=on,cache.no-flush=off,file=file_image1 \ -device scsi-hd,id=image1,drive=drive_image1,write-cache=on \ -chardev socket,id=char_virtiofs_fs,path=/tmp/vhostqemu \ -device vhost-user-fs-ccw,id=vufs_virtiofs_fs,chardev=char_virtiofs_fs,tag=myfs,queue-size=1024 \ -device virtio-net-ccw,mac=9a:9f:19:fe:49:13,id=idh8Nq82,netdev=idVx9QzC \ -netdev tap,id=idVx9QzC,vhost=on \ -nographic \ -rtc base=utc \ -boot strict=on \ -enable-kvm \ -device virtio-mouse-ccw,id=input_mouse1 \ -device virtio-keyboard-ccw,id=input_keyboard1 \ -monitor stdio 6. mount virtio-fs in guest [root@localhost ~]# mount -t virtiofs myfs /mnt/myfs mount -t virtiofs myfs /mnt/myfs [root@localhost ~]# df -h df -h Filesystem Size Used Avail Use% Mounted on devtmpfs 1.9G 0 1.9G 0% /dev tmpfs 1.9G 0 1.9G 0% /dev/shm tmpfs 749M 17M 733M 3% /run /dev/mapper/rhel-root 17G 3.9G 14G 23% / /dev/sda1 1014M 173M 842M 18% /boot tmpfs 375M 32K 375M 1% /run/user/0 myfs 70G 7.2G 63G 11% /mnt/myfs [root@localhost ~]# setfattr -n trusted.test /mnt/myfs setfattr -n trusted.test /mnt/myfs [root@localhost ~]# getfattr -m '' /mnt/myfs getfattr -m '' /mnt/myfs Actual results: Guest:guest hang Host:[root@l42 virttest]# /usr/libexec/virtiofsd --socket-path=/tmp/avocado-vt-vm1-viofs-virtiofsd.sock -o source=/root/avocado/data/avocado-vt/virtio_fs_test/ -o cache=none -o sandbox=chroot -o xattr -o xattrmap=':map::user.virtiofsd.:' [2022-03-28T10:31:37Z INFO virtiofsd] Waiting for vhost-user socket connection... [2022-03-28T10:31:37Z INFO virtiofsd] Client connected, servicing requests thread 'vring_worker' panicked at 'unterminated mapping: Error { cause: UnterminatedMapping, rule: None }', src/passthrough/mod.rs:2089:60 note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace Expected results: Guest: [root@L1 ~]# getfattr -m '' /mnt/Makefile getfattr: Removing leading '/' from absolute path names # file: mnt/Makefile trusted.test Host: [root@virtlab720 home]# getfattr /tmp/virtiofs_test/Makefile # file: tmp/virtiofs_test/Makefile user.virtiofsd.trusted.test Additional info: