Bug 1780400
| Summary: | [machines]requests are very slow when enabling libvirt polkit | ||
|---|---|---|---|
| Product: | Red Hat Enterprise Linux 7 | Reporter: | Ravindra Patil <ravpatil> |
| Component: | cockpit | Assignee: | Katerina Koukiou <kkoukiou> |
| Status: | CLOSED ERRATA | QA Contact: | Xianghua Chen <xchen> |
| Severity: | unspecified | Docs Contact: | |
| Priority: | medium | ||
| Version: | 7.7 | CC: | kkoukiou, laurent, mpitt, phrdina, sbalasub, xchen, yafu, ymao, yunyang |
| Target Milestone: | rc | Keywords: | Extras |
| Target Release: | 7.9 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2020-09-30 07:45:20 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Ravindra Patil
2019-12-05 21:17:01 UTC
I tried on RHEL7.7 and RHEL 8.2 with following version, but failed to reproduce this bug: RHEL7.7: cockpit-machines-195-1.el7.x86_64 libvirt-4.5.0-23.el7.x86_64 RHEL8.2: cockpit-machines-208-1.el8.noarch libvirt-4.5.0-35.module+el8.1.0+4227+b2722cb3.x86_64 Steps: 1. Login the cockpit web console with root, prepare a VM. 2. Create user sysadmin and add it to libvirt and libvirtdbus groups: # useradd -G libvirt,libvirtdbus sysadmin 3. Login the cockpit web console use sysadmin, can see and manage the VM (start/shutdown...). 4. Change configuration in libvirt.conf: # vi /etc/libvirt/libvirtd.conf access_drivers = [ "polkit" ] 5. Restart libvirtd: # systemctl restart libvirtd 6. Refresh the cockpit page, try to start/shutdown the vm, it can not manage the vm and there will be an error message"access denied:'QEMU' denied access" . It take less than 10s to see the VM in step 6, I think it's acceptable. And the result can display immediately after the 'virsh list --all' command. Please help to check the following steps and correct me if I missed some steps, thanks! Hello Chen The issue reported in intermittent. The problem becomes more apparent when there are more virtual machines on the host (in our case it's up to 22 vms on one host), what more Cockpit/virsh behavior is inconsistent, sometimes it displays VMs almost immediately, sometimes it can take up to two minutes... Is there any other way to configure cockpit so that non-root user could see installed VMs, interact with their consoles but could not manage (start/stop etc) other than polkit in libvirtd.conf ? Regards Ravindra Hi, I tried to create more than 10 vms on the host, and keep polkit configure as you described, it's indeed slower when using sysadmin to login than root. The vms is listed one by one, but only take about 10+s maybe, not so long as two minutes as you encountered. For the second question , I'm not quite sure , may be dev can give some advice? Sorry for late reply, I'm back from PTO now.
> Is there any other way to configure cockpit so that non-root user could see installed VMs, interact with their consoles but could not manage (start/stop etc) other than polkit in libvirtd.conf ?
No, there is not. Cockpit uses exactly the same system APIs as command line clients, and runs exactly the same kind of login session as e. g. ssh or gdm. It's not technically possible (nor desired) to have different privileges there. Of course in principle the UI could hide some elements, but through .e. g. the JavaScript console you always have access to everything.
As this is reproducible with virsh, it seems that this should be reassigned to libvirt?
After several days of investigation I have figured out where the issue is.
TL;DR:
Tt's in the way how cockpit uses DBus where on every call cockpit runs introspection which is slow.
Now the full summary. libvirt-dbus is written in a way that VMs, StoragePools and basically
all libvirt object are listed dynamically which means that it has to talk to libvirt every time
any object is requested. In case polkit is enabled for every single VM, StoragePool or any other
object libvirt has to talk to polkit to figure out if it should allow access or not.
The way how DBus works in general is that if you introspect some object it will give you an XML
with all interfaces and also its child nodes. So in case of libvirt dbus if you run
`gdbus introspect --system --dest org.libvirt --object-path /org/libvirt/QEMU`
you will get a list of all interfaces and also the list of nodes. Now if you run the same command
for a different object path:
`gdbus introspect --system --dest org.libvirt --object-path /org/libvirt/QEMU/domain`
That will give you all system VMs in the reply and if libvirt has polkit enabled it has to check
for every VM in that list if you can see it or not and that takes some time.
Now the issue that cockpit is hitting is that for every single GetXMLDesc on a VM object it will
run the introspection as well which means that for every single GetXMLDesc call libvirt has to
do number_of_all_vms + the API call checks with polkit.
There is no way how to change this behavior in libvirt-dbus as we have to report the nodes or in
libvirt as we have to check the access. However, there is a possible solution for cockpit itself
where you can change the way how you use DBus.
In addition to support this solution I've created a simple python script to demonstrate the solution:
-----------------------------------------------------------------------------------------------------
import dbus
introspect = False
bus = dbus.SystemBus()
connobj = bus.get_object('org.libvirt', '/org/libvirt/QEMU', introspect=introspect)
conn = dbus.Interface(connobj, 'org.libvirt.Connect')
for dompath in conn.ListDomains(dbus.types.UInt32(0)):
domobj = bus.get_object('org.libvirt', dompath, introspect=introspect)
dom = dbus.Interface(domobj, 'org.libvirt.Domain')
print(dom.GetXMLDesc(dbus.types.UInt32(0)))
-----------------------------------------------------------------------------------------------------
This scripts takes around 3 to 4 seconds to run on my host with polkit enabled and 65 VMs, but if
I change the `introspect` variable to True (which is a default BTW) it takes more 1 minut 10 seconds.
I was looking into cockpit documentations and there is probably a similar solution as well. From
the documentation [1] the `client.call()` method has an optional parametr where you can provide
signature of a method and in that case the introspection is skipped. The current code uses it
this way:
const TIMEOUT = { timeout: 30000 };
connection.call(vmPath, 'org.libvirt.Domain', 'GetXMLDesc', [0], TIMEOUT);
But if you change it to something like this:
const opts = { timeout: 30000, type: 'u' };
connection.call(vmPath, 'org.libvirt.Domain', 'GetXMLDesc', [0], opts);
You can save a lot o time as there will be no introspection.
The documentation states that the introspection is cached, which is nice, but every single call will
be always slow so even with the caching loading initial data will be slow.
Based on all of this I'm moving the BZ back to cockpit to figure out what to do next.
[1] <https://cockpit-project.org/guide/207/cockpit-dbus.html>
Oh wow, thanks Pavel, you're a genius! We indeed don't need the introspection. Sorry, I was totally not aware of that.. Katerina still has trouble with reproducing these absurdly slow times, but we now applied explict D-Bus signatures to all calls to libvirt-dbus. This is available in upstream release 214.1 (just releasing to Fedora 31/32/rawhide and COPR), and I backported the fix here: https://github.com/cockpit-project/cockpit/pull/13707 . Once that lands, I'll make a COPR package for RHEL 7. Bouncing this to Katerina, can you please have a look? Thank you! > Do you need me to file a new bug for the "vms can not be listed" issue? I'm not sure yet what to do about it on the cockpit side. I am a rather strong believer of reasonable timeouts, i. e. if the UI does not get a reply from libvirt after 30 seconds, then it gives up. I don't think a UI should block forever on someting. There's another school of thought that says "no timeouts", let's just wait however long it takes", so that on really busy servers you eventually get something. But this covers much more fundamental problems. I. e. if you are still concerned about this, you can file a new bug, but I don't want to promise that we'll "fix" it on the UI side. Honestly, I'd start with filing a bug for the issue below. Let's rather fix the slowness than removing the defenses against it. > do you think 10-20s to list the vms is acceptable? It's certainly a bad user experience, so filing another bug for that is a good idea IMHO. On your demo machine above there is no shortage of CPU or RAM, so it's not limited system resources. Things are genuinely just slow. That may be either a bug or unavoidable slowness in libvirt-dbus, or libvirt itself, or polkit, or we are doing too many calls at one time. Change to verified per comment 32 and comment 41. New bug1820493 and bug1820469. *** Bug 1820469 has been marked as a duplicate of this bug. *** I just tried this again on current RHEL 7.9, with the following package versions: cockpit-machines-195.12-1.el7_9.x86_64 libvirt-4.5.0-36.el7.x86_64 polkit-0.112-26.el7.x86_64 kernel-3.10.0-1158.el7.x86_64 I followed the setup of the description and comment #1: yum install -y cockpit cockpit-machines useradd -G libvirt,libvirtdbus user echo user:foobar | chpasswd echo 'access_drivers = [ "polkit" ]' >> /etc/libvirt/libvirtd.conf systemctl try-restart libvirtd firewall-cmd --add-service cockpit systemctl start cockpit.socket I created 25 system VMs with for i in `seq 25`; do virt-install --connect qemu:///system --memory 50 --pxe --network network=default --virt-type qemu --os-variant alpinelinux3.8 --disk none --wait 0 --name test$i; done I logged in as "user" several times. "Connecting to virtualization service" takes about 10 seconds, then the 25 VMs get shown. @Ravindra: Note that your version cockpit-machines-195-1.el7.x86_64.rpm is too old and does *not* have this fix yet -- it's fixed in version 195.12, see the attached erratum here (https://errata.devel.redhat.com/advisory/53832). Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (cockpit bug fix update), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2020:4085 |