Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.

Bug 1930037

Summary: subscription-manager-cockpit is broken on KVM guest images: [Errno 20] Not a directory: '/etc/pki/product'
Product: Red Hat Enterprise Linux 8 Reporter: Martin Pitt <mpitt>
Component: subscription-managerAssignee: Pino Toscano <ptoscano>
Status: CLOSED UPSTREAM QA Contact: Red Hat subscription-manager QE Team <rhsm-qe>
Severity: low Docs Contact:
Priority: unspecified    
Version: 8.4CC: cdonnell, csnyder, mmarusak, ptoscano, redakkan, thozza
Target Milestone: rcKeywords: Regression, Triaged
Target Release: 8.5Flags: pm-rhel: mirror+
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: subscription-manager-1.28.16-1.el8 Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of:
: 1931429 (view as bug list) Environment:
Last Closed: 2021-05-12 15:49:06 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Martin Pitt 2021-02-18 09:43:05 UTC
Description of problem: Until at least 8 days ago [1], with rhel-guest-image-8.4-594.x86_64.qcow2 [2], subscription-manager was enabled by default in dnf:

# grep enabled= /etc/yum/pluginconf.d/subscription-manager.conf
enabled=1

That was subscription-manager-1.28.11-1.el8.x86_64.

With the most recent RHEL 8.4 nightlies this now broke [3]

# grep enabled= /etc/yum/pluginconf.d/subscription-manager.conf
enabled=0

This breaks at least cockpit and apparently also subscription-manager-cockpit [4]


Version-Release number of selected component (if applicable):


How reproducible:


Steps to Reproduce:
1. Download current cloud image and some cloud-init.iso that will allow you to log in; using cockpit CI's here, but that's nothing special
curl -O http://download.eng.brq.redhat.com/rhel-8/nightly/RHEL-8/latest-RHEL-8.4/compose/BaseOS/x86_64/images/rhel-guest-image-8.4-649.x86_64.qcow2
curl -L -O https://github.com/cockpit-project/bots/raw/master/machine/cloud-init.iso
2. Boot the image
qemu-system-x86_64 -enable-kvm  -nographic -m 2048 -drive file="rhel-guest-image-8.4-649.x86_64.qcow2",if=virtio -snapshot -cdrom cloud-init.iso
3. Log in as root:foobar
4. grep enabled= /etc/yum/pluginconf.d/subscription
5. Open cockpit, check Health box
6. Check the status on the CLI:
subscription-manager status`
7. grep enabled= /etc/yum/pluginconf.d/subscription

Actual results:
4. sub-man is disabled in dnf by default
5. No complaint about "This system is not subscribed"
7. `subscription-manager status` *changed* /etc/yum/pluginconf.d/subscription and set enabled=1

Expected results:
4. sub-man should be enabled in dnf by default on RHEL (but not on Fedora/CentOS)
5. Shows "This system is not subscribed"
7. Calling a status command has no business changing config files in /etc. This confuses Ansible and other tools. such state belongs into /var.

Especially 7 is really upsetting, please don't do this.

If you insist on this for whatever reason (I don't claim I have insights in all the details here), please advise how cockpit and sub-man-cockpit are supposed to check if the OS requires subscriptions to be able to install/update packages.

Thank you!

Additional info:

[1] https://github.com/cockpit-project/bots/pull/1645
[2] http://download.eng.blr.redhat.com/rhel-8/nightly/RHEL-8/latest-RHEL-8.4/compose/BaseOS/x86_64/images//rhel-guest-image-8.4-594.x86_64.qcow2
[3] https://github.com/cockpit-project/bots/pull/1684
[4] see https://github.com/candlepin/subscription-manager/pull/2447 and https://github.com/candlepin/subscription-manager/pull/2448 for some related discussion

Comment 1 Martin Pitt 2021-02-18 09:45:51 UTC
This is apparently related to some changes in osbuild: https://issues.redhat.com/browse/COMPOSER-717 So please reassign to osbuild if that is the case. Thank you!

Comment 2 Martin Pitt 2021-02-18 09:53:03 UTC
> Version-Release number of selected component (if applicable):

Whoops, forgot to paste that: subscription-manager-1.28.12-1.el8.x86_64

Comment 3 Martin Pitt 2021-02-18 14:33:25 UTC
Tomas just kindly pointed out that enabled= in /etc/yum/pluginconf.d/product-id.conf has the exact same behaviour. That may also explain the weird errors that the subscription-manager tests get now:

   [Errno 20] Not a directory: '/etc/pki/product'

See test log: https://logs.cockpit-project.org/logs/pull-1684-20210218-111709-5c46a7d0-rhel-8-4-candlepin-subscription-manager/log.html#1-2
and screenshot: https://logs.cockpit-project.org/logs/pull-1684-20210218-111709-5c46a7d0-rhel-8-4-candlepin-subscription-manager/TestSubscriptions-testInsights-rhel-8-4-127.0.0.2-2301-FAIL.png

and indeed, when I set enabled=1 in product-id.conf as well (as I just did in https://github.com/cockpit-project/bots/pull/1684), then subscription-manager tests are happy again.

Comment 4 Tomáš Hozza 2021-02-18 15:15:41 UTC
Just a note, that the fact that RHSM DNF plugins are disabled on RHEL Guest images is not a Regression. This has been the situation since RHEL-6.

8.4 images produced before osbuild-composer has been used have these plugins also disabled (e.g. http://download.eng.brq.redhat.com/rhel-8/nightly/RHEL-8/RHEL-8.4.0-20201119.n.0/compose/BaseOS/x86_64/images/rhel-guest-image-8.4-294.x86_64.qcow2).

While 8.4 images produced since ~Dec 2020, until we fixed the discrepancy in ~Feb 2021 had RHSM DNF plugins enabled (e.g. the first still available image built by osbuild-composer where DNF plugins got enabled - http://download.eng.brq.redhat.com/rhel-8/nightly/RHEL-8/RHEL-8.4.0-20201209.n.0/compose/BaseOS/x86_64/images/rhel-guest-image-8.4-338.x86_64.qcow2).

Comment 5 Martin Pitt 2021-02-18 15:38:41 UTC
Thanks Tomas! So is this what happened then:

 - The qcow images had enabled=0 for a long time, until in December osbuild took over building them
 - That had an unexpected change that it set enabled=1 from December to ~ 1 week ago
 - A few days ago, osbuild switched it back to off

cockpit CI switching from virt-install to qcow cloud images happened right during these few weeks (we switched in January).

That would explain both your and my point of view, and clear the confusion. Assuming that is what happened, I still have some new questions, though:

 * I take it that we still expect users to install some packages on a cloud image, which will require a subscription. What is the *reason* why sub-man is disabled by default on cloud images? I.e. what does this dynamic enabled=0 → 1 achieve over a static enabled=1 config?

 * What is the official documented thing that a user has to invoke to enable sub-man on these images?

 * Cockpit is one such supported place that tells you that your system is not subscribed, and Cockpit is installed by default. So obviously, the current situation of not saying that is a bug. What *should* cockpit do to test that a system is not subscribed, but needs to be? We have tested for enabled=1 in /etc/yum/pluginconf.d/subscription for years, but obviously that is not the right thing. But what is?

Thanks!

Comment 6 Tomáš Hozza 2021-02-18 16:46:44 UTC
(In reply to Martin Pitt from comment #5)
> Thanks Tomas! So is this what happened then:
> 
>  - The qcow images had enabled=0 for a long time, until in December osbuild
> took over building them
>  - That had an unexpected change that it set enabled=1 from December to ~ 1
> week ago
>  - A few days ago, osbuild switched it back to off
> 
> cockpit CI switching from virt-install to qcow cloud images happened right
> during these few weeks (we switched in January).
> 
> That would explain both your and my point of view, and clear the confusion.
> Assuming that is what happened, I still have some new questions, though:

Yes, that is exactly what happened ;)


>  * I take it that we still expect users to install some packages on a cloud
> image, which will require a subscription. What is the *reason* why sub-man
> is disabled by default on cloud images? I.e. what does this dynamic
> enabled=0 → 1 achieve over a static enabled=1 config?

"cloud images" is an overloaded term, so I'll make more distinction.

The images that you used are (KVM) Guest images, which are intended for use only on KVM-based hypervisors. I was not able to determine the reason why RHSM DNF plugins have been disabled on KVM Guest images. My assumption would be that in the past, they were intended also for use with cloud providers, which is not the case any more (but I don't know). So I really don't know...

Disabling these plugins on images used with cloud providers (e.g. EC2 images) makes sense, because the content is delivered via other means, specifically via RHUI (Red Hat Update Infrastructure). While for any special content, one has to use sub-man, RHEL content with couple of addons is available via RHUI. If these DNF plugins are not disabled, then DNF produces misleading log messages on each command, that the system is not subscribed, while subscribing it is not needed to access content.

The reasoning is that if you really need to subscribe the system, then doing it via GUI/CLI will auto-enable these plugins and everything works.

>  * What is the official documented thing that a user has to invoke to enable
> sub-man on these images?

Nothing AFAIK. DNF plugins are auto-enabled on the very first subscription-manager command. Actually any command.

>  * Cockpit is one such supported place that tells you that your system is
> not subscribed, and Cockpit is installed by default. So obviously, the
> current situation of not saying that is a bug. What *should* cockpit do to
> test that a system is not subscribed, but needs to be? We have tested for
> enabled=1 in /etc/yum/pluginconf.d/subscription for years, but obviously
> that is not the right thing. But what is?

I'm not sub-man devel, but based on my information, the fact that the RHSM DNF plugins are enabled tells you nothing about the subscription status of the OS. I would say that probably checking the output of a specific sub-man command, but I will deffer the answer to actual members of the subscription-manager team... Also the situation is a bit complicated if the system uses RHUI or the latest RHSM feature - SCA (Simple Content Access)

> 
> Thanks!

Comment 7 Martin Pitt 2021-02-19 04:28:55 UTC
(In reply to Tomáš Hozza from comment #6)
> > What *should* cockpit do to test that a system is not subscribed, but needs to be?
> > We have tested for enabled=1 in /etc/yum/pluginconf.d/subscription for years, but obviously that is not the right thing. But what is?
>
> the fact that the RHSM DNF plugins are enabled tells you nothing about the subscription status of the OS.

Right, and that's also not what I meant. Cockpit obviously calls rhsm D-Bus to see if the system *is* subscribed or not -- but that is not sufficient, as sub-man says "wah wah not subscribed" on Fedora or CentOS as well. So we additionally need to check if the system *needs* to be subscribed, and that's a question to dnf, not to rhsm. We thought querying the dnf sub-man plugin enablement status was appropriate for that.

Hence my question to the sub-man developers: What is a robust replacement for that, if it's not the config files?

Thanks!

Comment 8 Tomáš Hozza 2021-02-19 08:58:06 UTC
(In reply to Martin Pitt from comment #7)
> Right, and that's also not what I meant. Cockpit obviously calls rhsm D-Bus
> to see if the system *is* subscribed or not -- but that is not sufficient,
> as sub-man says "wah wah not subscribed" on Fedora or CentOS as well. So we
> additionally need to check if the system *needs* to be subscribed, and
> that's a question to dnf, not to rhsm. We thought querying the dnf sub-man
> plugin enablement status was appropriate for that.

Thanks for the wider context, I now understand better the problem that you are trying to solve. What I meant mostly is that the RHSM DNF plugins are by default enabled when the subcsription-manager RPM is installed. So unless the image is built in a specific way (like the RHEL qcow2), the subscription-manager and product-id DNF plugins will be most probably enabled even if the system still needs subscribing.

Comment 9 Martin Pitt 2021-02-22 12:04:34 UTC
We just had a meeting to resolve this. The sub-man team's recommendation was to stop looking at the dnf config file, and instead use the RHSM ListInstalledProducts() call.

On current RHEL 8.4 cloud image:

# busctl call com.redhat.RHSM1 /com/redhat/RHSM1/Products com.redhat.RHSM1.Products ListInstalledProducts 'sa{sv}s' '' 0 ''
s "[[\"Red Hat Enterprise Linux for x86_64 Beta\", \"486\", \"8.4 Beta\", \"x86_64\", \"unknown\", [], \"\", \"\"]]"

where as on CentOS 8 stream or Fedora it returns

s "[]"

so that's as expected and good.

This needs to be done in subscription-manager-cockpit, so this bug report can stay on the subscription-manager component. I'll clone it to track making that change to Cockpit's Overview and Software Updates pages.

[1] https://www.candlepinproject.org/docs/subscription-manager/dbus_objects.html

Comment 10 Martin Pitt 2021-02-23 09:51:18 UTC
On the cockpit side (bug 1931429) I fixed this in https://github.com/cockpit-project/cockpit/pull/15401 . subscription-manager has an (outdated) copy of packagekit.js, so my first attempt was to update that to the version of that PR. But subscription-manager-cockpit does not use the affected watchRedHatSubscription() function at all, and its old code copy did not even have the code to check /etc/yum/pluginconf.d/subscription . Nor does the s-m-c code check anything in pluginconf.d/ anywhere else. Instead, cockpit/src/subscriptions-client.js already seems to use productsService.ListInstalledProducts(), thus in *theory* it should already work with the qcow images.

But as the test on the refreshed image shows, that is clearly not the case:

  https://logs.cockpit-project.org/logs/pull-1710-20210223-090659-d6b23c6c-rhel-8-4-candlepin-subscription-manager/log.html

They still all fail on "[Errno 20] Not a directory: '/etc/pki/product'", both in the logs and in the visible dialogs, e.g. 

  https://logs.cockpit-project.org/logs/pull-1710-20210223-090659-d6b23c6c-rhel-8-4-candlepin-subscription-manager/TestSubscriptions-testRegister-rhel-8-4-127.0.0.2-2301-FAIL.png

With that, I will leave this investigation to you as domain experts -- I'm afraid I have no idea how the enabled=0 default somehow magically removes /etc/pki/product/. Keep in mind that the integration test even explicitly creates that:

        def download_product(product):
            prod_id = product["id"]
            filename = os.path.join(self.tmpdir, "%s.pem" % prod_id)
            self.candlepin.download("/home/admin/candlepin/generated_certs/%s.pem" % prod_id, filename)
            m.upload([filename], "/etc/pki/product")

        download_product(PRODUCT_SNOWY)
        download_product(PRODUCT_SHARED)

Comment 11 Pino Toscano 2021-04-08 11:50:29 UTC
Hi Martin,

I'm trying to reproduce the issue reported in this bug, which has a semi-long and interesting discussion.

I did the following steps:

1) download a very recent RHEL qcow2 image: rhel-guest-image-8.5-286.x86_64.qcow2 (current 8.5 nightly)
   - I tried a 8.4 version of ~2 months ago as well, rhel-guest-image-8.4-846.x86_64.qcow2

As first check, I checked the content of the disk images, and both lack /etc/pki/product.

2) download the cloud-init.iso from cockpit-project.org
3) create a VM to boot the qcow2 disk with the cloud-init ISO
4) boot the VM
5) log in as root:foobar

At this point, I see that /etc/pki/product still does not exist.

6) enable cockpit: systemctl enable --now cockpit.socket
7) open https://address-of-vm:9090/ in a web browser, and only do log in without any action in cockpit

At this point, I see that /etc/pki/product still does not exist.

8) switch to the "Subscriptions" tab/section

After a brief spinning, the page correctly shows: Status: This system is currently not registered.
Also, now /etc/pki/product exists, and it is empty.
Furthermore, /var/log/rhsm/rhsm.log does not show anything about /etc/pki/product.

Alternative route from step 6:

6a) invoke any subscription-manager command, e.g.: subscription-manager status

Also, now /etc/pki/product exists, and it is empty.

From what I can see (src/subscription_manager/certdirectory.py), subscription-manager has been ensuring that the certificate directories exist for more than 10 years at this point.

Martin, is there anything that I missed? What's left for us to do here?
Thanks!

Comment 12 Martin Pitt 2021-04-08 15:47:45 UTC
Pino, thanks for checking the "real-life" behaviour -- I'm glad that it does not seem to break the actual product. The main issue is that this error appears on most of your CI tests for rhel-8.4, e.g. here: https://logs.cockpit-project.org/logs/pull-2557-20210408-110040-fee291b6-rhel-8-4/log.html

It does not appear as "red" failure as we marked that as a known failure in https://github.com/cockpit-project/bots/issues/1711 -- but it still prevents actually testing s-m-c on 8.4. (That's also the main reason to still test on 8.3, as there it works fine)

Comment 13 Pino Toscano 2021-04-08 16:30:42 UTC
(In reply to Martin Pitt from comment #12)
> Pino, thanks for checking the "real-life" behaviour -- I'm glad that it does
> not seem to break the actual product. The main issue is that this error
> appears on most of your CI tests for rhel-8.4, e.g. here:
> https://logs.cockpit-project.org/logs/pull-2557-20210408-110040-fee291b6-rhel-8-4/log.html

Thanks, all the former links of cockpit CI are 404, so an up-to-date log is definitely useful.

The failing code is:

    def list_all(self):
        all_items = []
        if not os.path.exists(self.path):
            return all_items

        for fn in os.listdir(self.path):
            p = (self.path, fn)
            all_items.append(p)
        return all_items

and os.listdir() that fails with ENOTDIR... so, given that the os.path.exists() for the same path returns True, this means that that path is not a directory, but most likely a regular file. Looking at the scripts used by integration tests, I see this bit that you also mentioned earlier (in comment 10):

        def download_product(product):
            prod_id = product["id"]
            filename = os.path.join(self.tmpdir, "%s.pem" % prod_id)
            self.candlepin.download("/home/admin/candlepin/generated_certs/%s.pem" % prod_id, filename)
            m.upload([filename], "/etc/pki/product")

Considering that /etc/pki/product does not exist, my fear is that m.upload() will copy the .pem file as /etc/pki/product file.
Let's see: https://github.com/candlepin/subscription-manager/pull/2558

Comment 14 Pino Toscano 2021-04-08 16:49:20 UTC
(In reply to Pino Toscano from comment #13)
> Considering that /etc/pki/product does not exist, my fear is that m.upload()
> will copy the .pem file as /etc/pki/product file.
> Let's see: https://github.com/candlepin/subscription-manager/pull/2558

Aaaand... it looks like it did it. Martin, would you be able to take a look at the PR?

Thanks!

Comment 15 Martin Pitt 2021-04-08 17:42:06 UTC
Thanks Pino! I'm relieved! I reviewed the PR, but of course I can't land it.

Comment 16 Pino Toscano 2021-04-09 07:10:11 UTC
Martin, thanks for the review!

I'm lowering the severity to low, since (at least so far) it seems only an issue in the integration tests and not in real life.

Comment 19 Martin Pitt 2021-05-12 15:15:25 UTC
Thanks! It's all good from my end now, so fine for me to drop this bug (closed/currentrelease or any other way).

Comment 20 Pino Toscano 2021-05-12 15:49:06 UTC
Thanks Martin for the confirmation!

Since it was only an issue in the CI scripts with no impact on the actual product, and it is already fixed upstream, I'm closing this bug as such.

Thanks everyone for the bug report and the feedback here provided!