Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.

Bug 1975414

Summary: OVN databases are not easily parseable
Product: Red Hat Enterprise Linux 8 Reporter: Adrián Moreno <amorenoz>
Component: sosAssignee: Pavel Moravec <pmoravec>
Status: CLOSED WORKSFORME QA Contact: Upgrades and Supportability <upgrades-and-supportability>
Severity: unspecified Docs Contact:
Priority: unspecified    
Version: 8.2CC: agk, bmr, dceara, plambri, sbradley, theute
Target Milestone: betaFlags: pm-rhel: mirror+
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-12-06 22:02:57 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Adrián Moreno 2021-06-23 15:50:53 UTC
Description of problem:

The way we currently archive the OVN database is by just copying the db file.

Since ovsdb-server is transaction based, the resulting file has uncollapsed transactions. This is not a problem if you want to recreate the DB using ovsdb-server. However, if another program (e.g: insights) wants to parse that file, it's really complicated since it has to implement the ovsdb-server transaction mechanism (which btw, might change from one version to another).

A possible solution would be to use "ovsdb-client backup".

How reproducible:
always

Steps to Reproduce:
1. run sos report on a OVN master node

Actual results:

db contains transactions and is difficult to parse

Expected results:

The DB should be easy to parse by insights so we can develop rules based on the content of the DB

Additional info:

Comment 1 Adrián Moreno 2021-06-23 16:17:15 UTC
Something on the lines of this patch would do the work, I think:

https://github.com/amorenoz/sos/commit/ec6cfaa1a824e04ab62b1fee4c69e1674f874174

I'm happy to send it upstream for review

Comment 2 Pavel Moravec 2021-06-25 07:36:55 UTC
If we will take the backup in another way/format, isn't it worth stopping to collect the DBs? (is that the https://github.com/sosreport/sos/blob/master/sos/report/plugins/ovn_central.py#L134-L143 ? ) I just want to prevent collecting duplicate data, if possible.

Also, can't the "ovndbclient backup .." command get stuck in either situation? Can't it alter anything or produce some huge output (probably not)?

Are you aware that sos report will collect "ovndbclient backup" std output "only"? *Usually*, backup tools create some archive file with the backup - you dont need to collect such a file, as whole backup is printed to stdout, right?

Comment 3 Adrián Moreno 2021-06-29 07:41:44 UTC
(In reply to Pavel Moravec from comment #2)
> If we will take the backup in another way/format, isn't it worth stopping to
> collect the DBs? (is that the
> https://github.com/sosreport/sos/blob/master/sos/report/plugins/ovn_central.
> py#L134-L143 ? ) I just want to prevent collecting duplicate data, if
> possible.

The difference between ovsdb-client backup and copying the DB is the the former would not include ephemeral columns. The only ephemeral columns in OVN are related to the `Connection`'s state. CCing Dumitru from the OVN team to comment on this.

If we need the Connection ephemeral information we can dump it separately or use some ovsdb-client or ovs-appctl commands to show such information (if available).

> Also, can't the "ovndbclient backup .." command get stuck in either
> situation? Can't it alter anything or produce some huge output (probably
> not)?
> 

I'd say ovsdb-client backup is less intrusive than compacting but probably more than copying.
About blocking: if the ovsdb-server is terribly busy it the transaction could take long. However, we could use the "--timeout" option to limit a maximum time for the transaction to finish and maybe fall back to "cp" if we run out of time.

> Are you aware that sos report will collect "ovndbclient backup" std output
> "only"? *Usually*, backup tools create some archive file with the backup -
> you dont need to collect such a file, as whole backup is printed to stdout,
> right?

Correct. In fact, "ovsdb-client backup" has to to be piped to a file (or the client complaints), sos report does that already so we should we good. The result file is generally smaller than the content of the database file because the uncompacted file contains transactions that make the file bigger and its parsing more complicated.

Comment 4 Dumitru Ceara 2021-06-30 07:23:38 UTC
(In reply to Adrián Moreno from comment #3)
> (In reply to Pavel Moravec from comment #2)
> > If we will take the backup in another way/format, isn't it worth stopping to
> > collect the DBs? (is that the
> > https://github.com/sosreport/sos/blob/master/sos/report/plugins/ovn_central.
> > py#L134-L143 ? ) I just want to prevent collecting duplicate data, if
> > possible.
> 
> The difference between ovsdb-client backup and copying the DB is the the
> former would not include ephemeral columns. The only ephemeral columns in
> OVN are related to the `Connection`'s state. CCing Dumitru from the OVN team
> to comment on this.
> 
> If we need the Connection ephemeral information we can dump it separately or
> use some ovsdb-client or ovs-appctl commands to show such information (if
> available).
> 
> > Also, can't the "ovndbclient backup .." command get stuck in either
> > situation? Can't it alter anything or produce some huge output (probably
> > not)?
> > 
> 
> I'd say ovsdb-client backup is less intrusive than compacting but probably
> more than copying.
> About blocking: if the ovsdb-server is terribly busy it the transaction
> could take long. However, we could use the "--timeout" option to limit a
> maximum time for the transaction to finish and maybe fall back to "cp" if we
> run out of time.
> 
> > Are you aware that sos report will collect "ovndbclient backup" std output
> > "only"? *Usually*, backup tools create some archive file with the backup -
> > you dont need to collect such a file, as whole backup is printed to stdout,
> > right?
> 
> Correct. In fact, "ovsdb-client backup" has to to be piped to a file (or the
> client complaints), sos report does that already so we should we good. The
> result file is generally smaller than the content of the database file
> because the uncompacted file contains transactions that make the file bigger
> and its parsing more complicated.

Another option, that doesn't interact with ovsdb-server on the target
at all is to copy the DB file and then run:

ovsdb-tool -h 
ovsdb-tool: Open vSwitch database management utility
usage: ovsdb-tool [OPTIONS] COMMAND [ARG...]
[...]
  compact [DB [DST]]      compact DB in-place (or to DST)

If the DB is clustered we might need to first turn the copy of the DB
file into "standalone".

E.g., assuming the copy of the DB file is in /tmp/ovnnb_db.db:

$ ovsdb-tool cluster-to-standalone /tmp/ovnnb_db.db.standalone /tmp/ovnnb_db.db
$ ovsdb-tool compact /tmp/ovnnb_db.db.standalone

At this point, ovnnnb_db.db.standalone is compacted, only has one
transaction record, and we didn't interact with ovsdb-server at all.

Hth,
Dumitru

Comment 5 Adrián Moreno 2021-06-30 13:49:49 UTC
Thanks Dumitru,

That's also a very good idea.
I also thought of it but a couple of complications came to mind:

- Looking at the sos report plugin, I don't see a clear order between the "cp" and the command executions.
Pavel, can you please clarify if we can ensure a command is run *after* the file is copied?

- In systems such as Openstack where ovs/ovn commands are not available on the host, sos is smart enough to run these commands inside a container the does have them (the command would actually be "$ podman exec -it {POD_RUNNING_OVN} ovsdb-tool ..."). However, the file is directly available at the host (via podman volume mount), so the 'cp' command can run on the host. Therefore, we would need to copy it into a directory that is also volume-mounted into the container, exec the commands from the container, and then copy it again to its final destination.

Pavel may correct me if I'm wrong but that might be a bit overcomplicated. At that point we might as well do all that processing at the debugging phase which also comes with its complications:

- Tools like insights-core do not support running external tooling to do processing, so automatic parsing won't work out of the box.
- To ensure DB integrity, we would need to run a version of ovsdb-tool that matches the version of ovsdb-server that created the file.

Comment 6 Pavel Moravec 2021-07-01 07:21:41 UTC
(In reply to Adrián Moreno from comment #5)
> Thanks Dumitru,
> 
> That's also a very good idea.
> I also thought of it but a couple of complications came to mind:
> 
> - Looking at the sos report plugin, I don't see a clear order between the
> "cp" and the command executions.
> Pavel, can you please clarify if we can ensure a command is run *after* the
> file is copied?
> 
> - In systems such as Openstack where ovs/ovn commands are not available on
> the host, sos is smart enough to run these commands inside a container the
> does have them (the command would actually be "$ podman exec -it
> {POD_RUNNING_OVN} ovsdb-tool ..."). However, the file is directly available
> at the host (via podman volume mount), so the 'cp' command can run on the
> host. Therefore, we would need to copy it into a directory that is also
> volume-mounted into the container, exec the commands from the container, and
> then copy it again to its final destination.
> 
> Pavel may correct me if I'm wrong but that might be a bit overcomplicated.
> At that point we might as well do all that processing at the debugging phase
> which also comes with its complications:
> 
> - Tools like insights-core do not support running external tooling to do
> processing, so automatic parsing won't work out of the box.
> - To ensure DB integrity, we would need to run a version of ovsdb-tool that
> matches the version of ovsdb-server that created the file.

If the backup tool supports generating the DB backup on a given path (which seems so), you can call it to generate the dump/backup directly in the sos working directory tree. Like:

https://github.com/sosreport/sos/blob/master/sos/report/plugins/lvm2.py#L24-L37

calls "lvmdump -d /var/tmp/sos.3egb2mby/sosreport-pmoravec-rhel8-2021-07-01-fwwhqdg/sos_commands/lvm2/lvmdump"

If one uses exec_cmd to generate the file and then add_cmd_output to compact the DB and let sosreport to count with the file, it should work.


Anyway, if the DB is small, it can be easier to collect its backup/dump in both formats..?

Comment 7 Adrián Moreno 2021-07-01 08:09:21 UTC
(In reply to Pavel Moravec from comment #6)

> 
> If the backup tool supports generating the DB backup on a given path (which
> seems so), you can call it to generate the dump/backup directly in the sos
> working directory tree. Like:
> 

The backup tool has to be run inside a container (via 'podman exec') so the sos working directory tree won't be accessible from within the container.
We would have to make the command output it through stdout so that sos can then place the output of the command in the corresponding file.

> 
> Anyway, if the DB is small, it can be easier to collect its backup/dump in
> both formats..?

Dumitru can answer that better than me.

Using a single command, I see two options: 

a) run "podman exec $CONTAINER_NAME ovsdb-client --timeout 10 --target ${...} backup "
Note that ovsdb-client backup will abort if STDOUT is a tty, so we have to omit the "-t" in "podman exec".
Pros: a single command should do the work
Cons: the command does interact with the database

b) run something like "podman exec -it $CONTAINER_NAME sh -c 'cp /var/lib/openvswitch/ovnnb_db{,_backup}.db && ovsdb-tool cluster-to-standalone /var/lib/openvswitch/ovnnb_db_backup{_sa,}.db && ovsdb-tool compact /var/lib/openvswitch/ovnnb_db_backup_sa.db && cat /var/lib/openvswitch/ovnnb_db_backup_sa.db'"
Pros: Single command no interaction with ovsdb-server
Cons: Is the filename (as placed in the sos archive tree) going to be that long?. We have to cleanup the generated files.


Dumitru: does the "cluster-to-standalone" command destroy any information that might be necessary for debugging? Is it better to keep both (clustered and standalone) dbs?

Comment 8 Dumitru Ceara 2021-07-01 08:24:33 UTC
(In reply to Adrián Moreno from comment #7)
> (In reply to Pavel Moravec from comment #6)
> 
> > 
> > If the backup tool supports generating the DB backup on a given path (which
> > seems so), you can call it to generate the dump/backup directly in the sos
> > working directory tree. Like:
> > 
> 
> The backup tool has to be run inside a container (via 'podman exec') so the
> sos working directory tree won't be accessible from within the container.
> We would have to make the command output it through stdout so that sos can
> then place the output of the command in the corresponding file.
> 
> > 
> > Anyway, if the DB is small, it can be easier to collect its backup/dump in
> > both formats..?
> 
> Dumitru can answer that better than me.

This is a part I didn't think of until now..  It's probably better to
collect both the original and the compacted version of the DB.

> 
> Using a single command, I see two options: 
> 
> a) run "podman exec $CONTAINER_NAME ovsdb-client --timeout 10 --target
> ${...} backup "
> Note that ovsdb-client backup will abort if STDOUT is a tty, so we have to
> omit the "-t" in "podman exec".
> Pros: a single command should do the work
> Cons: the command does interact with the database
> 
> b) run something like "podman exec -it $CONTAINER_NAME sh -c 'cp
> /var/lib/openvswitch/ovnnb_db{,_backup}.db && ovsdb-tool
> cluster-to-standalone /var/lib/openvswitch/ovnnb_db_backup{_sa,}.db &&
> ovsdb-tool compact /var/lib/openvswitch/ovnnb_db_backup_sa.db && cat
> /var/lib/openvswitch/ovnnb_db_backup_sa.db'"
> Pros: Single command no interaction with ovsdb-server
> Cons: Is the filename (as placed in the sos archive tree) going to be that
> long?. We have to cleanup the generated files.
> 
> 
> Dumitru: does the "cluster-to-standalone" command destroy any information
> that might be necessary for debugging? Is it better to keep both (clustered
> and standalone) dbs?

Good point, it does, and as a matter of fact I think "ovsdb-client backup"
does that too.  We'll be losing all RAFT logs (which are stored as json
records in the DB file in clustered format).

Unfortunately, if we want to have all debugging info we'd need both the
original DB file (clustered format, non-compacted) for full logs and the
standalone, compacted DB file for simplified parsing by external tools.

Comment 9 Pavel Moravec 2021-07-01 09:41:30 UTC
OK, if the DBs are small, then I think the #c1 / https://github.com/amorenoz/sos/commit/ec6cfaa1a824e04ab62b1fee4c69e1674f874174 seems a good approach.

Adrián, let me know if I shall do a PR per the commit, or if you will do it by yourself (either works fine for me).

Comment 10 Adrián Moreno 2021-07-01 10:25:26 UTC
(In reply to Pavel Moravec from comment #9)
> OK, if the DBs are small, then I think the #c1 /
> https://github.com/amorenoz/sos/commit/
> ec6cfaa1a824e04ab62b1fee4c69e1674f874174 seems a good approach.
> 

I guess the SB can get pretty big. Dumitru has worked on scale tests, do you have an order of magnitude?

Being in json format, gzip/xz will do a decent job compacting them.

As an example, on one of my sample databases I got the following figures:

uncompacted: 5.9M (820K gzipped)
compacted: 74K (13K gzipped)

> Adrián, let me know if I shall do a PR per the commit, or if you will do it
> by yourself (either works fine for me).

I'll send the PR, no worries.

Comment 11 Dumitru Ceara 2021-07-01 10:50:25 UTC
(In reply to Adrián Moreno from comment #10)
> (In reply to Pavel Moravec from comment #9)
> > OK, if the DBs are small, then I think the #c1 /
> > https://github.com/amorenoz/sos/commit/
> > ec6cfaa1a824e04ab62b1fee4c69e1674f874174 seems a good approach.
> > 
> 
> I guess the SB can get pretty big. Dumitru has worked on scale tests, do you
> have an order of magnitude?
> 
> Being in json format, gzip/xz will do a decent job compacting them.
> 
> As an example, on one of my sample databases I got the following figures:
> 
> uncompacted: 5.9M (820K gzipped)
> compacted: 74K (13K gzipped)

They can be even larger, e.g., from an OpenShift scale test:

uncompacted: 319M (66M gzipped)
compacted: 131M (27M gzipped)

Comment 12 Pavel Moravec 2021-07-01 12:16:43 UTC
Huh, such a big dump being collected twice..? I would prefer one collection then - recall it must take some time to generate the dump/backup, to pack it (no need to pack it "manually" as sosreport will tar.gz it), to transfer among computers, unpack,..

Could you please provide:
- a sequence of commands in bash (including podman exec .. stuff) how to collect a DB snapshot/dump/export and store it in some file
- what commands from the current sosreport will be obsolete (as they collect the DB backup now which would be duplicate)

or optionally the PR (that would need proper exec_cmd + add_copy_spec "orchestration" where I am offering the help, knowing sos API).

Comment 13 Adrián Moreno 2021-07-01 16:15:54 UTC
(In reply to Pavel Moravec from comment #12)
> Huh, such a big dump being collected twice..? I would prefer one collection
> then - recall it must take some time to generate the dump/backup, to pack it
> (no need to pack it "manually" as sosreport will tar.gz it), to transfer
> among computers, unpack,..
> 
> Could you please provide:
> - a sequence of commands in bash (including podman exec .. stuff) how to
> collect a DB snapshot/dump/export and store it in some file
> - what commands from the current sosreport will be obsolete (as they collect
> the DB backup now which would be duplicate)
> 
> or optionally the PR (that would need proper exec_cmd + add_copy_spec
> "orchestration" where I am offering the help, knowing sos API).

The sequence could be:

for db in ovnnb.db ovnsb.db; do
    cp /var/lib/openvswitch/$db /var/lib/openvswitch/backup_$db
    podman exec -it $CONTAINER ovsdb-tool compact /var/lib/openvswitch/backup_$db
    mv /var/lib/openvswitch/backup_$db /tmp/${SOS_TREE)/var/lib/openvswitch/$db
done

(this assumes /var/lib/openvswitch is volume mounted inside the container, which seems to be the same assumption made in current implementation)

If you point me to some available documentation about the sos API I can try implement this :)

Comment 14 Adrián Moreno 2021-07-02 15:53:44 UTC
Oh, wait, ovsdb-tool compact does not support clustered DBs.

So, either we make it standalone (and loose raft information), we copy it uncompacted (and quite big) or we interact the ovsdb-server (via "appctl ovsdb/compact" or "ovsdb-client backup").

Comment 15 Pavel Moravec 2021-08-02 07:24:25 UTC
(In reply to Adrián Moreno from comment #14)
> Oh, wait, ovsdb-tool compact does not support clustered DBs.
> 
> So, either we make it standalone (and loose raft information), we copy it
> uncompacted (and quite big) or we interact the ovsdb-server (via "appctl
> ovsdb/compact" or "ovsdb-client backup").

OK, looking forward for specific requirement / specification, then.

Comment 16 Adrián Moreno 2021-08-05 09:54:39 UTC
We discussed the same in must-gather scripts: https://github.com/openshift/must-gather/pull/245

Faced with the same problem we ended up just copying the full db file (which is the current behavior).

If @dceara I think we can reject this BZ and figure out a way to work around the parsing issues elsewhere.

Comment 17 Dumitru Ceara 2021-08-30 15:15:30 UTC
(In reply to Adrián Moreno from comment #16)
> We discussed the same in must-gather scripts:
> https://github.com/openshift/must-gather/pull/245
> 
> Faced with the same problem we ended up just copying the full db file (which
> is the current behavior).
> 
> If @dceara I think we can reject this BZ and figure out a way to
> work around the parsing issues elsewhere.

Unfortunately it looks like this is the way to go for now.

Comment 18 Pavel Moravec 2021-12-06 22:02:57 UTC
OK, so I am closing this BZ as WORKSFORME. Let me know if something else is required.