Bug 1975414
| Summary: | OVN databases are not easily parseable | ||
|---|---|---|---|
| Product: | Red Hat Enterprise Linux 8 | Reporter: | Adrián Moreno <amorenoz> |
| Component: | sos | Assignee: | Pavel Moravec <pmoravec> |
| Status: | CLOSED WORKSFORME | QA Contact: | Upgrades and Supportability <upgrades-and-supportability> |
| Severity: | unspecified | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 8.2 | CC: | agk, bmr, dceara, plambri, sbradley, theute |
| Target Milestone: | beta | Flags: | pm-rhel:
mirror+
|
| Target Release: | --- | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2021-12-06 22:02:57 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Adrián Moreno
2021-06-23 15:50:53 UTC
Something on the lines of this patch would do the work, I think: https://github.com/amorenoz/sos/commit/ec6cfaa1a824e04ab62b1fee4c69e1674f874174 I'm happy to send it upstream for review If we will take the backup in another way/format, isn't it worth stopping to collect the DBs? (is that the https://github.com/sosreport/sos/blob/master/sos/report/plugins/ovn_central.py#L134-L143 ? ) I just want to prevent collecting duplicate data, if possible. Also, can't the "ovndbclient backup .." command get stuck in either situation? Can't it alter anything or produce some huge output (probably not)? Are you aware that sos report will collect "ovndbclient backup" std output "only"? *Usually*, backup tools create some archive file with the backup - you dont need to collect such a file, as whole backup is printed to stdout, right? (In reply to Pavel Moravec from comment #2) > If we will take the backup in another way/format, isn't it worth stopping to > collect the DBs? (is that the > https://github.com/sosreport/sos/blob/master/sos/report/plugins/ovn_central. > py#L134-L143 ? ) I just want to prevent collecting duplicate data, if > possible. The difference between ovsdb-client backup and copying the DB is the the former would not include ephemeral columns. The only ephemeral columns in OVN are related to the `Connection`'s state. CCing Dumitru from the OVN team to comment on this. If we need the Connection ephemeral information we can dump it separately or use some ovsdb-client or ovs-appctl commands to show such information (if available). > Also, can't the "ovndbclient backup .." command get stuck in either > situation? Can't it alter anything or produce some huge output (probably > not)? > I'd say ovsdb-client backup is less intrusive than compacting but probably more than copying. About blocking: if the ovsdb-server is terribly busy it the transaction could take long. However, we could use the "--timeout" option to limit a maximum time for the transaction to finish and maybe fall back to "cp" if we run out of time. > Are you aware that sos report will collect "ovndbclient backup" std output > "only"? *Usually*, backup tools create some archive file with the backup - > you dont need to collect such a file, as whole backup is printed to stdout, > right? Correct. In fact, "ovsdb-client backup" has to to be piped to a file (or the client complaints), sos report does that already so we should we good. The result file is generally smaller than the content of the database file because the uncompacted file contains transactions that make the file bigger and its parsing more complicated. (In reply to Adrián Moreno from comment #3) > (In reply to Pavel Moravec from comment #2) > > If we will take the backup in another way/format, isn't it worth stopping to > > collect the DBs? (is that the > > https://github.com/sosreport/sos/blob/master/sos/report/plugins/ovn_central. > > py#L134-L143 ? ) I just want to prevent collecting duplicate data, if > > possible. > > The difference between ovsdb-client backup and copying the DB is the the > former would not include ephemeral columns. The only ephemeral columns in > OVN are related to the `Connection`'s state. CCing Dumitru from the OVN team > to comment on this. > > If we need the Connection ephemeral information we can dump it separately or > use some ovsdb-client or ovs-appctl commands to show such information (if > available). > > > Also, can't the "ovndbclient backup .." command get stuck in either > > situation? Can't it alter anything or produce some huge output (probably > > not)? > > > > I'd say ovsdb-client backup is less intrusive than compacting but probably > more than copying. > About blocking: if the ovsdb-server is terribly busy it the transaction > could take long. However, we could use the "--timeout" option to limit a > maximum time for the transaction to finish and maybe fall back to "cp" if we > run out of time. > > > Are you aware that sos report will collect "ovndbclient backup" std output > > "only"? *Usually*, backup tools create some archive file with the backup - > > you dont need to collect such a file, as whole backup is printed to stdout, > > right? > > Correct. In fact, "ovsdb-client backup" has to to be piped to a file (or the > client complaints), sos report does that already so we should we good. The > result file is generally smaller than the content of the database file > because the uncompacted file contains transactions that make the file bigger > and its parsing more complicated. Another option, that doesn't interact with ovsdb-server on the target at all is to copy the DB file and then run: ovsdb-tool -h ovsdb-tool: Open vSwitch database management utility usage: ovsdb-tool [OPTIONS] COMMAND [ARG...] [...] compact [DB [DST]] compact DB in-place (or to DST) If the DB is clustered we might need to first turn the copy of the DB file into "standalone". E.g., assuming the copy of the DB file is in /tmp/ovnnb_db.db: $ ovsdb-tool cluster-to-standalone /tmp/ovnnb_db.db.standalone /tmp/ovnnb_db.db $ ovsdb-tool compact /tmp/ovnnb_db.db.standalone At this point, ovnnnb_db.db.standalone is compacted, only has one transaction record, and we didn't interact with ovsdb-server at all. Hth, Dumitru Thanks Dumitru,
That's also a very good idea.
I also thought of it but a couple of complications came to mind:
- Looking at the sos report plugin, I don't see a clear order between the "cp" and the command executions.
Pavel, can you please clarify if we can ensure a command is run *after* the file is copied?
- In systems such as Openstack where ovs/ovn commands are not available on the host, sos is smart enough to run these commands inside a container the does have them (the command would actually be "$ podman exec -it {POD_RUNNING_OVN} ovsdb-tool ..."). However, the file is directly available at the host (via podman volume mount), so the 'cp' command can run on the host. Therefore, we would need to copy it into a directory that is also volume-mounted into the container, exec the commands from the container, and then copy it again to its final destination.
Pavel may correct me if I'm wrong but that might be a bit overcomplicated. At that point we might as well do all that processing at the debugging phase which also comes with its complications:
- Tools like insights-core do not support running external tooling to do processing, so automatic parsing won't work out of the box.
- To ensure DB integrity, we would need to run a version of ovsdb-tool that matches the version of ovsdb-server that created the file.
(In reply to Adrián Moreno from comment #5) > Thanks Dumitru, > > That's also a very good idea. > I also thought of it but a couple of complications came to mind: > > - Looking at the sos report plugin, I don't see a clear order between the > "cp" and the command executions. > Pavel, can you please clarify if we can ensure a command is run *after* the > file is copied? > > - In systems such as Openstack where ovs/ovn commands are not available on > the host, sos is smart enough to run these commands inside a container the > does have them (the command would actually be "$ podman exec -it > {POD_RUNNING_OVN} ovsdb-tool ..."). However, the file is directly available > at the host (via podman volume mount), so the 'cp' command can run on the > host. Therefore, we would need to copy it into a directory that is also > volume-mounted into the container, exec the commands from the container, and > then copy it again to its final destination. > > Pavel may correct me if I'm wrong but that might be a bit overcomplicated. > At that point we might as well do all that processing at the debugging phase > which also comes with its complications: > > - Tools like insights-core do not support running external tooling to do > processing, so automatic parsing won't work out of the box. > - To ensure DB integrity, we would need to run a version of ovsdb-tool that > matches the version of ovsdb-server that created the file. If the backup tool supports generating the DB backup on a given path (which seems so), you can call it to generate the dump/backup directly in the sos working directory tree. Like: https://github.com/sosreport/sos/blob/master/sos/report/plugins/lvm2.py#L24-L37 calls "lvmdump -d /var/tmp/sos.3egb2mby/sosreport-pmoravec-rhel8-2021-07-01-fwwhqdg/sos_commands/lvm2/lvmdump" If one uses exec_cmd to generate the file and then add_cmd_output to compact the DB and let sosreport to count with the file, it should work. Anyway, if the DB is small, it can be easier to collect its backup/dump in both formats..? (In reply to Pavel Moravec from comment #6) > > If the backup tool supports generating the DB backup on a given path (which > seems so), you can call it to generate the dump/backup directly in the sos > working directory tree. Like: > The backup tool has to be run inside a container (via 'podman exec') so the sos working directory tree won't be accessible from within the container. We would have to make the command output it through stdout so that sos can then place the output of the command in the corresponding file. > > Anyway, if the DB is small, it can be easier to collect its backup/dump in > both formats..? Dumitru can answer that better than me. Using a single command, I see two options: a) run "podman exec $CONTAINER_NAME ovsdb-client --timeout 10 --target ${...} backup " Note that ovsdb-client backup will abort if STDOUT is a tty, so we have to omit the "-t" in "podman exec". Pros: a single command should do the work Cons: the command does interact with the database b) run something like "podman exec -it $CONTAINER_NAME sh -c 'cp /var/lib/openvswitch/ovnnb_db{,_backup}.db && ovsdb-tool cluster-to-standalone /var/lib/openvswitch/ovnnb_db_backup{_sa,}.db && ovsdb-tool compact /var/lib/openvswitch/ovnnb_db_backup_sa.db && cat /var/lib/openvswitch/ovnnb_db_backup_sa.db'" Pros: Single command no interaction with ovsdb-server Cons: Is the filename (as placed in the sos archive tree) going to be that long?. We have to cleanup the generated files. Dumitru: does the "cluster-to-standalone" command destroy any information that might be necessary for debugging? Is it better to keep both (clustered and standalone) dbs? (In reply to Adrián Moreno from comment #7) > (In reply to Pavel Moravec from comment #6) > > > > > If the backup tool supports generating the DB backup on a given path (which > > seems so), you can call it to generate the dump/backup directly in the sos > > working directory tree. Like: > > > > The backup tool has to be run inside a container (via 'podman exec') so the > sos working directory tree won't be accessible from within the container. > We would have to make the command output it through stdout so that sos can > then place the output of the command in the corresponding file. > > > > > Anyway, if the DB is small, it can be easier to collect its backup/dump in > > both formats..? > > Dumitru can answer that better than me. This is a part I didn't think of until now.. It's probably better to collect both the original and the compacted version of the DB. > > Using a single command, I see two options: > > a) run "podman exec $CONTAINER_NAME ovsdb-client --timeout 10 --target > ${...} backup " > Note that ovsdb-client backup will abort if STDOUT is a tty, so we have to > omit the "-t" in "podman exec". > Pros: a single command should do the work > Cons: the command does interact with the database > > b) run something like "podman exec -it $CONTAINER_NAME sh -c 'cp > /var/lib/openvswitch/ovnnb_db{,_backup}.db && ovsdb-tool > cluster-to-standalone /var/lib/openvswitch/ovnnb_db_backup{_sa,}.db && > ovsdb-tool compact /var/lib/openvswitch/ovnnb_db_backup_sa.db && cat > /var/lib/openvswitch/ovnnb_db_backup_sa.db'" > Pros: Single command no interaction with ovsdb-server > Cons: Is the filename (as placed in the sos archive tree) going to be that > long?. We have to cleanup the generated files. > > > Dumitru: does the "cluster-to-standalone" command destroy any information > that might be necessary for debugging? Is it better to keep both (clustered > and standalone) dbs? Good point, it does, and as a matter of fact I think "ovsdb-client backup" does that too. We'll be losing all RAFT logs (which are stored as json records in the DB file in clustered format). Unfortunately, if we want to have all debugging info we'd need both the original DB file (clustered format, non-compacted) for full logs and the standalone, compacted DB file for simplified parsing by external tools. OK, if the DBs are small, then I think the #c1 / https://github.com/amorenoz/sos/commit/ec6cfaa1a824e04ab62b1fee4c69e1674f874174 seems a good approach. Adrián, let me know if I shall do a PR per the commit, or if you will do it by yourself (either works fine for me). (In reply to Pavel Moravec from comment #9) > OK, if the DBs are small, then I think the #c1 / > https://github.com/amorenoz/sos/commit/ > ec6cfaa1a824e04ab62b1fee4c69e1674f874174 seems a good approach. > I guess the SB can get pretty big. Dumitru has worked on scale tests, do you have an order of magnitude? Being in json format, gzip/xz will do a decent job compacting them. As an example, on one of my sample databases I got the following figures: uncompacted: 5.9M (820K gzipped) compacted: 74K (13K gzipped) > Adrián, let me know if I shall do a PR per the commit, or if you will do it > by yourself (either works fine for me). I'll send the PR, no worries. (In reply to Adrián Moreno from comment #10) > (In reply to Pavel Moravec from comment #9) > > OK, if the DBs are small, then I think the #c1 / > > https://github.com/amorenoz/sos/commit/ > > ec6cfaa1a824e04ab62b1fee4c69e1674f874174 seems a good approach. > > > > I guess the SB can get pretty big. Dumitru has worked on scale tests, do you > have an order of magnitude? > > Being in json format, gzip/xz will do a decent job compacting them. > > As an example, on one of my sample databases I got the following figures: > > uncompacted: 5.9M (820K gzipped) > compacted: 74K (13K gzipped) They can be even larger, e.g., from an OpenShift scale test: uncompacted: 319M (66M gzipped) compacted: 131M (27M gzipped) Huh, such a big dump being collected twice..? I would prefer one collection then - recall it must take some time to generate the dump/backup, to pack it (no need to pack it "manually" as sosreport will tar.gz it), to transfer among computers, unpack,.. Could you please provide: - a sequence of commands in bash (including podman exec .. stuff) how to collect a DB snapshot/dump/export and store it in some file - what commands from the current sosreport will be obsolete (as they collect the DB backup now which would be duplicate) or optionally the PR (that would need proper exec_cmd + add_copy_spec "orchestration" where I am offering the help, knowing sos API). (In reply to Pavel Moravec from comment #12) > Huh, such a big dump being collected twice..? I would prefer one collection > then - recall it must take some time to generate the dump/backup, to pack it > (no need to pack it "manually" as sosreport will tar.gz it), to transfer > among computers, unpack,.. > > Could you please provide: > - a sequence of commands in bash (including podman exec .. stuff) how to > collect a DB snapshot/dump/export and store it in some file > - what commands from the current sosreport will be obsolete (as they collect > the DB backup now which would be duplicate) > > or optionally the PR (that would need proper exec_cmd + add_copy_spec > "orchestration" where I am offering the help, knowing sos API). The sequence could be: for db in ovnnb.db ovnsb.db; do cp /var/lib/openvswitch/$db /var/lib/openvswitch/backup_$db podman exec -it $CONTAINER ovsdb-tool compact /var/lib/openvswitch/backup_$db mv /var/lib/openvswitch/backup_$db /tmp/${SOS_TREE)/var/lib/openvswitch/$db done (this assumes /var/lib/openvswitch is volume mounted inside the container, which seems to be the same assumption made in current implementation) If you point me to some available documentation about the sos API I can try implement this :) Oh, wait, ovsdb-tool compact does not support clustered DBs. So, either we make it standalone (and loose raft information), we copy it uncompacted (and quite big) or we interact the ovsdb-server (via "appctl ovsdb/compact" or "ovsdb-client backup"). (In reply to Adrián Moreno from comment #14) > Oh, wait, ovsdb-tool compact does not support clustered DBs. > > So, either we make it standalone (and loose raft information), we copy it > uncompacted (and quite big) or we interact the ovsdb-server (via "appctl > ovsdb/compact" or "ovsdb-client backup"). OK, looking forward for specific requirement / specification, then. We discussed the same in must-gather scripts: https://github.com/openshift/must-gather/pull/245 Faced with the same problem we ended up just copying the full db file (which is the current behavior). If @dceara I think we can reject this BZ and figure out a way to work around the parsing issues elsewhere. (In reply to Adrián Moreno from comment #16) > We discussed the same in must-gather scripts: > https://github.com/openshift/must-gather/pull/245 > > Faced with the same problem we ended up just copying the full db file (which > is the current behavior). > > If @dceara I think we can reject this BZ and figure out a way to > work around the parsing issues elsewhere. Unfortunately it looks like this is the way to go for now. OK, so I am closing this BZ as WORKSFORME. Let me know if something else is required. |