Fedora Account System
Red Hat Associate
Red Hat Customer
Created attachment 315962 [details] Index tables after loading data The attached patch makes yum-metadata-parser index SQLite tables after loading them - inserting into indexed tables is slower than inserting to unindexed ones and later indexing them. The diff is a bit hard to read, but it just moves index creation out of create table functions into ones of their own, there are no other changes. This provides a nice speedup, below are results on my x86_64 AMD64 3200+, 2G RAM, F-9 box when doing createrepos on a couple of days old Rawhide i386 repo, taken from createrepo -v, customized so that it additionally outputs timestamps for the SQLite db operations only (w/o bzip2 and other things done after the db is created). Values below are Xs / Ys so that X includes the bzip2 and friends part of createrepo, Y only the db operations part. Unpatched: - other: 8s / 2s - filelists: 48s / 32s - primary: 17s / 9s = total: 73s / 43s Patched: - other: 8s / 2s (no changes) - filelists: 29s / 12s (much faster) - primary: 14s / 6s (somewhat faster) = total: 51s / 20s (quite a bit faster) As an additional bonus, the patched version creates somewhat smaller databases: Unpatched: 11M repodata/filelists.sqlite.bz2 3.7M repodata/other.sqlite.bz2 7.0M repodata/primary.sqlite.bz2 Patched: 11M repodata/filelists.sqlite.bz2 3.5M repodata/other.sqlite.bz2 6.4M repodata/primary.sqlite.bz2
Created attachment 315963 [details] Create indexes with IF NOT EXISTS Companion patch to be applied on top of the previous one: I'm not sure if there's a scenario in which indexes would already exist after applying the previous patch, but I suppose creating them with IF NOT EXISTS does not hurt in any case.
Just to make sure I understand this - does this mean the indexes will get created on the client machine? Doesn't that mean we're offloading all the time onto the user?
Or are you just making the indexes in the repo-side AFTER all the other data has been inserted first?
(In reply to comment #3) > Or are you just making the indexes in the repo-side AFTER all the other data > has been inserted first? Yes.
Committed to upstream, thanks.