diff options
author | Derrick Stolee <stolee@gmail.com> | 2018-07-12 15:39:29 -0400 |
---|---|---|
committer | Junio C Hamano <gitster@pobox.com> | 2018-07-20 11:27:28 -0700 |
commit | fe1ed56f5e482507b54a4fb491273f122c5fd9ea (patch) | |
tree | 2f1e2521f84234aaae3eba17b459fe5d26155e56 /bulk-checkin.h | |
parent | midx: read pack names into array (diff) | |
download | tgif-fe1ed56f5e482507b54a4fb491273f122c5fd9ea.tar.xz |
midx: sort and deduplicate objects from packfiles
Before writing a list of objects and their offsets to a multi-pack-index,
we need to collect the list of objects contained in the packfiles. There
may be multiple copies of some objects, so this list must be deduplicated.
It is possible to artificially get into a state where there are many
duplicate copies of objects. That can create high memory pressure if we
are to create a list of all objects before de-duplication. To reduce
this memory pressure without a significant performance drop,
automatically group objects by the first byte of their object id. Use
the IDX fanout tables to group the data, copy to a local array, then
sort.
Copy only the de-duplicated entries. Select the duplicate based on the
most-recent modified time of a packfile containing the object.
Signed-off-by: Derrick Stolee <dstolee@microsoft.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
Diffstat (limited to 'bulk-checkin.h')
0 files changed, 0 insertions, 0 deletions