Commit-Graphs, Multi-Pack Indexes, Maintenance, Garbage Collection, and Performance: Diagnostics, Failure Modes, Security, and Performance
Diagnose wrong-layer tuning, corrupted auxiliary indexes, dangerous pruning, maintenance overlap, misleading benchmarks, and unnecessary large-repository features while preserving recovery evidence.
Learning objectives
- Preserve repository evidence before performance remediation.
- Distinguish auxiliary commit-graph/MIDX corruption from object corruption.
- Reject aggressive pruning as a default performance tactic.
- Coordinate maintenance ownership instead of overlapping independent GC jobs.
- Benchmark under comparable cache/workload conditions and keep security/trust separate.
1. Diagnose performance without destroying recovery evidence
- Preserve evidence: refs, reflogs, HEAD, object counts, pack list, integrity checks, configuration, Git version, and timing method.
- Identify the affected layer: worktree/index, commit graph, object lookup, network, or background maintenance.
- Measure the smallest representative command.
- Apply the least destructive correction.
- Verify repository integrity and repeat the same measurement.
2. Capture a diagnostic baseline
git --version
git status --short --branch
git rev-parse HEAD
git show-ref
git reflog -10 --all
git count-objects -vH
git fsck --full
git config --list --show-origin --show-scope
# Auxiliary-structure checks when present:
git commit-graph verify
git multi-pack-index verify
Not every repository has a commit-graph or MIDX. A missing optional acceleration structure is different from a corrupted one.
3. Failure mode — tuning packfiles when status is slow
A team sees a 6-second git status in a working tree
containing hundreds of thousands of paths, then runs aggressive
repacking. If count-objects already shows a healthy
pack layout, object compression is unlikely to fix directory
scanning.
git count-objects -vH
time git status --porcelain >/dev/null
git update-index --test-untracked-cache
The correction is to measure working-tree/index costs and consider untracked-cache/FSMonitor or sparse-working-tree choices, not to rewrite storage layout blindly.
4. Intentionally broken example — corrupt a copied/disposable commit-graph and interpret verification
git commit-graph write --reachable
git commit-graph verify
cp .git/objects/info/commit-graph ../commit-graph.backup
printf 'broken graph data\n' > .git/objects/info/commit-graph
git commit-graph verify
echo "verify exit=$?"
Interpretation: a non-zero verify result says the
auxiliary commit-graph is invalid. Normal commands that try to read
that corrupt graph—including git fsck --full on some
Git versions—may also fail before they can prove the underlying
commits are healthy. Preserve the broken file, then bypass
commit-graph reading only for the source-object cross-check:
git -c core.commitGraph=false fsck --full
git -c core.commitGraph=false log -5 --oneline
If those object-based checks succeed, restore the known-good lab backup or regenerate the acceleration structure according to your incident plan. This does not make the corrupt graph harmless; it isolates which layer failed. Never “repair” a valuable repository by guessing bytes in Git metadata.
5. Verify pack indexes before blaming object corruption
git count-objects -vH
for idx in .git/objects/pack/*.idx; do
git verify-pack "$idx"
done
git fsck --full
A failing pack/index check is a storage-integrity issue. Preserve the affected files and obtain a trustworthy copy/backup before deleting or repacking evidence.
6. Failure mode — aggressive cleanup destroys recovery opportunities
git gc --prune=now,
git prune --expire now, forced reflog expiry, or broad
history deletion. They can shorten/remove recovery paths and can be
dangerous with concurrent writers.
Instead, inspect reachability and reflogs first:
git reflog --all
git fsck --unreachable
git count-objects -vH
If storage cleanup is genuinely required, use normal grace periods or an explicitly reviewed retention policy. Modern Git's cruft-pack design exists partly to keep unreachable objects compact while preserving age information.
7. git gc --aggressive is usually not worth guessing
with
Current Git documentation says aggressive GC spends much more time recomputing deltas and is usually not worth using without tailored benchmarks. “Aggressive” means more optimization effort, not “safer” or “more correct.”
# Prefer measurement and normal maintenance.
git count-objects -vH
git maintenance run --task=gc
Even the GC maintenance task should run in a planned window on repositories where foreground latency matters.
8. Failure mode — overlapping maintenance with critical automation
Current git maintenance run takes an object-database
lock to prevent two maintenance runs from colliding. The
documentation separately warns that ad-hoc git gc does
not coordinate through that same maintenance lock. Running both
scheduling systems can create unpredictable I/O contention and extra
concurrency risk.
Choose one orchestration model. If possible, schedule GC through
git maintenance run --task=gc rather than layering
independent GC jobs over maintenance.
9. Failure mode — comparing cold-cache “before” with warm-cache “after”
Repeated history/status commands can warm OS page cache, Git's own auxiliary data, antivirus caches, and filesystem metadata. An apparent 5× speedup may disappear after reboot or when the cache state is equalized.
Use multiple repetitions, report median/range, label warm/cold conditions, and keep the repository/workload identical. For CI, measure end-to-end job time as well as individual Git commands.
10. Failure mode — enabling every large-repository feature on a small repository
Auxiliary indexes, background schedulers, FSMonitor processes, and extra configuration have operational cost. If a repository is small and commands already complete below human-perception thresholds, defaults can be the fastest system overall because they minimize complexity and failure modes.
11. Auxiliary structures are rebuildable; commits are authoritative
Commit-graph and MIDX files are derived lookup structures.
Verification failure should lead to object/ref preservation and
reconstruction from trustworthy object data—not history rewriting.
git fsck --full, representative object reads, and ref
checks help separate auxiliary corruption from object corruption.
12. Security relevance — maintenance can touch remotes and execute in privileged contexts
The incremental strategy's prefetch task contacts registered remotes. Background schedulers may run with the user's credentials/environment. Therefore do not register untrusted repositories/remotes blindly, do not log credentials in remote URLs, and treat system/service-account maintenance configuration as a security boundary.
13. Performance features do not establish integrity or trust
A valid MIDX or commit-graph only says the auxiliary structure is internally consistent with object data; it does not prove commits are authorized, signed, malware-free, or policy-compliant. Chapter 21's authenticity/authorization/integrity distinctions remain separate from performance correctness.
14. Symptom → layer → least-destructive correction
| Symptom | Likely layer | Next step |
|---|---|---|
| Slow status, healthy pack count | worktree/index | measure tracked/untracked scan; test cache/FSMonitor support |
| Commit-graph verify fails, fsck clean | auxiliary graph | preserve evidence then regenerate graph |
| verify-pack/fsck fails | object/pack storage | preserve affected files; restore from trustworthy copy |
| Maintenance overlaps deployments | scheduling/I/O | consolidate scheduler ownership and move expensive tasks |
| No measured latency | none | keep defaults |
15. Knowledge check
Question 1. Why does a corrupt commit-graph not automatically mean corrupt commits?
Question 2. Why can git gc --prune=now be
dangerous?
Question 3. Why should maintenance and separate scheduled GC not be layered casually?
Question 4. What is wrong with a one-shot before/after benchmark?
Question 5. Does a clean commit-graph/MIDX prove repository authenticity?
16. Summary
Performance incidents need evidence preservation just like security incidents. Diagnose the layer, verify auxiliary structures separately from objects, avoid aggressive expiry, coordinate maintenance ownership, and benchmark consistently. Optimization should never destroy the data needed to understand the problem.
Authoritative references
git-gc
git-maintenance
git-commit-graph
git-verify-pack
cruft packs
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.