Chapter 22Lesson 04~150 minutes

Commit-Graphs, Multi-Pack Indexes, Maintenance, Garbage Collection, and Performance: Diagnostics, Failure Modes, Security, and Performance

Diagnose wrong-layer tuning, corrupted auxiliary indexes, dangerous pruning, maintenance overlap, misleading benchmarks, and unnecessary large-repository features while preserving recovery evidence.

DiagnosticsRecovery-awareBenchmarkingCorruption

Learning objectives

  • Preserve repository evidence before performance remediation.
  • Distinguish auxiliary commit-graph/MIDX corruption from object corruption.
  • Reject aggressive pruning as a default performance tactic.
  • Coordinate maintenance ownership instead of overlapping independent GC jobs.
  • Benchmark under comparable cache/workload conditions and keep security/trust separate.

1. Diagnose performance without destroying recovery evidence

  1. Preserve evidence: refs, reflogs, HEAD, object counts, pack list, integrity checks, configuration, Git version, and timing method.
  2. Identify the affected layer: worktree/index, commit graph, object lookup, network, or background maintenance.
  3. Measure the smallest representative command.
  4. Apply the least destructive correction.
  5. Verify repository integrity and repeat the same measurement.

2. Capture a diagnostic baseline

git --version
git status --short --branch
git rev-parse HEAD
git show-ref
git reflog -10 --all
git count-objects -vH
git fsck --full
git config --list --show-origin --show-scope

# Auxiliary-structure checks when present:
git commit-graph verify
git multi-pack-index verify

Not every repository has a commit-graph or MIDX. A missing optional acceleration structure is different from a corrupted one.

3. Failure mode — tuning packfiles when status is slow

A team sees a 6-second git status in a working tree containing hundreds of thousands of paths, then runs aggressive repacking. If count-objects already shows a healthy pack layout, object compression is unlikely to fix directory scanning.

git count-objects -vH
time git status --porcelain >/dev/null
git update-index --test-untracked-cache

The correction is to measure working-tree/index costs and consider untracked-cache/FSMonitor or sparse-working-tree choices, not to rewrite storage layout blindly.

4. Intentionally broken example — corrupt a copied/disposable commit-graph and interpret verification

Disposable repository only. This example intentionally corrupts Git metadata. Preserve refs and make a byte-for-byte backup first.
git commit-graph write --reachable
git commit-graph verify

cp .git/objects/info/commit-graph ../commit-graph.backup
printf 'broken graph data\n' > .git/objects/info/commit-graph

git commit-graph verify
echo "verify exit=$?"

Interpretation: a non-zero verify result says the auxiliary commit-graph is invalid. Normal commands that try to read that corrupt graph—including git fsck --full on some Git versions—may also fail before they can prove the underlying commits are healthy. Preserve the broken file, then bypass commit-graph reading only for the source-object cross-check:

git -c core.commitGraph=false fsck --full
git -c core.commitGraph=false log -5 --oneline

If those object-based checks succeed, restore the known-good lab backup or regenerate the acceleration structure according to your incident plan. This does not make the corrupt graph harmless; it isolates which layer failed. Never “repair” a valuable repository by guessing bytes in Git metadata.

5. Verify pack indexes before blaming object corruption

git count-objects -vH
for idx in .git/objects/pack/*.idx; do
  git verify-pack "$idx"
done
git fsck --full

A failing pack/index check is a storage-integrity issue. Preserve the affected files and obtain a trustworthy copy/backup before deleting or repacking evidence.

6. Failure mode — aggressive cleanup destroys recovery opportunities

Red zone — do not run these as performance experiments on valuable repositories: git gc --prune=now, git prune --expire now, forced reflog expiry, or broad history deletion. They can shorten/remove recovery paths and can be dangerous with concurrent writers.

Instead, inspect reachability and reflogs first:

git reflog --all
git fsck --unreachable
git count-objects -vH

If storage cleanup is genuinely required, use normal grace periods or an explicitly reviewed retention policy. Modern Git's cruft-pack design exists partly to keep unreachable objects compact while preserving age information.

7. git gc --aggressive is usually not worth guessing with

Current Git documentation says aggressive GC spends much more time recomputing deltas and is usually not worth using without tailored benchmarks. “Aggressive” means more optimization effort, not “safer” or “more correct.”

# Prefer measurement and normal maintenance.
git count-objects -vH
git maintenance run --task=gc

Even the GC maintenance task should run in a planned window on repositories where foreground latency matters.

8. Failure mode — overlapping maintenance with critical automation

Current git maintenance run takes an object-database lock to prevent two maintenance runs from colliding. The documentation separately warns that ad-hoc git gc does not coordinate through that same maintenance lock. Running both scheduling systems can create unpredictable I/O contention and extra concurrency risk.

Choose one orchestration model. If possible, schedule GC through git maintenance run --task=gc rather than layering independent GC jobs over maintenance.

9. Failure mode — comparing cold-cache “before” with warm-cache “after”

Repeated history/status commands can warm OS page cache, Git's own auxiliary data, antivirus caches, and filesystem metadata. An apparent 5× speedup may disappear after reboot or when the cache state is equalized.

Use multiple repetitions, report median/range, label warm/cold conditions, and keep the repository/workload identical. For CI, measure end-to-end job time as well as individual Git commands.

10. Failure mode — enabling every large-repository feature on a small repository

Auxiliary indexes, background schedulers, FSMonitor processes, and extra configuration have operational cost. If a repository is small and commands already complete below human-perception thresholds, defaults can be the fastest system overall because they minimize complexity and failure modes.

11. Auxiliary structures are rebuildable; commits are authoritative

Commit-graph and MIDX files are derived lookup structures. Verification failure should lead to object/ref preservation and reconstruction from trustworthy object data—not history rewriting. git fsck --full, representative object reads, and ref checks help separate auxiliary corruption from object corruption.

12. Security relevance — maintenance can touch remotes and execute in privileged contexts

The incremental strategy's prefetch task contacts registered remotes. Background schedulers may run with the user's credentials/environment. Therefore do not register untrusted repositories/remotes blindly, do not log credentials in remote URLs, and treat system/service-account maintenance configuration as a security boundary.

13. Performance features do not establish integrity or trust

A valid MIDX or commit-graph only says the auxiliary structure is internally consistent with object data; it does not prove commits are authorized, signed, malware-free, or policy-compliant. Chapter 21's authenticity/authorization/integrity distinctions remain separate from performance correctness.

14. Symptom → layer → least-destructive correction

Symptom Likely layer Next step
Slow status, healthy pack count worktree/index measure tracked/untracked scan; test cache/FSMonitor support
Commit-graph verify fails, fsck clean auxiliary graph preserve evidence then regenerate graph
verify-pack/fsck fails object/pack storage preserve affected files; restore from trustworthy copy
Maintenance overlaps deployments scheduling/I/O consolidate scheduler ownership and move expensive tasks
No measured latency none keep defaults

15. Knowledge check

Question 1. Why does a corrupt commit-graph not automatically mean corrupt commits?

Question 2. Why can git gc --prune=now be dangerous?

Question 3. Why should maintenance and separate scheduled GC not be layered casually?

Question 4. What is wrong with a one-shot before/after benchmark?

Question 5. Does a clean commit-graph/MIDX prove repository authenticity?

16. Summary

Performance incidents need evidence preservation just like security incidents. Diagnose the layer, verify auxiliary structures separately from objects, avoid aggressive expiry, coordinate maintenance ownership, and benchmark consistently. Optimization should never destroy the data needed to understand the problem.

Next

Run a production-style checkpoint with measurable invariants

Lesson 5 generates enough history/packs to observe changes, writes/validates acceleration structures, runs safe maintenance, verifies integrity, and ends with separate developer/CI/server recommendations.

Authoritative references

 git-gc
 git-maintenance
 git-commit-graph
 git-verify-pack
 cruft packs

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.