Chapter 22Lesson 01~125 minutes

Commit-Graphs, Multi-Pack Indexes, Maintenance, Garbage Collection, and Performance: Concepts, Architecture, and Mental Model

Build the mental model for loose objects, packfiles and deltas, reachability/cruft retention, commit-graph generation data, multi-pack indexes, maintenance strategies, and evidence-driven Git performance.

Object storageCommit-graphMIDXMaintenance

Learning objectives

  • Separate working-tree, commit-graph, object-lookup, transfer, and maintenance performance layers.
  • Explain loose objects, packfiles, deltas, reachability, and cruft retention.
  • Explain how commit-graphs and generation data accelerate history traversal.
  • Explain multi-pack indexes as lookup structures over multiple packfiles.
  • Treat maintenance/GC as measured housekeeping with recovery-aware retention.

1. Performance problems begin with the question “which layer is slow?”

Chapter 21 treated the object database as evidence that must not be destroyed casually. This chapter asks a different operational question: as a repository grows, how does Git keep object lookup, history traversal, and routine housekeeping efficient without turning maintenance into a risky manual ritual?

A slow git status, a slow history query, a repository with thousands of loose objects, and a repository with hundreds of packfiles are different problems. They touch different layers. Before tuning anything, identify whether the cost comes from the working tree/index, commit-graph traversal, object lookup, storage layout, network transfer, or a maintenance job competing with foreground work.

2. Start with read-only measurements

git --version
git status --short --branch
git rev-list --count --all
git count-objects -vH
git config --show-origin --get core.commitGraph
git config --show-origin --get core.multiPackIndex
git config --show-origin --get maintenance.strategy
git config --show-origin --get gc.auto

Expected observations: the history count tells you scale in commits; count-objects separates loose objects from packed objects and reports pack count/storage; configuration queries show whether explicit overrides exist. No output from a config query means “no explicit value at this scope,” not necessarily “feature disabled,” because several features have documented defaults.

3. Git starts cheaply: new objects are usually loose

When Git creates a new blob, tree, commit, or tag object, it can store that object as an individual compressed file in the object database. This is called a loose object. Loose creation is responsive because Git does not need to reorganize the entire repository each time you commit.

That design deliberately defers optimization. A repository that keeps accumulating data eventually benefits from grouping many objects into packfiles.

Object-storage lifecycle
flowchart TD
A[New loose objects] -->|repack / maintenance| B[Packfiles]
B --> C[Pack indexes]
C --> D[Optional multi-pack index]
A --> E[Reachability decision]
B --> E
E -->|reachable| K[Retained repository data]
E -->|unreachable but within retention| U[Cruft / grace-period retention]
U -->|expiry policy later| X[Eligible for removal]

The first arrow is storage optimization, not history rewriting. The reachability branch is a retention decision: whether an object is reachable from refs, reflogs, the index, and other roots determines whether it is protected as active/recovery data.

4. Packfiles reduce object-store overhead and can delta-compress similar content

A packfile stores many Git objects in one file, with a companion index that maps object IDs to pack locations. Packing reduces filesystem overhead and can store some objects as deltas against similar objects. A delta is a compact representation describing how to reconstruct one object from another; it does not change the logical object ID or content.

More aggressive delta search can consume more CPU and memory. That is why “maximize compression” is not automatically the best operational choice.

5. Reachability is the boundary between active history and potential garbage

An object is reachable when Git can traverse to it from a root such as a branch/tag ref or another protected repository state. Unreachable objects can remain useful for recovery—for example after an amend, rebase, or deleted temporary ref—so Git normally applies retention/grace rules rather than deleting them immediately.

Garbage collection therefore does more than “delete old files.” It evaluates storage, refs/reflogs, reachability, retention windows, pack layout, and auxiliary structures.

6. Cruft packs keep unreachable objects compact while preserving expiry information

Modern Git can store unreachable objects in a separate cruft pack with per-object modification-time metadata. This avoids keeping all unreachable objects loose merely to remember their ages. Current git gc uses cruft packs by default when expiring unreachable objects.

Recovery implication: an unreachable object is not “useless.” It may be the only surviving copy of lost work. Aggressive expiration shortens recovery opportunities and can also race with concurrent writers. Never begin performance troubleshooting by expiring recovery evidence.

7. The commit-graph accelerates graph questions without replacing commits

A commit-graph file is an auxiliary serialized structure containing graph information derived from commit objects. Git can use it to answer reachability and ancestry questions more efficiently. The commits remain the source of truth; deleting/rebuilding a commit-graph does not rewrite commit objects.

git commit-graph write --reachable
git commit-graph verify

The write operation creates/refreshes the acceleration structure. verify checks it against the object database.

8. Generation data gives Git useful bounds during commit traversal

Commit-graph files can store a generation number: derived ordering information that helps Git avoid walking parts of the history graph that cannot affect an answer. Current Git defaults to generation-data version 2, which uses corrected commit dates. This is graph acceleration—not a claim that author or committer wall-clock timestamps are perfectly ordered.

9. Changed-path Bloom filters can accelerate path history

With --changed-paths or the corresponding configuration, a commit-graph can include probabilistic Bloom-filter data describing paths changed by commits. These filters help history queries such as git log -- path skip commits that definitely did not touch the path. A Bloom filter may produce a false positive, but not the false-negative behavior that would silently omit a relevant commit.

Computing filters costs time and storage, so the tradeoff matters most when path-history queries are frequent and repositories are large.

10. A multi-pack index is one lookup layer over many packfiles

If the object directory contains many packs, Git would otherwise need to consult multiple individual pack indexes. A multi-pack index (MIDX) provides one index over objects in multiple packfiles. It does not by itself merge the packfiles or change object IDs.

Pack lookup with a MIDX
flowchart TD
Q[Object ID lookup] --> M[Multi-pack index]
M --> P1[Pack A]
M --> P2[Pack B]
M --> P3[Pack C]
P1 --> O[Object bytes]
P2 --> O
P3 --> O

The MIDX tells Git which pack contains the desired object. Separate multi-pack-index repack/expire operations can later consolidate packs or remove packs no longer referenced by the MIDX.

11. git maintenance coordinates optimization tasks

Git intentionally keeps foreground commands responsive and defers full-repository optimizations. git maintenance provides task-oriented maintenance such as commit-graph updates, loose-object packing, prefetch, incremental repacking, pack-refs, and GC-related work.

Current Git distinguishes strategies. The current manual recommends the geometric strategy for large repositories for manual maintenance, while incremental is designed around smaller scheduled tasks that avoid data deletion. Exact strategy defaults have evolved, so production runbooks should record the deployed Git version.

12. Garbage collection is one maintenance approach, not the universal tuning command

git gc performs several housekeeping operations and can repack data, prune according to expiry policy, expire some metadata, pack refs, and update auxiliary structures. Current Git documentation says manual git gc is usually unnecessary in ordinary porcelain workflows and advises using the maintenance framework when combining background maintenance with other tasks.

For large repositories, repeatedly forcing full GC can consume much more CPU/I/O than incremental tasks while adding little user-visible improvement.

13. Measure the correct layer before choosing a feature

Symptom Likely layer Useful evidence Possible feature
status scans many files working tree/index tracked/untracked counts, repeated status timing untracked cache / FSMonitor where supported
ancestry/log graph walks are expensive commit graph commit count/query timings commit-graph
many packfiles increase object lookup cost object database count-objects -vH MIDX / incremental repack
many loose objects object database loose-object count loose-objects task / repack
clone/fetch transfers too much network/object transfer transfer size/history shape partial/shallow strategies from Chapter 14

14. DevOps connection — repository latency becomes pipeline latency

History queries, checkout/status operations, fetches, and repository maintenance all appear inside developer loops and CI jobs. A disciplined operating model captures a baseline, selects the narrowest acceleration, verifies integrity afterward, and schedules expensive work so that maintenance does not compete with deployment-critical automation.

15. Knowledge check

Question 1. Does writing a commit-graph rewrite commit objects?

Question 2. What problem does a multi-pack index solve?

Question 3. Why can unreachable objects still matter?

Question 4. Which layer should you investigate first when git status is slow in a huge working tree?

Question 5. Why is aggressive pruning a bad first performance experiment?

16. Summary

Git performance is layered. Loose objects become packs; pack indexes and MIDX accelerate object lookup; commit-graphs accelerate ancestry/history traversal; maintenance coordinates safe optimization; and reachability/cruft retention protect recovery. Measure first, optimize the right layer, and never confuse “smaller object store” with “safer repository.”

Next

Build and measure a repository maintenance lab

Lesson 2 creates multiple packs, writes/verifies a MIDX and commit-graph, inspects pack statistics, registers isolated maintenance configuration, and runs safe maintenance/GC operations on generated data.

Authoritative references

 git-maintenance
 git-commit-graph
 git-multi-pack-index
 git-gc
 git-count-objects

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.