Commit-Graphs, Multi-Pack Indexes, Maintenance, Garbage Collection, and Performance: Concepts, Architecture, and Mental Model
Build the mental model for loose objects, packfiles and deltas, reachability/cruft retention, commit-graph generation data, multi-pack indexes, maintenance strategies, and evidence-driven Git performance.
Learning objectives
- Separate working-tree, commit-graph, object-lookup, transfer, and maintenance performance layers.
- Explain loose objects, packfiles, deltas, reachability, and cruft retention.
- Explain how commit-graphs and generation data accelerate history traversal.
- Explain multi-pack indexes as lookup structures over multiple packfiles.
- Treat maintenance/GC as measured housekeeping with recovery-aware retention.
1. Performance problems begin with the question “which layer is slow?”
Chapter 21 treated the object database as evidence that must not be destroyed casually. This chapter asks a different operational question: as a repository grows, how does Git keep object lookup, history traversal, and routine housekeeping efficient without turning maintenance into a risky manual ritual?
A slow git status, a slow history query, a repository
with thousands of loose objects, and a repository with hundreds of
packfiles are different problems. They touch different layers.
Before tuning anything, identify whether the cost comes from the
working tree/index, commit-graph traversal, object lookup, storage
layout, network transfer, or a maintenance job competing with
foreground work.
2. Start with read-only measurements
git --version
git status --short --branch
git rev-list --count --all
git count-objects -vH
git config --show-origin --get core.commitGraph
git config --show-origin --get core.multiPackIndex
git config --show-origin --get maintenance.strategy
git config --show-origin --get gc.auto
Expected observations: the history count tells you
scale in commits; count-objects separates loose objects
from packed objects and reports pack count/storage; configuration
queries show whether explicit overrides exist. No output from a
config query means “no explicit value at this scope,” not
necessarily “feature disabled,” because several features have
documented defaults.
3. Git starts cheaply: new objects are usually loose
When Git creates a new blob, tree, commit, or tag object, it can store that object as an individual compressed file in the object database. This is called a loose object. Loose creation is responsive because Git does not need to reorganize the entire repository each time you commit.
That design deliberately defers optimization. A repository that keeps accumulating data eventually benefits from grouping many objects into packfiles.
flowchart TD A[New loose objects] -->|repack / maintenance| B[Packfiles] B --> C[Pack indexes] C --> D[Optional multi-pack index] A --> E[Reachability decision] B --> E E -->|reachable| K[Retained repository data] E -->|unreachable but within retention| U[Cruft / grace-period retention] U -->|expiry policy later| X[Eligible for removal]
The first arrow is storage optimization, not history rewriting. The reachability branch is a retention decision: whether an object is reachable from refs, reflogs, the index, and other roots determines whether it is protected as active/recovery data.
4. Packfiles reduce object-store overhead and can delta-compress similar content
A packfile stores many Git objects in one file, with a companion index that maps object IDs to pack locations. Packing reduces filesystem overhead and can store some objects as deltas against similar objects. A delta is a compact representation describing how to reconstruct one object from another; it does not change the logical object ID or content.
More aggressive delta search can consume more CPU and memory. That is why “maximize compression” is not automatically the best operational choice.
5. Reachability is the boundary between active history and potential garbage
An object is reachable when Git can traverse to it from a root such as a branch/tag ref or another protected repository state. Unreachable objects can remain useful for recovery—for example after an amend, rebase, or deleted temporary ref—so Git normally applies retention/grace rules rather than deleting them immediately.
Garbage collection therefore does more than “delete old files.” It evaluates storage, refs/reflogs, reachability, retention windows, pack layout, and auxiliary structures.
6. Cruft packs keep unreachable objects compact while preserving expiry information
Modern Git can store unreachable objects in a separate
cruft pack with per-object modification-time
metadata. This avoids keeping all unreachable objects loose merely
to remember their ages. Current git gc uses cruft packs
by default when expiring unreachable objects.
7. The commit-graph accelerates graph questions without replacing commits
A commit-graph file is an auxiliary serialized structure containing graph information derived from commit objects. Git can use it to answer reachability and ancestry questions more efficiently. The commits remain the source of truth; deleting/rebuilding a commit-graph does not rewrite commit objects.
git commit-graph write --reachable
git commit-graph verify
The write operation creates/refreshes the acceleration structure.
verify checks it against the object database.
8. Generation data gives Git useful bounds during commit traversal
Commit-graph files can store a generation number: derived ordering information that helps Git avoid walking parts of the history graph that cannot affect an answer. Current Git defaults to generation-data version 2, which uses corrected commit dates. This is graph acceleration—not a claim that author or committer wall-clock timestamps are perfectly ordered.
9. Changed-path Bloom filters can accelerate path history
With --changed-paths or the corresponding
configuration, a commit-graph can include probabilistic Bloom-filter
data describing paths changed by commits. These filters help history
queries such as git log -- path skip commits that
definitely did not touch the path. A Bloom filter may produce a
false positive, but not the false-negative behavior that would
silently omit a relevant commit.
Computing filters costs time and storage, so the tradeoff matters most when path-history queries are frequent and repositories are large.
10. A multi-pack index is one lookup layer over many packfiles
If the object directory contains many packs, Git would otherwise need to consult multiple individual pack indexes. A multi-pack index (MIDX) provides one index over objects in multiple packfiles. It does not by itself merge the packfiles or change object IDs.
flowchart TD Q[Object ID lookup] --> M[Multi-pack index] M --> P1[Pack A] M --> P2[Pack B] M --> P3[Pack C] P1 --> O[Object bytes] P2 --> O P3 --> O
The MIDX tells Git which pack contains the desired object. Separate
multi-pack-index repack/expire operations
can later consolidate packs or remove packs no longer referenced by
the MIDX.
11. git maintenance coordinates optimization tasks
Git intentionally keeps foreground commands responsive and defers
full-repository optimizations. git maintenance provides
task-oriented maintenance such as commit-graph updates, loose-object
packing, prefetch, incremental repacking, pack-refs, and GC-related
work.
Current Git distinguishes strategies. The current manual recommends
the geometric strategy for large repositories for
manual maintenance, while incremental is designed
around smaller scheduled tasks that avoid data deletion. Exact
strategy defaults have evolved, so production runbooks should record
the deployed Git version.
12. Garbage collection is one maintenance approach, not the universal tuning command
git gc performs several housekeeping operations and can
repack data, prune according to expiry policy, expire some metadata,
pack refs, and update auxiliary structures. Current Git
documentation says manual git gc is usually unnecessary
in ordinary porcelain workflows and advises using the maintenance
framework when combining background maintenance with other tasks.
For large repositories, repeatedly forcing full GC can consume much more CPU/I/O than incremental tasks while adding little user-visible improvement.
13. Measure the correct layer before choosing a feature
| Symptom | Likely layer | Useful evidence | Possible feature |
|---|---|---|---|
status scans many files |
working tree/index | tracked/untracked counts, repeated status timing | untracked cache / FSMonitor where supported |
| ancestry/log graph walks are expensive | commit graph | commit count/query timings | commit-graph |
| many packfiles increase object lookup cost | object database | count-objects -vH |
MIDX / incremental repack |
| many loose objects | object database | loose-object count | loose-objects task / repack |
| clone/fetch transfers too much | network/object transfer | transfer size/history shape | partial/shallow strategies from Chapter 14 |
14. DevOps connection — repository latency becomes pipeline latency
History queries, checkout/status operations, fetches, and repository maintenance all appear inside developer loops and CI jobs. A disciplined operating model captures a baseline, selects the narrowest acceleration, verifies integrity afterward, and schedules expensive work so that maintenance does not compete with deployment-critical automation.
15. Knowledge check
Question 1. Does writing a commit-graph rewrite commit objects?
Question 2. What problem does a multi-pack index solve?
Question 3. Why can unreachable objects still matter?
Question 4. Which layer should you investigate first when
git status is slow in a huge working tree?
Question 5. Why is aggressive pruning a bad first performance experiment?
16. Summary
Git performance is layered. Loose objects become packs; pack indexes and MIDX accelerate object lookup; commit-graphs accelerate ancestry/history traversal; maintenance coordinates safe optimization; and reachability/cruft retention protect recovery. Measure first, optimize the right layer, and never confuse “smaller object store” with “safer repository.”
Authoritative references
git-maintenance
git-commit-graph
git-multi-pack-index
git-gc
git-count-objects
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.