Git Object Model: Blobs, Trees, Commits, Tags, and the Object Database: Concepts, Architecture, and Mental Model
Build a beginner-first mental model of Git's content-addressed object database: blobs, trees, commits, annotated tags, parent edges, movable refs, and object-format-aware identifiers.
Learning objectives
- Explain content-addressable storage and the role of object IDs without assuming SHA-1-only repositories.
- Distinguish blob, tree, commit, and annotated-tag objects and the relationships among them.
- Explain commit parent edges as the basis of Git's directed acyclic history graph.
- Separate immutable objects from movable refs such as branches and tags.
- Use read-only Git plumbing to inspect object type, size, content, tree edges, and object format.
1. Why Git needs an object database
Chapter 01 treated Git history as a graph of snapshots. Chapter 02 made the Git executable and configuration inputs predictable. This chapter now answers the next question: what is a snapshot made of inside Git?
If a repository merely stored one complete physical directory copy for every commit, large histories would become unnecessarily expensive and many later Git behaviors—branching, deduplication, recovery, tagging, and object transfer—would be hard to explain. Git instead stores content as typed objects addressed by object IDs, then uses references such as branch names to point into the resulting graph.
2. Content-addressable storage — identify data by what it contains
An object ID (OID) is the hash-derived identifier Git uses for an object in a repository's object format. For a blob, Git hashes a typed header together with the content Git will store. If the stored content is identical under the same object format, the resulting blob OID is identical; if the stored content changes, the OID changes.
This gives Git two important properties:
- Addressing: another object can refer to this object by OID.
- Integrity checking: Git can detect when stored bytes do not match the identifier expected for them.
3. The four object types you need first
| Object | What it represents | Important references inside it |
|---|---|---|
| blob | File content as stored by Git | None; a blob does not carry its filename |
| tree | A directory-like snapshot: names, modes, and object IDs | Blobs and nested trees |
| commit | A recorded project snapshot plus history metadata | One root tree and zero or more parent commits |
| tag | An annotated tag object with tagger/message metadata | A target object, commonly a commit |
A normal file's name is therefore not “inside the blob.” The tree entry provides the name and points at the blob. The same blob can appear under different names or in multiple commits.
4. Mental model — refs point to objects; objects point to other objects
flowchart TD B[refs/heads/trunk] --> C2[commit C2] C2 -->|tree| T2[tree T2] C2 -->|parent| C1[commit C1] C1 -->|tree| T1[tree T1] T2 -->|stable.txt| S[blob S] T2 -->|changing.txt| X2[blob X2] T1 -->|stable.txt| S T1 -->|changing.txt| X1[blob X1] R[refs/tags/v1.0] --> A[tag object A] A -->|object| C2
The two tree objects reuse the same stable.txt blob
because that stored content did not change. They point to different
blobs for changing.txt. The branch is not a commit
object—it is a movable reference whose current value names
C2. The annotated tag reference points to a separate
tag object, which in turn names its target commit.
5. Commit parents create the directed history graph
A root commit has no parent. A normal follow-up commit usually has one parent. A merge commit can have multiple parents. Parent links point backward toward earlier commits, so ordinary Git history is modeled as a directed acyclic graph (DAG): following parent edges takes you toward ancestors rather than creating loops.
git cat-file -p HEAD
On an ordinary commit, the pretty-printed object begins with a
tree line, then a parent line when the
commit is not a root, followed by author/committer metadata and the
commit message. Those parent OIDs are the actual historical edges;
timestamps are metadata and do not replace ancestry.
6. Immutable object identity versus movable names
Once a valid object exists, its OID identifies that exact object content under the repository's object format. Git does not “edit commit C1 in place.” Operations that appear to alter an old commit create replacement objects with different OIDs and move refs to the new graph.
References are the movable naming layer. Common examples include
refs/heads/trunk, refs/tags/v1.0, and
remote-tracking refs. A branch name can move from one commit OID to
another while both commit objects may remain in the object database
for some time.
7. Object-format awareness — stop assuming 40 hexadecimal characters
Current Git repositories can use different storage object formats. Discover the repository's format instead of inferring it from a string length:
git rev-parse --show-object-format=storage
git rev-parse --verify HEAD
git rev-parse --short HEAD
In a SHA-1 repository the full OID is traditionally 40 hexadecimal
characters; a SHA-256 repository uses a different full
representation. Scripts should treat an OID as an opaque identifier
returned by Git, not as a fixed-width field invented by the script.
When a human-friendly abbreviation is needed, let Git compute a
unique abbreviation with --short.
8. Read-only inspection commands are your microscope
git rev-parse --show-object-format=storage
git rev-parse HEAD
git cat-file -t HEAD
git cat-file -s HEAD
git cat-file -p HEAD
git ls-tree HEAD
git show --no-patch --format=fuller HEAD
These commands let you ask Git to resolve names and inspect existing objects. None needs to write a new object. This chapter deliberately begins with them so plumbing does not become synonymous with dangerous repository surgery.
9. Why this matters in DevOps
Build provenance often pins an immutable commit OID. Artifact caches can reuse results keyed by source identity. Recovery work depends on knowing whether an object still exists and whether any ref/reflog can reach it. Release tags add another naming/provenance layer. Troubleshooting corrupted or incomplete repositories requires distinguishing object storage from references and working-tree files. All of those workflows become safer when the object graph is explicit.
10. Micro-lab — inspect one existing commit without changing anything
Use a disposable repository from an earlier chapter or create a fresh repository with one commit. Then:
git status --short --branch
git rev-parse --show-object-format=storage
git rev-parse HEAD
git cat-file -t HEAD
git cat-file -p HEAD
git rev-parse HEAD^{tree}
git ls-tree HEAD
Predict before running the last two commands: will
HEAD^{tree} resolve to the same OID as
HEAD? It should not. One names a commit object; the
other peels that commit to its root tree object.
11. Knowledge check
Question 1. Where is a tracked filename stored: in the blob or in a tree entry?
Question 2. Is a branch such as trunk itself a
commit object?
Question 3. Why can two commits reuse the same blob object?
Question 4. Does knowing a commit OID prove who authored it?
Question 5. Why should automation call
git rev-parse --show-object-format instead of
assuming 40 hex characters?
12. Summary
Git history is a graph built from typed, content-addressed objects. Blobs hold content, trees provide names and directory structure, commits point to trees and parent commits, and annotated tag objects point to targets. Refs are movable names over that immutable object layer.
Authoritative references
Git glossary
Git repository layout
git-cat-file
git-rev-parse
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.