Git Object Model: Blobs, Trees, Commits, Tags, and the Object Database: Diagnostics, Failure Modes, Security, and Performance
Diagnose object-model failures safely: filename/blob confusion, snapshot duplication myths, ref/object confusion, invalid or ambiguous object names, trust misconceptions, and dangerous manual object-store edits.
Learning objectives
- Apply an evidence-first diagnostic sequence to object/ref problems.
- Distinguish paths/tree entries from blob identity and refs from target objects.
- Verify object names safely and diagnose invalid or ambiguous abbreviations.
- Use git fsck and related inspection commands without equating dangling objects with corruption.
- Recognize why direct object-file editing and aggressive pruning are inappropriate first responses.
1. Evidence-first object diagnostics
- Preserve evidence: exact name/OID, command, error, Git version, repository path.
-
Inspect refs/history:
status,show-ref,log,rev-parse. -
Inspect objects through Git:
cat-file,ls-tree,fsckwhen integrity/connectivity is genuinely in question. - Identify the layer: working-tree path, tree entry, blob, commit, tag, ref, or physical storage.
- Choose the least destructive correction and verify again.
Most object-model confusion is solved by asking Git what a name
resolves to. Manual editing of .git/objects is not a
diagnostic shortcut.
2. Failure mode — treating a filename as blob identity
Create one commit with the same content under two names:
git rev-parse HEAD:path/to/file
git ls-tree -r HEAD -- path/to/file
A tree entry owns the path name. A blob is content. Renaming a file without changing stored content can result in a different tree but the same blob OID. Git may later detect renames by comparing content; there is no permanent “this blob's filename” field.
3. Failure mode — assuming each commit physically duplicates every file
git rev-parse HEAD~1:stable.txt
git rev-parse HEAD:stable.txt
git rev-parse HEAD~1^{tree}
git rev-parse HEAD^{tree}
If the file is unchanged, the blob OID can be reused even though the commit and often root-tree OIDs differ. Git's pack storage may later compress representations further, so “commit = copied folder” is wrong both logically and physically.
4. Failure mode — treating a branch as an object
git show-ref --heads
git rev-parse --symbolic-full-name HEAD
git rev-parse HEAD
git cat-file -t HEAD
show-ref pairs ref names with OIDs.
cat-file -t HEAD reports the type of the resolved
target object, normally commit; it does not prove “HEAD
is a commit file.” The name is resolved first.
5. Intentionally broken example — an object name that is too short or invalid
Git accepts sufficiently unique hexadecimal prefixes, but not arbitrary strings or ambiguous/nonexistent prefixes. In a disposable repository try:
git rev-parse --verify deadbeef^{object}
git cat-file -t deadbeef
Unless your repository happens to contain an object matching that prefix, Git reports that the revision/object is unknown or invalid. The repair is not to fabricate a longer string. Resolve from a trustworthy name or ask Git for a unique abbreviation:
git rev-parse --verify HEAD
git rev-parse --short=12 HEAD
For untrusted input, also use --end-of-options with
rev-parse --verify so a malicious name is not parsed as
an option.
6. Failure mode — treating OIDs as signatures
A commit's object identity covers the commit object's bytes, which include tree/parent IDs, author/committer metadata, and message. That does not mean the named author cryptographically signed it. Anyone who can create a commit can place arbitrary author metadata into a new commit object.
Later security lessons distinguish object integrity, cryptographic signatures, trusted signer identity, server authorization, and review/approval policy.
7. Failure mode — editing .git/objects manually
Use supported commands:
git cat-file -e HEAD^{object}
git cat-file -t HEAD
git cat-file -p HEAD
git fsck --full
git fsck checks object connectivity/validity and is
appropriate when corruption is genuinely suspected. Interpret
dangling/unreachable reports carefully—they are not automatically
corruption.
8. Failure mode — confusing “unreachable” with “missing”
An unreachable object may still exist but have no current ref path leading to it. A missing object is absent when another reachable object/reference requires it. These lead to very different recovery strategies.
git fsck --full
git reflog
git show-ref
Preserve the repository before cleanup if recovery matters. Do not
immediately prune unreachable objects just because
fsck lists them.
9. Red-zone operations that reduce recovery options
git gc --prune=now, direct
object deletion, destructive resets, force updates to shared refs,
or history-rewrite tools. Preserve refs/reflogs/objects until you
understand the incident.
10. Security analysis that belongs to the object model
- Integrity: object IDs let Git detect mismatched/corrupt object contents.
- Authenticity: requires signatures or another trusted identity mechanism; an author field alone is not proof.
- Authorization: a server decides who may read/update refs; an OID does not grant permission.
- Confidentiality: content placed in Git history may persist in objects even after a visible ref moves; secret response requires credential rotation plus coordinated history remediation.
11. Performance analysis that belongs here
Do not count loose files and conclude “one file per object forever.” Git can pack objects and use deltas to reduce storage and improve access. Conversely, a huge number of loose objects can be a maintenance signal. Use:
git count-objects -vH
Chapter 22 teaches maintenance tuning. This lesson's key point is that physical storage representation is not the same thing as logical object identity.
12. Symptom → object-model question
| Symptom | Ask first | Read-only evidence |
|---|---|---|
| “This file's hash changed after rename” | Are you hashing path-filtered working content or stored blob? | rev-parse REV:path, ls-tree |
| “Branch object disappeared” | Was the ref moved/deleted while commit object remains? |
show-ref, reflog,
cat-file
|
| “OID doesn't resolve” | Invalid/ambiguous prefix or missing object? |
rev-parse --verify, cat-file -e
|
| “Repository is large” | Loose objects, packs, history, binaries, or working tree? | count-objects -vH |
| “Commit says Alice authored it” | What establishes trust? | signature/platform/policy evidence, not OID alone |
13. Knowledge check
Question 1. Why can a blob be reused under two different filenames?
Question 2. git fsck reports a dangling blob. Is
the repository necessarily corrupt?
Question 3. What is wrong with manually changing bytes under
.git/objects?
Question 4. Why use
rev-parse --verify --end-of-options for untrusted
object-name input?
Question 5. Does an author name in a commit object prove that person created the commit?
14. Summary
Paths live in trees, branches live in refs, logical objects can outlive refs, abbreviations require uniqueness, and physical storage is an implementation layer. Diagnose through Git's object/ref commands and preserve recovery evidence before maintenance or cleanup.
Authoritative references
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.