Git Object Model: Blobs, Trees, Commits, Tags, and the Object Database: Configuration, Design Choices, and Tradeoffs
Make object architecture and storage tradeoffs explicit: repository hash format, unique abbreviations, loose versus packed storage, reachability and garbage collection, and integrity versus trust.
Learning objectives
- Discover and reason about a repository's object format instead of assuming 40-character IDs.
- Use Git-computed unique abbreviations rather than fixed string truncation.
- Separate the logical object model from loose/packed physical storage.
- Explain reachability and its consequences for retention and recovery.
- Distinguish hash-based object integrity from authentication, authorization, and approval.
1. Object format is repository architecture, not a display preference
The repository's object format determines how Git computes and represents object IDs for stored objects. Current Git supports SHA-1 and SHA-256 repository object formats, but current documentation states that there is still no direct interoperability between SHA-1 and SHA-256 repositories. Choosing an object format is therefore an architectural compatibility decision made when the repository is created—not something to flip casually in an existing project.
2. Discover the format from Git
git rev-parse --show-object-format=storage
git rev-parse --show-object-format=input
git rev-parse --show-object-format=output
Most scripts only need the storage format plus opaque OIDs returned by Git. Do not infer the algorithm from hard-coded length checks when Git can report the format directly.
3. Optional disposable SHA-256 experiment
Git Bash, Bash, or zsh
mkdir sha256-object-lab
cd sha256-object-lab
git init --object-format=sha256 -b trunk
git config user.name "SHA256 Lab"
git config user.email "sha256@example.invalid"
printf "hello\n" > hello.txt
git add hello.txt
git commit -m "Create SHA-256 snapshot"
git rev-parse --show-object-format=storage
git rev-parse HEAD
cd ..
rm -rf sha256-object-lab
PowerShell file-creation alternative
New-Item -ItemType Directory sha256-object-lab | Out-Null
Set-Location sha256-object-lab
git init --object-format=sha256 -b trunk
git config user.name "SHA256 Lab"
git config user.email "sha256@example.invalid"
Set-Content hello.txt 'hello'
git add hello.txt
git commit -m "Create SHA-256 snapshot"
git rev-parse --show-object-format=storage
git rev-parse HEAD
Set-Location ..
Remove-Item -Recurse -Force sha256-object-lab
The purpose is not “SHA-256 is always better, switch now.” It is to make format-dependence visible and reinforce that scripts should not be wired to SHA-1 widths.
4. Abbreviated object names are conveniences, not new identifiers
git rev-parse HEAD
git rev-parse --short HEAD
git rev-parse --short=12 HEAD
--short asks Git to find a unique prefix with at least
the requested length. The default minimum is influenced by
core.abbrev. A short prefix is useful for humans, logs,
and UI, but long-lived machine contracts should store full OIDs
unless the consuming protocol explicitly defines abbreviation
semantics.
5. Collision-safe display means letting Git check uniqueness
A prefix such as abc1234 is only useful if it resolves
uniquely in the relevant repository. As repositories grow, a
formerly convenient prefix length may be too short. Avoid scripts
that truncate an OID with string slicing:
BAD IDEA (pseudo-code):
short = oid[0:7]
BETTER:
git rev-parse --short=12 <object-name>
The second form asks Git to preserve uniqueness rather than assuming seven or twelve characters are always enough.
6. Loose objects versus packed storage
New objects are often first written as individual (“loose”) object files under the object store. Git can later pack many objects together and delta-compress representations for storage/performance. The logical object model does not change. A commit still refers to its tree by object ID whether the object is loose, in a packfile, or represented internally as a delta.
git count-objects -vH
git cat-file -t HEAD
git cat-file -p HEAD
Use Git commands to inspect logical objects. Do not build
application logic that assumes every OID maps to a loose file at
.git/objects/xx/yyyy…; objects may be packed, borrowed
from alternates, or otherwise stored through supported repository
mechanisms.
7. Reachability is the retention concept that matters
An object is reachable when Git can walk to it from a root such as a reference (and, in recovery contexts, other retention roots such as reflogs). Reachability is why a branch ref can keep an entire ancestry graph alive: the branch points to a commit, that commit points to its tree and parents, those trees point to blobs/trees, and so on.
When a ref moves away from an old commit, the old object is not necessarily deleted immediately. Reflogs and garbage-collection grace periods can preserve recovery opportunities. Later chapters cover recovery in depth.
8. Garbage collection is storage maintenance, not “delete everything not on screen”
Current git gc performs housekeeping such as repacking
and handling unreachable objects. Its pruning behavior uses grace
periods; immediate pruning reduces recovery options and can be
unsafe around concurrent writers.
git gc --prune=now in normal learning
repositories or as a troubleshooting reflex.
It shortens the safety window that later recovery techniques may
rely on.
9. Hash integrity versus identity and trust
If Git expects object OID X and the stored object's bytes no longer hash to X under the repository format, that is an integrity problem. But a valid object can still contain a malicious change, a false author field, or content introduced by someone who lacked organizational approval.
| Question | OID/hash helps? | What else may be needed? |
|---|---|---|
| Did these stored bytes change unexpectedly? | Yes, integrity checking is relevant | Repository health checks / trusted copies |
| Who authored this commit metadata? | No proof by itself | Identity policy, signatures, platform records |
| Was this change approved for production? | No | Review/authorization/CI policy |
| Is this release label immutable by policy? | No | Protected refs and release governance |
10. Configuration scope: what belongs globally and what belongs to the repository
Object format is repository architecture and is recorded by
repository state. New-repository defaults such as
init.defaultObjectFormat can be configured, but a
personal global preference can create compatibility surprises if
teammates expect the conventional format. Treat such defaults like
infrastructure architecture: make them explicit in bootstrap
automation and team documentation.
Human abbreviation preferences such as core.abbrev can
be user-level, but automation should still request an explicit
format/length appropriate to its contract.
11. Git core versus hosting/server capabilities
A hosting platform may or may not support every repository object format or migration path that your local Git binary supports. Local capability is not proof of server compatibility. Before choosing a non-default object format for a shared repository, verify the actual remote server, CI runners, libraries, IDEs, mirroring tools, archival systems, and deployment integrations.
12. Decision table
| Decision | Prefer | Reason |
|---|---|---|
| Store OIDs in a CI provenance record | Full OID from Git | Avoid ambiguous/truncated machine identifiers |
| Display OIDs in a human dashboard | Git-computed unique abbreviation | Readable while checking uniqueness |
| Inspect object type/content | cat-file/ls-tree |
Works across loose/packed logical storage |
| Create shared repo with SHA-256 | Only after compatibility verification | Repository/server/tooling interoperability matters |
| Recover old work | Preserve reachability/reflogs first | Aggressive pruning reduces options |
13. Knowledge check
Question 1. Why is oid[:7] a poor universal
abbreviation strategy?
Question 2. If an object moves from loose storage into a packfile, does its logical OID/type change?
Question 3. Why can an unreachable object still exist for a while?
Question 4. Does a valid commit OID prove organizational approval?
Question 5. What should a script use to discover the repository hash algorithm?
git rev-parse --show-object-format=storage, rather
than guessing from OID length.
14. Summary
Object format affects compatibility, abbreviation is only a presentation layer, packed storage is an implementation optimization, and reachability determines what history Git considers rooted. Hash-derived OIDs protect object identity/integrity semantics but do not replace signatures or organizational trust controls.
Authoritative references
git-init
git-rev-parse
gitrepository-layout
git-count-objects
git-gc
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.