Chapter 03Lesson 03~80 minutes

Git Object Model: Blobs, Trees, Commits, Tags, and the Object Database: Configuration, Design Choices, and Tradeoffs

Make object architecture and storage tradeoffs explicit: repository hash format, unique abbreviations, loose versus packed storage, reachability and garbage collection, and integrity versus trust.

SHA-1 / SHA-256AbbreviationReachabilityStorage

Learning objectives

  • Discover and reason about a repository's object format instead of assuming 40-character IDs.
  • Use Git-computed unique abbreviations rather than fixed string truncation.
  • Separate the logical object model from loose/packed physical storage.
  • Explain reachability and its consequences for retention and recovery.
  • Distinguish hash-based object integrity from authentication, authorization, and approval.

1. Object format is repository architecture, not a display preference

The repository's object format determines how Git computes and represents object IDs for stored objects. Current Git supports SHA-1 and SHA-256 repository object formats, but current documentation states that there is still no direct interoperability between SHA-1 and SHA-256 repositories. Choosing an object format is therefore an architectural compatibility decision made when the repository is created—not something to flip casually in an existing project.

2. Discover the format from Git

git rev-parse --show-object-format=storage
git rev-parse --show-object-format=input
git rev-parse --show-object-format=output

Most scripts only need the storage format plus opaque OIDs returned by Git. Do not infer the algorithm from hard-coded length checks when Git can report the format directly.

3. Optional disposable SHA-256 experiment

Compatibility experiment only. Do not convert a production repository just to follow this section. Use a separate disposable repository, and verify your installed Git supports the requested format.

Git Bash, Bash, or zsh

mkdir sha256-object-lab
cd sha256-object-lab
git init --object-format=sha256 -b trunk
git config user.name "SHA256 Lab"
git config user.email "sha256@example.invalid"
printf "hello\n" > hello.txt
git add hello.txt
git commit -m "Create SHA-256 snapshot"
git rev-parse --show-object-format=storage
git rev-parse HEAD
cd ..
rm -rf sha256-object-lab

PowerShell file-creation alternative

New-Item -ItemType Directory sha256-object-lab | Out-Null
Set-Location sha256-object-lab
git init --object-format=sha256 -b trunk
git config user.name "SHA256 Lab"
git config user.email "sha256@example.invalid"
Set-Content hello.txt 'hello'
git add hello.txt
git commit -m "Create SHA-256 snapshot"
git rev-parse --show-object-format=storage
git rev-parse HEAD
Set-Location ..
Remove-Item -Recurse -Force sha256-object-lab

The purpose is not “SHA-256 is always better, switch now.” It is to make format-dependence visible and reinforce that scripts should not be wired to SHA-1 widths.

4. Abbreviated object names are conveniences, not new identifiers

git rev-parse HEAD
git rev-parse --short HEAD
git rev-parse --short=12 HEAD

--short asks Git to find a unique prefix with at least the requested length. The default minimum is influenced by core.abbrev. A short prefix is useful for humans, logs, and UI, but long-lived machine contracts should store full OIDs unless the consuming protocol explicitly defines abbreviation semantics.

5. Collision-safe display means letting Git check uniqueness

A prefix such as abc1234 is only useful if it resolves uniquely in the relevant repository. As repositories grow, a formerly convenient prefix length may be too short. Avoid scripts that truncate an OID with string slicing:

BAD IDEA (pseudo-code):
short = oid[0:7]

BETTER:
git rev-parse --short=12 <object-name>

The second form asks Git to preserve uniqueness rather than assuming seven or twelve characters are always enough.

6. Loose objects versus packed storage

New objects are often first written as individual (“loose”) object files under the object store. Git can later pack many objects together and delta-compress representations for storage/performance. The logical object model does not change. A commit still refers to its tree by object ID whether the object is loose, in a packfile, or represented internally as a delta.

git count-objects -vH
git cat-file -t HEAD
git cat-file -p HEAD

Use Git commands to inspect logical objects. Do not build application logic that assumes every OID maps to a loose file at .git/objects/xx/yyyy…; objects may be packed, borrowed from alternates, or otherwise stored through supported repository mechanisms.

7. Reachability is the retention concept that matters

An object is reachable when Git can walk to it from a root such as a reference (and, in recovery contexts, other retention roots such as reflogs). Reachability is why a branch ref can keep an entire ancestry graph alive: the branch points to a commit, that commit points to its tree and parents, those trees point to blobs/trees, and so on.

When a ref moves away from an old commit, the old object is not necessarily deleted immediately. Reflogs and garbage-collection grace periods can preserve recovery opportunities. Later chapters cover recovery in depth.

8. Garbage collection is storage maintenance, not “delete everything not on screen”

Current git gc performs housekeeping such as repacking and handling unreachable objects. Its pruning behavior uses grace periods; immediate pruning reduces recovery options and can be unsafe around concurrent writers.

Do not use git gc --prune=now in normal learning repositories or as a troubleshooting reflex. It shortens the safety window that later recovery techniques may rely on.

9. Hash integrity versus identity and trust

If Git expects object OID X and the stored object's bytes no longer hash to X under the repository format, that is an integrity problem. But a valid object can still contain a malicious change, a false author field, or content introduced by someone who lacked organizational approval.

Question OID/hash helps? What else may be needed?
Did these stored bytes change unexpectedly? Yes, integrity checking is relevant Repository health checks / trusted copies
Who authored this commit metadata? No proof by itself Identity policy, signatures, platform records
Was this change approved for production? No Review/authorization/CI policy
Is this release label immutable by policy? No Protected refs and release governance

10. Configuration scope: what belongs globally and what belongs to the repository

Object format is repository architecture and is recorded by repository state. New-repository defaults such as init.defaultObjectFormat can be configured, but a personal global preference can create compatibility surprises if teammates expect the conventional format. Treat such defaults like infrastructure architecture: make them explicit in bootstrap automation and team documentation.

Human abbreviation preferences such as core.abbrev can be user-level, but automation should still request an explicit format/length appropriate to its contract.

11. Git core versus hosting/server capabilities

A hosting platform may or may not support every repository object format or migration path that your local Git binary supports. Local capability is not proof of server compatibility. Before choosing a non-default object format for a shared repository, verify the actual remote server, CI runners, libraries, IDEs, mirroring tools, archival systems, and deployment integrations.

12. Decision table

Decision Prefer Reason
Store OIDs in a CI provenance record Full OID from Git Avoid ambiguous/truncated machine identifiers
Display OIDs in a human dashboard Git-computed unique abbreviation Readable while checking uniqueness
Inspect object type/content cat-file/ls-tree Works across loose/packed logical storage
Create shared repo with SHA-256 Only after compatibility verification Repository/server/tooling interoperability matters
Recover old work Preserve reachability/reflogs first Aggressive pruning reduces options

13. Knowledge check

Question 1. Why is oid[:7] a poor universal abbreviation strategy?

Question 2. If an object moves from loose storage into a packfile, does its logical OID/type change?

Question 3. Why can an unreachable object still exist for a while?

Question 4. Does a valid commit OID prove organizational approval?

Question 5. What should a script use to discover the repository hash algorithm?

14. Summary

Object format affects compatibility, abbreviation is only a presentation layer, packed storage is an implementation optimization, and reachability determines what history Git considers rooted. Hash-derived OIDs protect object identity/integrity semantics but do not replace signatures or organizational trust controls.

Next

Diagnose object-model misconceptions before they become destructive fixes

Lesson 4 deliberately breaks assumptions about filenames, branches, shortened hashes, trust, and manual object-store editing, then uses evidence-first commands to diagnose each case.

Authoritative references

 git-init
 git-rev-parse
 gitrepository-layout
 git-count-objects
 git-gc

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.