Chapter 05Lesson 01~70 minutes

Commits, Diffs, History Inspection, Revision Ranges, and Message Discipline: Concepts, Architecture, and Mental Model

Understand commits as snapshot-and-parent objects, diffs as computed state comparisons, revision selectors and ranges, author versus committer metadata, path-limited history, and message discipline as operational data.

Commit anatomyDiff semanticsRevision rangesMessages

Learning objectives

  • Explain which fields make up a commit object and why message/metadata changes create new commit identities.
  • Distinguish stored snapshots from computed patch/diff views.
  • Use parent selectors and revision ranges to describe graph locations and commit sets precisely.
  • Distinguish author identity/date from committer identity/date.
  • Explain how message quality affects review, release automation, incident forensics, and later bisect workflows.

1. The practical problem — a history can exist and still be hard to use

Chapter 03 showed that a commit is an immutable object pointing to a tree and parent commits. Chapter 04 showed that a new commit records the index, not every file currently visible in the working tree. This chapter asks what makes that recorded history useful after the moment of creation.

A repository with hundreds of technically valid commits can still be operationally poor if the commits mix unrelated changes, carry messages such as “fix stuff,” or are inspected with range expressions the team does not understand. Reviewers, release tooling, incident responders, and future maintainers need to answer exact questions: what snapshot did this commit record, what came before it, what changed between two states, which commits are unique to one line of development, and why was the change made?

2. Commit identity comes from the entire commit object

A normal commit object records a root tree OID, zero or more parent commit OIDs, author identity/date, committer identity/date, and the commit message. The commit OID identifies that exact object content. Change the tree, parent set, identity metadata, timestamps, or message and the resulting commit object is different.

git cat-file -p HEAD
git show --no-patch --format=raw HEAD

The first command exposes the object's stored fields directly. The second gives a commit-oriented view. Neither should be confused with the patch that git show normally prints after the message.

3. A commit stores a snapshot relationship, not a patch file

Git commits point to trees. A textual patch is computed when two states are compared. For an ordinary one-parent commit, git show COMMIT usually presents the commit metadata/message followed by a diff between its parent tree and the commit tree. That patch is a useful view of the change, but it is not a patch blob embedded inside the commit object.

Snapshot objects and computed comparisons
flowchart TD
P[Parent commit P] --> PT[Tree P]
C[Commit C] --> CT[Tree C]
C -->|parent| P
PT -. git diff P C computes .-> D[Patch / comparison output]
CT -. git diff P C computes .-> D
C -. git show formats metadata + comparison .-> V[Human review view]

The solid arrows are stored object relationships. The dotted arrows are views computed at command time. This distinction matters later when diff algorithms or display options change the patch presentation without changing either commit.

4. Author and committer are different roles, with different dates

The author identifies the person/time originally associated with writing the change. The committer identifies the person/time that particular commit object was created in this history. In a simple local commit they are often the same. They can differ after applying somebody else's patch, rebasing, cherry-picking, importing history, or deliberately setting metadata.

git show --no-patch --format=fuller HEAD
git log -1 --format='author=%an <%ae> %aI%ncommitter=%cn <%ce> %cI'

Do not sort operational events by author date and assume it is the time a commit entered a branch. The parent graph defines ancestry; author and committer timestamps are metadata that answer different questions.

5. Single revisions and parent selectors

A revision such as HEAD, a branch name, tag, or full OID identifies an object that Git can resolve. Parent syntax lets you navigate the graph without hard-coding OIDs:

Expression Meaning
HEAD^ or HEAD^1 First parent of HEAD
M^2 Second parent of merge commit M
HEAD~2 Follow first-parent links twice
HEAD^{tree} Peel the commit to its tree object

These are object-selection expressions. They are different from revision set ranges used by history traversal.

6. Revision ranges describe sets of commits

Commands such as git log and git rev-list walk commit reachability. For those commands:

  • A..B means commits reachable from B, excluding commits reachable from A.
  • A...B means the symmetric difference: commits reachable from either side but not from both.
git rev-list --oneline A..B
git rev-list --left-right --oneline A...B

The three-dot form is especially useful before integration because it exposes commits unique to both sides. It is not simply “a bigger diff.”

7. Important exception: dotted notation in git diff is endpoint comparison syntax

git diff is not a history-set listing command. git diff A B compares the two endpoint trees; git diff A..B is effectively the same endpoint comparison. git diff A...B means “compare the merge base of A and B to B.”

git diff A B
git diff A..B
git diff A...B
Do not transfer the git log A...B set definition mechanically to git diff A...B. Always ask whether the command is traversing a commit set or comparing snapshots.

8. Path-limited history answers a narrower question

git log --oneline -- src/service.py
git log -p -- src/service.py
git log --full-history --oneline -- src/service.py

The -- makes the path boundary explicit. Path filtering can simplify history, especially around merges, so an apparently “missing” commit may be a history-simplification effect rather than lost data. When a forensic question matters, compare an unfiltered graph with the filtered view and learn the simplification options rather than trusting one short command blindly.

9. Commit messages are durable operational metadata

Git requires no universal sentence style beyond a commit message unless you explicitly allow an empty one, but teams often adopt conventions because humans and automation consume history. A useful message usually contains a concise subject that states the change and, when the reason is not obvious, a body that explains why, constraints, or consequences.

Weak message More useful message
update config Raise worker timeout for slow archive uploads
fix tests Stabilize retry test by controlling the clock
files changed Reject unsigned deployment manifests before upload

Issue IDs, type prefixes, signoffs, trailers, and line-length conventions are project policy, not intrinsic Git truth. Teach the reason for the convention and avoid cargo-cult formatting.

10. Why readable history is operational data

Release automation may generate notes from subjects or tags. Incident responders may use log, show, or later bisect to isolate a regression. Review systems display commit messages and diffs. Build provenance records commit OIDs. A clean history reduces the number of ambiguous decisions humans and tools must make under pressure.

11. Read-only inspection ladder

git status --short --branch
git rev-parse --verify HEAD
git show --no-patch --format=fuller HEAD
git log --graph --decorate --oneline --all
git rev-list --count HEAD
git diff
git diff --staged
git diff HEAD^ HEAD

These commands establish working state, the current commit, metadata, graph shape, commit-set traversal, and the three common diff endpoints before any new commit is created.

12. Knowledge check

Question 1. Is the patch printed by git show stored inside the commit object?

Question 2. What is the difference between author date and committer date?

Question 3. For git log, what does A..B select?

Question 4. Why is git diff A...B not the same concept as git log A...B?

Question 5. Why should a message explain why when the reason is not obvious?

13. Summary

Commits are snapshot-plus-parent objects with identity metadata; diffs are computed comparisons; parent selectors navigate one graph location; ranges select commit sets for traversal; path filters narrow/simplify history; and messages turn technically valid history into operationally useful history.

Next

Build a divergent history and query it precisely

Lesson 2 creates a disposable branch graph, compares working tree/index/HEAD and arbitrary commits, exercises two-dot/three-dot semantics with rev-list, and inspects a merge commit through both parent edges.

Authoritative references

 git-commit
 gitrevisions
 git-diff
 git-log
 pretty-formats

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.