Checkpoint Lab — Sparse Checkout, Partial Clone, Shallow Clone, LFS, and Large Repository Access
Checkpoint large-repository access with full, depth-limited, sparse/sparse-index, and blobless partial clones, optional LFS mechanics, measurable verification, and developer-versus-CI recommendations.
Learning objectives
- Predict and verify the history/object/working-tree differences between four clone/access modes.
- Demonstrate depth, deepen, and unshallow behavior against one source tip.
- Correct a sparse build dependency by extending the cone rather than abandoning sparsity.
- Verify blobless partial-clone promisor state and controlled demand fetching.
- Write workload-specific recommendations that preserve correctness while reducing measured cost.
1. Checkpoint scenario — four clones, one source, one large generated file
You are designing checkout policy for a synthetic monorepo.
Developers mainly work in services/api and
docs; CI has a fast current-tip build and a separate
release-analysis job; one generated binary represents the kind of
file that may justify LFS. You will measure full, shallow, sparse,
and blobless partial clones and then write a policy recommendation.
2. Predictions before creating optimized clones
- Will sparse checkout materially reduce the full clone's Git object pack size by itself?
- Will a depth-3 clone report exactly the same reachable commit count as the source?
- Can a blobless partial clone have complete commit history while reporting missing objects?
-
If sparse index is active, can excluded directories appear as
directory entries in
git ls-files --sparse?
3. Create the disposable synthetic source and remote
Git Bash, Bash, or zsh
mkdir git-large-checkpoint
cd git-large-checkpoint
git init -b trunk source
cd source
git config user.name "Large Access Checkpoint"
git config user.email "large-access@example.invalid"
mkdir -p services/api services/web services/worker libs/common tools/codegen docs assets history
printf "api=v1\n" > services/api/app.txt
printf "web=v1\n" > services/web/app.txt
printf "worker=v1\n" > services/worker/app.txt
printf "common=v1\n" > libs/common/lib.txt
printf "codegen=v1\n" > tools/codegen/tool.txt
printf "# Build and Operations\n" > docs/guide.md
printf "history-01\n" > history/churn.txt
python - <<'PY'
from pathlib import Path
Path("assets/generated-large.bin").write_bytes(b"A" * (2 * 1024 * 1024))
PY
git add .
git commit -m "Checkpoint baseline"
for n in 2 3 4 5 6 7 8 9 10 11 12; do
printf "api=v%s\n" "$n" >> services/api/app.txt
printf "history-%02d\n" "$n" > history/churn.txt
git add services/api/app.txt history/churn.txt
git commit -m "Checkpoint history step $n"
done
cd ..
git clone --bare source origin.git
git --git-dir=origin.git config uploadpack.allowFilter true
PowerShell setup alternative
New-Item -ItemType Directory git-large-checkpoint | Out-Null
Set-Location git-large-checkpoint
git init -b trunk source
Set-Location source
git config user.name "Large Access Checkpoint"
git config user.email "large-access@example.invalid"
'services/api','services/web','services/worker','libs/common','tools/codegen','docs','assets','history' | ForEach-Object {
New-Item -ItemType Directory -Force $_ | Out-Null
}
Set-Content services/api/app.txt 'api=v1'
Set-Content services/web/app.txt 'web=v1'
Set-Content services/worker/app.txt 'worker=v1'
Set-Content libs/common/lib.txt 'common=v1'
Set-Content tools/codegen/tool.txt 'codegen=v1'
Set-Content docs/guide.md '# Build and Operations'
Set-Content history/churn.txt 'history-01'
[IO.File]::WriteAllBytes('assets/generated-large.bin', [Text.Encoding]::ASCII.GetBytes(('A' * (2MB))))
git add .
git commit -m "Checkpoint baseline"
2..12 | ForEach-Object {
Add-Content services/api/app.txt "api=v$_"
Set-Content history/churn.txt ("history-{0:D2}" -f $_)
git add services/api/app.txt history/churn.txt
git commit -m "Checkpoint history step $_"
}
Set-Location ..
git clone --bare source origin.git
git --git-dir=origin.git config uploadpack.allowFilter true
The second PowerShell byte-generation line is intentionally explicit but can be memory-heavy on older shells; using any editor/script to generate a disposable 2-MB file is acceptable. The exact byte pattern is irrelevant—the file only creates measurable blob pressure.
4. Record source truth before optimizing
SOURCE_COMMITS=$(git -C source rev-list --count HEAD)
SOURCE_TIP=$(git -C source rev-parse HEAD)
printf "source_commits=%s\nsource_tip=%s\n" "$SOURCE_COMMITS" "$SOURCE_TIP"
git -C source count-objects -vH
git --git-dir=origin.git config --get uploadpack.allowFilter
Expected commit count is twelve. Save the exact tip OID; every clone
should initially identify the same trunk tip even when
its local history/object completeness differs.
5. Full clone baseline
git clone --no-local origin.git full
FULL_COMMITS=$(git -C full rev-list --count HEAD)
git -C full rev-parse HEAD
git -C full rev-parse --is-shallow-repository
git -C full count-objects -vH
Verify that the tip equals SOURCE_TIP, commit count
equals SOURCE_COMMITS, and the repository is not
shallow.
6. Shallow clone and boundary verification
git clone --no-local --depth=3 origin.git shallow
SHALLOW_COMMITS=$(git -C shallow rev-list --count HEAD)
git -C shallow rev-parse HEAD
git -C shallow rev-parse --is-shallow-repository
git -C shallow log --oneline --decorate
Verify prediction 2: the tip should match the source but reachable history should be three commits, not twelve.
git -C shallow fetch --deepen=3
git -C shallow rev-list --count HEAD
git -C shallow fetch --unshallow
git -C shallow rev-list --count HEAD
git -C shallow rev-parse --is-shallow-repository
7. Sparse checkout plus sparse index
git clone --no-local origin.git sparse
git -C sparse sparse-checkout init --cone --sparse-index
git -C sparse sparse-checkout set services/api docs
git -C sparse sparse-checkout list
git -C sparse config --get index.sparse
git -C sparse ls-files --sparse
git -C sparse count-objects -vH
Verify prediction 1: object storage remains broadly
comparable to the full clone because this was still a full object
clone. Verify prediction 4:
ls-files --sparse can show excluded directories as
sparse directory entries.
8. Discover and add a build dependency to the cone
The API build requires tools/codegen. Prove the file
exists in history but is absent from the sparse working tree:
git -C sparse ls-tree -r --name-only HEAD -- tools/codegen
test ! -f sparse/tools/codegen/tool.txt
git -C sparse sparse-checkout add tools/codegen
test -f sparse/tools/codegen/tool.txt
git -C sparse sparse-checkout list
git -C sparse sparse-checkout reapply
This turns a build failure into a documented checkout-profile correction rather than disabling sparsity without diagnosis.
9. Blobless partial clone before checkout
git clone --no-local --filter=blob:none --no-checkout origin.git partial
git -C partial rev-parse HEAD
git -C partial rev-list --count HEAD
git -C partial config --get remote.origin.promisor
git -C partial config --get remote.origin.partialclonefilter
git -C partial count-objects -vH
git -C partial rev-list --objects --all --missing=print
Verify prediction 3: the partial clone can have all twelve commits while still showing missing blobs. Its initial pack should be much smaller than the full clone in this synthetic repository.
10. Trigger controlled demand fetching
git -C partial switch trunk
git -C partial rev-list --objects --all --missing=print
git -C partial show HEAD~8:history/churn.txt
git -C partial rev-list --objects --all --missing=print
The first switch fetches current checkout blobs. The historical show can fetch an older churn-file blob. Compare the missing-object list before and after rather than assuming every checkout fully hydrates history.
11. Optional LFS checkpoint or pointer analysis
git lfs version
If Git LFS is installed, use the full clone:
git -C full lfs install --local
git -C full lfs track "*.bin"
git -C full add .gitattributes assets/generated-large.bin
git -C full commit -m "Represent generated binary with Git LFS"
git -C full lfs ls-files
git -C full show HEAD:assets/generated-large.bin
If LFS is unavailable, inspect this conceptual pointer instead:
version https://git-lfs.github.com/spec/v1
oid sha256:bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb
size 2097152
Record in your recommendation that LFS requires extension support plus durable LFS object storage; a pointer by itself is not the payload.
12. Build the comparison table from measured evidence
printf "full_commits=%s\n" "$(git -C full rev-list --count HEAD)"
printf "shallow_after_unshallow=%s\n" "$(git -C shallow rev-list --count HEAD)"
printf "sparse_commits=%s\n" "$(git -C sparse rev-list --count HEAD)"
printf "partial_commits=%s\n" "$(git -C partial rev-list --count HEAD)"
git -C full count-objects -vH
git -C sparse count-objects -vH
git -C partial count-objects -vH
Do not compare exact pack sizes across machines as a grading criterion; compression and Git implementation details vary. Compare mechanism-level direction and repository semantics.
13. Write developer-versus-CI recommendations
| Persona/job | Recommended mode | Required guardrail |
|---|---|---|
| API developer with offline history needs | Full history + cone sparse checkout + sparse index | Document required shared/tool directories |
| Fast current-tip compile CI | Potential shallow clone | Test that build/version logic does not require old tags/merge bases |
| Release/changelog CI | Full history/tags | Do not inherit depth defaults from generic checkout templates |
| Disk-constrained online investigator | Partial clone, optionally sparse | Promisor remote availability and prefetch plan |
| Large-binary project | LFS/artifact storage | Object-retention/backup and extension availability |
14. Verification checklist
- Source has twelve commits and a recorded exact tip OID.
- Full clone has complete history and ordinary object storage.
- Depth-3 clone initially has three reachable commits and reports shallow=true.
- Deepen increases available ancestry; unshallow restores full history.
- Sparse checkout populates API/docs but initially excludes unrelated directories.
- Sparse index is enabled and excluded directories can be represented sparsely.
-
Adding
tools/codegento the sparse cone materializes its required file. -
Blobless partial clone records origin as a promisor with
blob:none. - Partial clone initially reports missing objects while retaining complete commit count.
- Historical content access can reduce missing-object count by demand fetching.
- LFS is either exercised with a local extension or explicitly analyzed as pointer/external-storage mechanics.
- Recommendation distinguishes developer, current-tip CI, release CI, and large-binary workloads.
15. Cleanup
Git Bash / Bash / zsh
cd ..
pwd
rm -rf git-large-checkpoint
PowerShell
Set-Location ..
Get-Location
Remove-Item -Recurse -Force git-large-checkpoint
16. Knowledge check
Question 1. Why did the sparse clone keep roughly full object storage?
Question 2. Why did the partial clone keep all twelve commits while missing blobs?
Question 3. A release job fails to find an old tag in a depth-limited checkout. Which correction comes first?
Question 4. Why add tools/codegen to the sparse
cone instead of disabling sparse checkout?
Question 5. What is the final policy principle of this checkpoint?
17. What Chapter 14 adds to a production Git operating model
You can now classify large-repository cost by layer, select sparse working/index access, partial object transfer, shallow ancestry, or LFS payload externalization deliberately, and write CI/developer policies that state what information is intentionally absent.
18. Chapter checkpoint summary
“Smaller clone” is not the goal. The goal is the smallest and fastest repository view that still answers every correctness question required by that workflow. Measure first, reduce one layer intentionally, and preserve a documented path back to fuller data.
Authoritative references
git-sparse-checkout
sparse-index
partial-clone
git-clone
fetch options
Git LFS specification
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.