Sparse Checkout, Partial Clone, Shallow Clone, LFS, and Large Repository Access: Guided Hands-On Workflow and Core Operations
Measure full, sparse, shallow, and blobless partial clones in a disposable synthetic repository; deepen/unshallow history, exercise sparse-index operations, and follow an optional/no-account Git LFS path.
Learning objectives
- Create a reproducible synthetic repository with moderate history and a generated large file.
- Use sparse-checkout init/set/add/reapply/disable and verify physical/index effects.
- Create, deepen, and unshallow a bounded-history clone.
- Demonstrate a local server-compatible blobless partial clone and demand fetching.
- Exercise Git LFS locally when installed or analyze pointer mechanics without requiring storage.
1. Build one synthetic repository for every comparison
The lab uses only disposable local repositories. It creates several service directories, documentation, a large generated test file, and enough history to make shallow and partial behavior visible.
Git Bash, Bash, or zsh
mkdir git-large-access-lab
cd git-large-access-lab
git init -b trunk source
cd source
git config user.name "Large Repo Lab"
git config user.email "large-repo-lab@example.invalid"
mkdir -p services/api services/web libs/common docs assets history
printf "api=v1\n" > services/api/app.txt
printf "web=v1\n" > services/web/app.txt
printf "common=v1\n" > libs/common/lib.txt
printf "# Operations Guide\n" > docs/guide.md
printf "history-01\n" > history/churn.txt
python - <<'PY'
from pathlib import Path
Path("assets/large-test.bin").write_bytes((b"0123456789abcdef" * 131072))
PY
git add .
git commit -m "Create synthetic monorepo baseline"
for n in 2 3 4 5 6 7 8 9 10 11 12; do
printf "api=v%s\n" "$n" >> services/api/app.txt
printf "history-%02d\n" "$n" > history/churn.txt
git add services/api/app.txt history/churn.txt
git commit -m "Synthetic history step $n"
done
cd ..
git clone --bare source origin.git
git --git-dir=origin.git config uploadpack.allowFilter true
PowerShell setup alternative
New-Item -ItemType Directory git-large-access-lab | Out-Null
Set-Location git-large-access-lab
git init -b trunk source
Set-Location source
git config user.name "Large Repo Lab"
git config user.email "large-repo-lab@example.invalid"
'services/api','services/web','libs/common','docs','assets','history' | ForEach-Object {
New-Item -ItemType Directory -Force $_ | Out-Null
}
Set-Content services/api/app.txt 'api=v1'
Set-Content services/web/app.txt 'web=v1'
Set-Content libs/common/lib.txt 'common=v1'
Set-Content docs/guide.md '# Operations Guide'
Set-Content history/churn.txt 'history-01'
[IO.File]::WriteAllBytes('assets/large-test.bin', [Text.Encoding]::ASCII.GetBytes(('0123456789abcdef' * 131072)))
git add .
git commit -m "Create synthetic monorepo baseline"
2..12 | ForEach-Object {
Add-Content services/api/app.txt "api=v$_"
Set-Content history/churn.txt ("history-{0:D2}" -f $_)
git add services/api/app.txt history/churn.txt
git commit -m "Synthetic history step $_"
}
Set-Location ..
git clone --bare source origin.git
git --git-dir=origin.git config uploadpack.allowFilter true
uploadpack.allowFilter=true is set only on the
disposable local bare server so the partial-clone exercise can
demonstrate protocol filtering. Production servers/hosting products
decide independently whether filtering is supported.
2. Establish the source baseline
git -C source rev-list --count HEAD
git -C source count-objects -vH
git --git-dir=origin.git config --get uploadpack.allowFilter
Expected: twelve commits on trunk.
Object-size values vary by Git version/compression and should be
compared relatively rather than hard-coded.
3. Full clone — the comparison baseline
git clone --no-local origin.git full
git -C full rev-parse --is-shallow-repository
git -C full rev-list --count HEAD
git -C full count-objects -vH
git -C full status --short --branch
--no-local deliberately disables Git's local
hardlink/copy shortcut so this local lab behaves more like a
transport clone. The result has complete history and ordinary file
blobs.
4. Sparse checkout — reduce populated files while keeping the full object store
git clone --no-local origin.git sparse
git -C sparse sparse-checkout init --cone --sparse-index
git -C sparse sparse-checkout set services/api docs
git -C sparse sparse-checkout list
git -C sparse config --get core.sparseCheckout
git -C sparse config --get core.sparseCheckoutCone
git -C sparse config --get index.sparse
git -C sparse ls-files --sparse
git -C sparse count-objects -vH
In cone mode, root-level files remain populated as parent-pattern
content, while directories such as services/web and
assets are normally absent from the working directory.
count-objects remains comparable to the full clone
because sparse checkout did not ask the server to omit objects.
5. Measure populated files using the shell, not
git ls-files
git ls-files describes index paths and therefore still
knows about excluded content. Count physical files separately:
Git Bash / Bash / zsh
find sparse -path 'sparse/.git' -prune -o -type f -print | wc -l
test -f sparse/services/api/app.txt
test ! -f sparse/services/web/app.txt
test ! -f sparse/assets/large-test.bin
PowerShell
(Get-ChildItem sparse -File -Recurse -Force | Where-Object FullName -NotMatch '[\\/]\.git[\\/]').Count
Test-Path sparse/services/api/app.txt
Test-Path sparse/services/web/app.txt
Test-Path sparse/assets/large-test.bin
6. Grow the sparse cone, reapply it, then disable it
git -C sparse sparse-checkout add libs/common
git -C sparse sparse-checkout list
git -C sparse sparse-checkout reapply
git -C sparse status --short --branch
git -C sparse sparse-checkout disable
git -C sparse config --get core.sparseCheckout || true
add includes another directory without replacing
existing selections. reapply reapplies sparsity after
operations that may have materialized paths or after cleanup of
paths Git previously could not sparsify.
disable repopulates the full working tree.
7. Shallow clone — truncate commit history, not file contents
git clone --no-local --depth=3 origin.git shallow
git -C shallow rev-parse --is-shallow-repository
git -C shallow rev-list --count HEAD
git -C shallow log --oneline --decorate
Expected: true and three reachable
commits on the cloned branch. The working tree still contains the
current versions of all selected normal paths; shallow clone is
about ancestry depth.
8. Deepen by three commits, then restore full history
git -C shallow fetch --deepen=3
git -C shallow rev-list --count HEAD
git -C shallow fetch --unshallow
git -C shallow rev-parse --is-shallow-repository
git -C shallow rev-list --count HEAD
After deepening, the count should increase relative to three. After
unshallowing from this complete local source, the count returns to
twelve and --is-shallow-repository reports false.
9. Partial clone — omit blobs before checkout
git clone --no-local --filter=blob:none --no-checkout origin.git partial
git -C partial config --get remote.origin.promisor
git -C partial config --get remote.origin.partialclonefilter
git -C partial count-objects -vH
git -C partial rev-list --objects --all --missing=print
The promisor setting should be true and the filter should be
blob:none. Lines beginning with ? in the
last command identify objects absent from the local object store.
The initial object pack should be materially smaller than the
full-clone pack in this synthetic repository because the large test
blob and other file blobs were omitted.
10. Checkout and historical inspection can demand-fetch missing blobs
git -C partial switch trunk
git -C partial rev-list --objects --all --missing=print
git -C partial show HEAD~8:history/churn.txt
git -C partial rev-list --objects --all --missing=print
Checkout requires current working-tree blobs, so Git fetches some of
them. Historical show can then fetch an older version
that remained absent. This is the operational tradeoff of partial
clone: less up-front transfer in exchange for possible later network
I/O.
11. Partial clone requires server/protocol support
If a server does not advertise/support filtering, a client cannot
rely on --filter to save transfer. This lab explicitly
enabled filtering on the local upload-pack server.
Hosting providers may support different filters or policies; verify
them rather than assuming equivalence.
12. Check whether Git LFS is installed before using LFS commands
git lfs version
If the command is unavailable, skip to the pointer-analysis section. The mandatory lesson does not require installing an extension or using paid storage.
13. Optional local LFS exercise when the extension is installed
git -C full lfs install --local
git -C full lfs track "*.bin"
git -C full add .gitattributes assets/large-test.bin
git -C full commit -m "Track generated binary through Git LFS"
git -C full lfs ls-files
git -C full show HEAD:.gitattributes
git -C full show HEAD:assets/large-test.bin
The working-tree assets/large-test.bin contains binary
payload bytes, but the committed Git blob should display an LFS
pointer. git lfs track updates
.gitattributes; that file must be committed so
collaborators receive the filter policy.
14. No-LFS path — analyze a valid pointer shape conceptually
version https://git-lfs.github.com/spec/v1
oid sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
size 2097152
A pointer is UTF-8 text with a specification version, content OID, and size. The sample is not evidence that an LFS server possesses the referenced payload. This distinction is central to diagnosing “pointer file present but object unavailable.”
15. Challenge — choose the mechanism from the measured bottleneck
-
A developer needs only
services/apiin a monorepo but requires complete history offline. Which mechanism addresses working-tree scale without truncating history? - A CI lint job needs only the latest three commits but every current file. Sparse or shallow?
- A clone has complete commit history but historical blobs are fetched only when inspected. Which mechanism is active?
- A 4-GB video should not live as ordinary Git blobs across every revision. Which extension addresses that payload model?
- Sparse checkout reduced populated files but clone transfer stayed large. What mechanism could address object transfer?
16. Cleanup
Git Bash / Bash / zsh
cd ..
pwd
rm -rf git-large-access-lab
PowerShell
Set-Location ..
Get-Location
Remove-Item -Recurse -Force git-large-access-lab
17. Knowledge check
Question 1. Why use --no-local in the local clone
comparisons?
Question 2. Why can git ls-files still show paths
excluded from a sparse working tree?
Question 3. What configuration proves a clone is using origin as a promisor remote?
remote.origin.promisor=true, with the filter recorded
under remote.origin.partialclonefilter.
Question 4. Why does shallow clone affect changelog/version logic?
Question 5. Why must .gitattributes be committed
when using Git LFS?
18. Summary
You measured four access models against one source repository: full clone as the baseline, sparse checkout/index for working-tree/index scale, shallow clone for bounded ancestry, and partial clone for deferred object transfer. Git LFS remains an optional extension whose pointer mechanics can be understood without any hosted storage.
Authoritative references
git-sparse-checkout
git-clone
fetch options
partial-clone
Git LFS
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.