Chapter 14Lesson 02~145 minutes

Sparse Checkout, Partial Clone, Shallow Clone, LFS, and Large Repository Access: Guided Hands-On Workflow and Core Operations

Measure full, sparse, shallow, and blobless partial clones in a disposable synthetic repository; deepen/unshallow history, exercise sparse-index operations, and follow an optional/no-account Git LFS path.

Hands-on scale labSparse indexDepthPromisor objects

Learning objectives

  • Create a reproducible synthetic repository with moderate history and a generated large file.
  • Use sparse-checkout init/set/add/reapply/disable and verify physical/index effects.
  • Create, deepen, and unshallow a bounded-history clone.
  • Demonstrate a local server-compatible blobless partial clone and demand fetching.
  • Exercise Git LFS locally when installed or analyze pointer mechanics without requiring storage.

1. Build one synthetic repository for every comparison

The lab uses only disposable local repositories. It creates several service directories, documentation, a large generated test file, and enough history to make shallow and partial behavior visible.

Git Bash, Bash, or zsh

mkdir git-large-access-lab
cd git-large-access-lab
git init -b trunk source
cd source
git config user.name "Large Repo Lab"
git config user.email "large-repo-lab@example.invalid"

mkdir -p services/api services/web libs/common docs assets history
printf "api=v1\n" > services/api/app.txt
printf "web=v1\n" > services/web/app.txt
printf "common=v1\n" > libs/common/lib.txt
printf "# Operations Guide\n" > docs/guide.md
printf "history-01\n" > history/churn.txt

python - <<'PY'
from pathlib import Path
Path("assets/large-test.bin").write_bytes((b"0123456789abcdef" * 131072))
PY

git add .
git commit -m "Create synthetic monorepo baseline"

for n in 2 3 4 5 6 7 8 9 10 11 12; do
  printf "api=v%s\n" "$n" >> services/api/app.txt
  printf "history-%02d\n" "$n" > history/churn.txt
  git add services/api/app.txt history/churn.txt
  git commit -m "Synthetic history step $n"
done

cd ..
git clone --bare source origin.git
git --git-dir=origin.git config uploadpack.allowFilter true

PowerShell setup alternative

New-Item -ItemType Directory git-large-access-lab | Out-Null
Set-Location git-large-access-lab
git init -b trunk source
Set-Location source
git config user.name "Large Repo Lab"
git config user.email "large-repo-lab@example.invalid"

'services/api','services/web','libs/common','docs','assets','history' | ForEach-Object {
  New-Item -ItemType Directory -Force $_ | Out-Null
}
Set-Content services/api/app.txt 'api=v1'
Set-Content services/web/app.txt 'web=v1'
Set-Content libs/common/lib.txt 'common=v1'
Set-Content docs/guide.md '# Operations Guide'
Set-Content history/churn.txt 'history-01'
[IO.File]::WriteAllBytes('assets/large-test.bin', [Text.Encoding]::ASCII.GetBytes(('0123456789abcdef' * 131072)))

git add .
git commit -m "Create synthetic monorepo baseline"

2..12 | ForEach-Object {
  Add-Content services/api/app.txt "api=v$_"
  Set-Content history/churn.txt ("history-{0:D2}" -f $_)
  git add services/api/app.txt history/churn.txt
  git commit -m "Synthetic history step $_"
}

Set-Location ..
git clone --bare source origin.git
git --git-dir=origin.git config uploadpack.allowFilter true

uploadpack.allowFilter=true is set only on the disposable local bare server so the partial-clone exercise can demonstrate protocol filtering. Production servers/hosting products decide independently whether filtering is supported.

2. Establish the source baseline

git -C source rev-list --count HEAD
git -C source count-objects -vH
git --git-dir=origin.git config --get uploadpack.allowFilter

Expected: twelve commits on trunk. Object-size values vary by Git version/compression and should be compared relatively rather than hard-coded.

3. Full clone — the comparison baseline

git clone --no-local origin.git full
git -C full rev-parse --is-shallow-repository
git -C full rev-list --count HEAD
git -C full count-objects -vH
git -C full status --short --branch

--no-local deliberately disables Git's local hardlink/copy shortcut so this local lab behaves more like a transport clone. The result has complete history and ordinary file blobs.

4. Sparse checkout — reduce populated files while keeping the full object store

git clone --no-local origin.git sparse
git -C sparse sparse-checkout init --cone --sparse-index
git -C sparse sparse-checkout set services/api docs

git -C sparse sparse-checkout list
git -C sparse config --get core.sparseCheckout
git -C sparse config --get core.sparseCheckoutCone
git -C sparse config --get index.sparse
git -C sparse ls-files --sparse
git -C sparse count-objects -vH

In cone mode, root-level files remain populated as parent-pattern content, while directories such as services/web and assets are normally absent from the working directory. count-objects remains comparable to the full clone because sparse checkout did not ask the server to omit objects.

5. Measure populated files using the shell, not git ls-files

git ls-files describes index paths and therefore still knows about excluded content. Count physical files separately:

Git Bash / Bash / zsh

find sparse -path 'sparse/.git' -prune -o -type f -print | wc -l
test -f sparse/services/api/app.txt
test ! -f sparse/services/web/app.txt
test ! -f sparse/assets/large-test.bin

PowerShell

(Get-ChildItem sparse -File -Recurse -Force | Where-Object FullName -NotMatch '[\\/]\.git[\\/]').Count
Test-Path sparse/services/api/app.txt
Test-Path sparse/services/web/app.txt
Test-Path sparse/assets/large-test.bin

6. Grow the sparse cone, reapply it, then disable it

git -C sparse sparse-checkout add libs/common
git -C sparse sparse-checkout list
git -C sparse sparse-checkout reapply
git -C sparse status --short --branch

git -C sparse sparse-checkout disable
git -C sparse config --get core.sparseCheckout || true

add includes another directory without replacing existing selections. reapply reapplies sparsity after operations that may have materialized paths or after cleanup of paths Git previously could not sparsify. disable repopulates the full working tree.

7. Shallow clone — truncate commit history, not file contents

git clone --no-local --depth=3 origin.git shallow
git -C shallow rev-parse --is-shallow-repository
git -C shallow rev-list --count HEAD
git -C shallow log --oneline --decorate

Expected: true and three reachable commits on the cloned branch. The working tree still contains the current versions of all selected normal paths; shallow clone is about ancestry depth.

8. Deepen by three commits, then restore full history

git -C shallow fetch --deepen=3
git -C shallow rev-list --count HEAD

git -C shallow fetch --unshallow
git -C shallow rev-parse --is-shallow-repository
git -C shallow rev-list --count HEAD

After deepening, the count should increase relative to three. After unshallowing from this complete local source, the count returns to twelve and --is-shallow-repository reports false.

9. Partial clone — omit blobs before checkout

git clone --no-local --filter=blob:none --no-checkout origin.git partial

git -C partial config --get remote.origin.promisor
git -C partial config --get remote.origin.partialclonefilter
git -C partial count-objects -vH
git -C partial rev-list --objects --all --missing=print

The promisor setting should be true and the filter should be blob:none. Lines beginning with ? in the last command identify objects absent from the local object store. The initial object pack should be materially smaller than the full-clone pack in this synthetic repository because the large test blob and other file blobs were omitted.

10. Checkout and historical inspection can demand-fetch missing blobs

git -C partial switch trunk
git -C partial rev-list --objects --all --missing=print

git -C partial show HEAD~8:history/churn.txt
git -C partial rev-list --objects --all --missing=print

Checkout requires current working-tree blobs, so Git fetches some of them. Historical show can then fetch an older version that remained absent. This is the operational tradeoff of partial clone: less up-front transfer in exchange for possible later network I/O.

11. Partial clone requires server/protocol support

If a server does not advertise/support filtering, a client cannot rely on --filter to save transfer. This lab explicitly enabled filtering on the local upload-pack server. Hosting providers may support different filters or policies; verify them rather than assuming equivalence.

12. Check whether Git LFS is installed before using LFS commands

git lfs version

If the command is unavailable, skip to the pointer-analysis section. The mandatory lesson does not require installing an extension or using paid storage.

13. Optional local LFS exercise when the extension is installed

git -C full lfs install --local
git -C full lfs track "*.bin"
git -C full add .gitattributes assets/large-test.bin
git -C full commit -m "Track generated binary through Git LFS"

git -C full lfs ls-files
git -C full show HEAD:.gitattributes
git -C full show HEAD:assets/large-test.bin

The working-tree assets/large-test.bin contains binary payload bytes, but the committed Git blob should display an LFS pointer. git lfs track updates .gitattributes; that file must be committed so collaborators receive the filter policy.

14. No-LFS path — analyze a valid pointer shape conceptually

version https://git-lfs.github.com/spec/v1
oid sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
size 2097152

A pointer is UTF-8 text with a specification version, content OID, and size. The sample is not evidence that an LFS server possesses the referenced payload. This distinction is central to diagnosing “pointer file present but object unavailable.”

15. Challenge — choose the mechanism from the measured bottleneck

  1. A developer needs only services/api in a monorepo but requires complete history offline. Which mechanism addresses working-tree scale without truncating history?
  2. A CI lint job needs only the latest three commits but every current file. Sparse or shallow?
  3. A clone has complete commit history but historical blobs are fetched only when inspected. Which mechanism is active?
  4. A 4-GB video should not live as ordinary Git blobs across every revision. Which extension addresses that payload model?
  5. Sparse checkout reduced populated files but clone transfer stayed large. What mechanism could address object transfer?

16. Cleanup

Confirm the disposable lab path before recursive deletion.

Git Bash / Bash / zsh

cd ..
pwd
rm -rf git-large-access-lab

PowerShell

Set-Location ..
Get-Location
Remove-Item -Recurse -Force git-large-access-lab

17. Knowledge check

Question 1. Why use --no-local in the local clone comparisons?

Question 2. Why can git ls-files still show paths excluded from a sparse working tree?

Question 3. What configuration proves a clone is using origin as a promisor remote?

Question 4. Why does shallow clone affect changelog/version logic?

Question 5. Why must .gitattributes be committed when using Git LFS?

18. Summary

You measured four access models against one source repository: full clone as the baseline, sparse checkout/index for working-tree/index scale, shallow clone for bounded ancestry, and partial clone for deferred object transfer. Git LFS remains an optional extension whose pointer mechanics can be understood without any hosted storage.

Next

Turn mechanisms into repository and CI policy

Lesson 3 examines their configuration surfaces, compatibility requirements, commit-graph interactions, LFS attributes, and Scalar as an optional large-repository orchestration tool.

Authoritative references

 git-sparse-checkout
 git-clone
 fetch options
 partial-clone
 Git LFS

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.