Chapter 01Lesson 01~70 minutes

Build Automation Foundations, Reproducibility, Build Graphs, and the JVM Toolchain Ecosystem: Concepts, Architecture, and Mental Model

Build a precise mental model of JVM build automation: what is an input, what is generated, how Maven lifecycles and Gradle task graphs schedule work, and where versions, repositories, caches, plugins, and CI agents enter the trust boundary.

Build AutomationReproducibilityJVMMavenGradle

Learning objectives

  • Explain a build as a transformation from declared inputs and tool versions to tested, packaged outputs rather than as a single compiler command.
  • Distinguish the JVM, JDK, build-tool runtime, project toolchain, build model, dependency graph, execution graph, repositories, caches, and artifacts.
  • Compare Maven’s convention-driven lifecycle with Gradle’s configurable task graph without pretending their models are identical.
  • Inspect build identity and model state before changing files or clearing caches.
  • Recognize reproducibility and supply-chain boundaries that matter when the same build moves from a laptop to CI.
Version baseline — verified 2026-08-23. The examples use JDK 21, Apache Maven 3.9.16, Apache Maven Wrapper 3.3.4, and Gradle 9.7.1. Maven 3.9.16 is the current recommended Maven 3 release; Maven 4.0.0-rc-6 is still a preview and is intentionally not the production baseline here. Gradle 9 requires JVM 17 or newer to run, so JDK 21 gives both tools a common supported runtime. The versions are teaching pins, not timeless “latest” claims.

1. The real problem: source code is not the deliverable

A Java repository may contain only a few source files, but a production build must answer many more questions. Which JDK compiled the code? Which dependencies and plugins were resolved? Which tests ran? Which generated resources entered the package? Which build-tool version interpreted the configuration? Which repository supplied each byte? Which environment variables, files, clocks, locale settings, or network services influenced the result? A build that cannot answer those questions may still produce a JAR, but it is difficult to reproduce, audit, or trust.

Think of build automation as a controlled transformation. The source tree is only part of the input set. Configuration, tool versions, dependency metadata, plugin code, compiler options, test data, and selected environment values can all become inputs. The output set can include class files, reports, JARs, checksums, metadata, and publication records. A reliable build makes that relationship visible.

Build automation as a transformation with explicit trust boundaries
flowchart TD
  A["Source + tests"] --> M["Build model"]
  C["Config + declared inputs"] --> M
  J["JDK / toolchain"] --> E["Execution engine"]
  M --> E
  R["Dependency & plugin repositories"] --> E
  E --> T["Compile + test + package"]
  T --> O["Artifacts + reports + metadata"]
  E <--> L["Local dependency/cache state"]
  O --> CI["CI / promotion boundary"]

The arrows are important. Repository content flows into the build through dependency and plugin resolution. Local caches can accelerate that flow, but they do not become the source of truth. The JDK and build engine interpret the project model. CI should execute the same declared build rather than secretly replacing it with a second, unrelated build definition.

2. Separate the state stores before you learn commands

Beginners often call every directory involved in a build “the cache.” That collapses several independent mechanisms and makes troubleshooting destructive. Maven’s local repository, Gradle’s dependency cache, Gradle’s task output history, a project target/ or build/ directory, a remote artifact repository, and a CI cache all have different owners and purposes.

State Typical examples What it means Source of truth?
Project definition pom.xml, settings.gradle.kts, build.gradle.kts Declared build model and configuration Yes, when committed
Source/test inputs src/main/java, src/test/java Human-authored program and test inputs Yes
Generated project output target/, build/ Compilation, test and packaging products No; regenerate
Local dependency state ~/.m2/repository, Gradle User Home caches Downloaded artifacts and metadata used to avoid repeated network work No; a performance/local availability layer
Published repository Maven Central or an organization repository manager Shared distribution source for versioned modules/plugins Authoritative only according to repository policy
CI cache/artifact CI cache key, uploaded test report, JAR artifact Pipeline acceleration or retained evidence Depends on the object; cache is not a release source of truth

When a build fails, ask which state store is involved before deleting anything. If a local dependency cache is suspected, reproducing the failure with an isolated fresh cache is stronger evidence than destroying your normal developer state.

3. JVM, JDK, build-tool runtime, and project toolchain are different identities

The JVM executes Java bytecode. A JDK supplies a JVM plus development tools such as javac, jar, and diagnostics. Maven and Gradle are themselves programs that run on a JVM. Separately, a project can select a Java toolchain for compilation or testing. Therefore “we use Java 21” is ambiguous until you state whether Java 21 runs the build tool, compiles the project, runs tests, or is the target runtime.

java -version
javac -version
mvn -v          # bootstrap/global Maven, if installed
gradle -v       # bootstrap/global Gradle, if installed
java -version
javac -version
mvn -v
gradle -v

No mutation occurs here. The output is evidence about the current machine. Later, project wrappers will make the build-tool version project-owned instead of relying on whichever executable happens to be first on PATH.

4. A build has more than one graph

A dependency graph connects the project to external modules needed to compile, test, or run. An execution graph connects units of work: compiler invocations, test execution, packaging, reporting, and publication. A multi-module build adds a project/reactor graph. Those graphs influence each other, but they are not the same object.

For example, adding JUnit changes the test classpath dependency graph. That does not directly mean “run tests.” Maven’s lifecycle bindings or Gradle’s task dependencies are what cause a test execution step to run. Conversely, a packaging task can exist even when no new external dependency is added.

5. Maven’s mental model: project model + lifecycle + plugin goals

Maven centers the build around the Project Object Model (POM). The POM describes coordinates, dependencies, packaging, properties, plugins, and other model state. Maven defines standard build lifecycles. When you invoke a lifecycle phase such as verify, Maven walks the lifecycle up to that phase and executes plugin goals bound to the encountered phases for the project’s packaging.

mvn -v
mvn help:effective-pom
mvn help:effective-settings
mvn dependency:tree

help:effective-pom is especially important because the POM you typed is not always the entire model Maven uses: inheritance, defaults, active profiles, and plugin management can contribute effective configuration. The dependency tree is another model view; it is not a list of files you should edit by hand.

6. Gradle’s mental model: initialization + configuration + task execution

Gradle evaluates settings and build logic to create projects and tasks. Its lifecycle has initialization, configuration, and execution phases. During configuration, Gradle creates/configures the model and determines the task graph required by the requested work. During execution, selected tasks run according to dependencies and ordering constraints.

gradle -v
gradle projects
gradle tasks
gradle help
gradle build --dry-run

--dry-run helps you inspect selected task ordering without executing task actions. It does not prove that dependency resolution, tests, or packaging would succeed. It is one view of the execution plan.

7. Maven and Gradle are not interchangeable command syntaxes

Question Maven Gradle
Primary project model XML POM with conventions, inheritance and profiles Build/settings scripts and typed model APIs
Execution structure Lifecycle phases with plugin goals bound to phases Task graph selected from requested tasks and dependencies
Typical build invocation ./mvnw verify ./gradlew build
Repeated unchanged build Lifecycle goals generally run again; dependency downloads may already be warm Tasks with correctly declared inputs/outputs can become UP-TO-DATE or cache hits
Generated directory convention target/ build/
Strength Strong conventions and predictable lifecycle vocabulary Flexible model/task composition and rich incremental execution
Risk when misunderstood Assuming phase names are independent commands or ignoring effective model/plugin bindings Hiding inputs in arbitrary build logic or doing work during configuration

A useful comparison maps concepts, not spellings. mvn verify and gradle build are both common verification-oriented entry points, but they do not imply identical internal work or identical artifacts.

8. Reproducibility means controlling inputs, not merely running clean

A clean build removes generated project output, but it does not automatically reset remote repository state, downloaded dependency metadata, the JDK, the build-tool version, plugin versions, locale, time zone, environment variables, or external services. Reproducibility asks whether the same declared source and controlled inputs produce the same intended result in a suitably isolated environment.

For byte-level comparison, a checksum such as SHA-256 gives an exact identity for one artifact. Matching checksums across two controlled runs is useful evidence, but it is only as strong as the scope of the experiment. Two builds in the same warm workspace do not prove that a fresh CI agent will resolve exactly the same graph.

sha256sum target/*.jar
sha256sum build/libs/*.jar
Get-FileHash .\target\*.jar -Algorithm SHA256
Get-FileHash .\build\libs\*.jar -Algorithm SHA256

9. Dependency and plugin resolution is code execution trust

A dependency can become code compiled into or executed by your application. A build plugin can execute during the build itself. Annotation processors, Gradle plugins, Maven plugins/extensions, init scripts, and repository metadata can therefore influence the build before a final application ever starts.

That is why repository URLs, wrapper distribution URLs, plugin versions, dependency versions, verification metadata, signing keys, and CI credentials are security-sensitive build state. “It is only a build file” is the wrong trust model: build definitions are executable automation.

Safety boundary: do not paste repository passwords, signing keys, private tokens, or CI secrets into a POM, Gradle script, command history, screenshot, or lesson lab. Use synthetic values and local file:// repositories until a later chapter explicitly teaches publishing credentials.

10. DevOps connection: CI should reproduce the project build, not invent a second one

A pipeline agent is another machine with another filesystem, JDK, network path, cache state, and credential context. A project wrapper and explicit toolchain make more of that environment project-controlled. CI then becomes an orchestrator that checks out a revision, supplies approved external inputs, invokes the project build, retains evidence, and promotes an identified artifact.

A healthy boundary looks like this: the repository owns the build model and wrapper version; CI owns ephemeral execution infrastructure and secret injection; an artifact repository owns published immutable versions; policy determines which inputs and outputs are allowed to cross those boundaries.

11. Mini-lab — classify build observations before changing anything

In any disposable Java project you already have, record ten observations: JDK runtime, compiler version, Maven or Gradle version, wrapper presence, project model file, dependency repository declaration, generated-output directory, one dependency coordinate, one test report location, and one package output. Label each observation as project source, project configuration, toolchain, external repository input, local cache, generated output, or CI/external state.

Verification: you should be able to say which observations are safe to delete and regenerate, which are committed source-of-truth inputs, and which live outside the repository. Cleanup: none; this lab is read-only.

Knowledge check

A Maven build works after deleting target/. Does that prove it is reproducible on a fresh CI agent?

A Gradle build reports :compileJava UP-TO-DATE. What has been proven?

Why is a plugin repository part of the security boundary?

Is mvn verify just Maven’s spelling of gradle build?

A project wrapper pins the build tool. What important identity can still differ between machines?

Summary

  • A build is a transformation from explicit source, configuration, tool, dependency, plugin, and environment inputs to tested and packaged outputs.
  • Dependency graphs, project/reactor graphs, and execution graphs are related but distinct.
  • Maven’s POM/lifecycle/plugin-goal model and Gradle’s configured task graph are different abstractions, not command aliases.
  • Wrappers control build-tool distribution; they do not automatically control the JDK, repositories, or every environmental input.
  • Generated outputs and local caches are not the same as committed source-of-truth configuration.
  • Reproducibility and supply-chain security require exact input identity, trustworthy resolution, and evidence such as test results and artifact checksums.
Next lesson

Build the same tiny JVM behavior through two controlled build paths

Lesson 2 creates disposable Maven and Gradle projects, bootstraps wrappers, isolates local dependency/cache state, runs tests and packaging, and inspects the resulting execution evidence.

Primary sources and version notes

Version-sensitive statements in this lesson were checked on 2026-08-23. Re-check these primary sources when reusing the examples after a Maven, Gradle, JDK, or plugin upgrade.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.