Chapter 13Lesson 05~150 minutes

Checkpoint Lab — Metrics: Reliability, Maintainability, Security, Complexity, and Size

Produce a two-revision metric interpretation dossier with predictions, task evidence, mode-aware metric keys, raw/rated/new/overall comparisons, limitations, and an engineering recommendation that does not rely on a vanity metric.

CheckpointMetric dossierTwo revisionsLimitationsRecommendation

Learning objectives

  • Execute a complete two-revision metric experiment in a disposable Community Build project and preserve the evidence packet.
  • Predict and independently verify at least two source/indexing/metric/policy state changes before and after the controlled revision.
  • Record mode-aware reliability, maintainability, and security metrics without confusing MQR and Standard Experience terminology.
  • Explain unchanged metrics as carefully as changed metrics and state what each result does not prove.
  • Write an engineering recommendation supported by multiple signals and explicit limitations rather than a single vanity metric.
  • Connect metric interpretation to Chapter 14: Coverage, Test Execution Data, Duplication, and External Analyzer Reports.

1. Checkpoint scenario and success criteria

Use the Chapter 13 fixture—or an equivalent disposable project—to produce a metric interpretation dossier for two controlled revisions. Success is not “all metrics improve.” Success is proving what changed, why it changed, which signals did not change, what the values mean, and what they do not prove.

The dossier must distinguish source state, measurement inputs, server processing, metric definitions, policy context, and human interpretation.

2. Assumptions and preflight

Assumption Mandatory evidence
Community Build / scanner Server version output + sonar-scanner --version
Instance mode MQR or Standard Experience recorded from administration/configuration evidence
Source Full Git history and exact baseline/change SHA
Project identity academy-sq-ch13, stable project key/name
Profile/policy Assigned quality profile, New Code definition, gate assignment
Coverage pytest/pytest-cov versions and preserved XML for both revisions
Plugins/integrations None required; record any installed custom analyzer/plugin that could alter metrics
Credentials Project-analysis token + read-only Browse user token; optional project-admin token for baseline setup; values never stored
git --version
python --version
sonar-scanner --version
curl -fsS "$SONAR_HOST_URL/api/server/version"
git status --short
git rev-parse HEAD
git rev-parse --is-shallow-repository

3. Write predictions before the second revision

At minimum predict two state changes independently. A strong packet predicts several:

State Prediction Independent verification
Revision SHA changes once git rev-parse HEAD
Indexed LOC Increases ncloc + scanner indexing log
Complexity Increases complexity / cognitive_complexity
Overall coverage Decreases coverage XML + Sonar coverage
New coverage Low for changed code Specific-analysis baseline + new_coverage
Reliability/security rating No deliberate degradation Mode-appropriate rating keys
Maintainability Issue/effort may change; grade may remain same Mode-appropriate issue/effort/ratio/rating keys

4. Execute the two-revision workflow

  1. Run the baseline tests/coverage and commit the baseline.
  2. Run baseline Sonar analysis and wait for CE success.
  3. Discover available metric keys and capture baseline-metrics.json.
  4. Pin the baseline analysis as the New Code reference using the Chapter 12 method.
  5. Add the deliberately complex, untested routing function and commit once.
  6. Regenerate coverage evidence without adding tests for the new function.
  7. Run the second Sonar analysis and wait for CE success.
  8. Capture change-metrics.json using the same helper and same metric-key discovery logic.
  9. Do not modify profile, gate, scope, baseline, instance mode, or scanner/server versions between the two analyses.

5. Required evidence packet

evidence/
  assumptions.md
  instance-mode.txt
  server-version.txt
  scanner-version.txt
  python-tooling.txt
  baseline-revision.txt
  change-revision.txt
  baseline-coverage.xml
  change-coverage.xml
  baseline-report-task.txt
  change-report-task.txt
  baseline-ce-task.json
  change-ce-task.json
  metric-catalog.json
  baseline-metrics.json
  change-metrics.json
  metric-delta.txt
  new-code-definition.json
  project-profile-and-gate.md
  scope-and-indexing.md
  interpretation.md
  limitations.md
  recommendation.md
  cleanup.txt

Optional evidence includes screenshots, but the packet must remain interpretable without them. Do not store bearer tokens, passwords, cookies, or unrelated source.

6. Interpretation worksheet

Metric Before After Expected? Causal explanation Does not prove
ncloc … … Yes/No Changed indexed source size Productivity/value
complexity … … Yes/No Added decision branches Runtime speed
cognitive_complexity … … Yes/No Nested/compound flow Production reliability by itself
coverage … … Yes/No Uncovered executable new code Correctness of covered behavior
new_coverage … … Yes/No Specific-analysis New Code population Overall legacy test health
Maintainability rating … … Explain Debt ratio band Absence of difficult design
Reliability/security rating … … Explain Mode-specific issue impact/severity Absence of all defects/vulnerabilities

7. Write a recommendation that uses multiple signals

A strong recommendation resembles:

Revision B added a new routing function that increased LOC and both structural complexity measures while decreasing New Code coverage. Reliability and security ratings did not degrade under the current MQR/Standard profile, which is consistent with the change being primarily structural. Maintainability issue/effort evidence should be reviewed at the rule level, but the letter rating alone is not sufficient to dismiss the complexity increase. Before merging a comparable production change, add focused tests for routing branches and consider refactoring the nested decision structure if the active rule/profile flags it or if change/test difficulty warrants it. No claim about runtime performance is supported by these metrics.

The recommendation uses source, complexity, coverage, ratings, and limitations together. It does not claim “score went down, therefore bad engineer.”

8. Mandatory limitations note

  • Static analysis cannot prove runtime correctness or complete security.
  • Rule/analyzer updates can change issue-derived metrics across time.
  • Coverage depends on external test evidence and does not prove test quality.
  • Complexity is language/analyzer-specific and is not runtime performance.
  • LOC is scope/language-dependent and is not productivity.
  • Debt effort is an estimate derived from rule remediation functions.
  • New Code values depend on the recorded baseline/reference.
  • Cross-project comparison requires a separate comparability contract.

9. Guarded cleanup and credential revocation

  1. Copy evidence out of .scannerwork and verify both CE tasks are preserved.
  2. Revoke the project-analysis token, Browse API token, and temporary project-admin token if created.
  3. Delete the disposable project only after preserving the metric dossier and only if the key is uniquely the lab project.
  4. Delete the local fixture/virtual environment when reproduction is no longer needed.
  5. Do not delete shared SonarQube database/search state, global profiles/gates, unrelated projects, scanner caches, or infrastructure volumes as checkpoint cleanup.

10. What Chapter 13 adds to a governed SonarQube operating model

The organization can now state metrics as reproducible evidence: metric key, definition, population, source revision, mode/profile, analysis task, raw value, derived rating, trend window, and decision context. This prevents dashboards from becoming unexamined scoreboards and makes metric changes diagnosable.

Chapter 14—Coverage, Test Execution Data, Duplication, and External Analyzer Reports—takes four important evidence families that appeared here and teaches their import formats, provenance, path mapping, detection rules, and failure modes in depth.

Knowledge check

What makes the two-revision experiment causal rather than merely correlational?

Why must unchanged metrics be discussed?

Which value in the dossier is safe to use as an individual developer productivity score?

Why retain both New Code and Overall Code measures?

The maintainability rating stays A but cognitive complexity rises. What is the proper response?

What topic comes next?

Next lesson — Next chapter

Coverage, Test Execution Data, Duplication, and External Analyzer Reports

Chapter 14 moves from interpretation to evidence ingestion: how external test/analyzer data and duplication detection become trustworthy SonarQube measures.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-07. Mandatory executable examples target SonarQube Community Build 26.9.0.129388 and SonarScanner CLI 8.1.0.6389. New Community Build instances use MQR Mode by default, but upgraded instances can remain in Standard Experience; therefore examples discover available metric keys and record the actual instance mode rather than hard-coding one issue/rating vocabulary. Core structural keys used in both modes include ncloc, lines, complexity, cognitive_complexity, coverage, and duplicated_lines_density. MQR and Standard issue/rating families are treated as distinct. Re-check primary documentation and the instance’s built-in Web API before automating against another release.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.