Checkpoint Lab — Metrics: Reliability, Maintainability, Security, Complexity, and Size
Produce a two-revision metric interpretation dossier with predictions, task evidence, mode-aware metric keys, raw/rated/new/overall comparisons, limitations, and an engineering recommendation that does not rely on a vanity metric.
Learning objectives
- Execute a complete two-revision metric experiment in a disposable Community Build project and preserve the evidence packet.
- Predict and independently verify at least two source/indexing/metric/policy state changes before and after the controlled revision.
- Record mode-aware reliability, maintainability, and security metrics without confusing MQR and Standard Experience terminology.
- Explain unchanged metrics as carefully as changed metrics and state what each result does not prove.
- Write an engineering recommendation supported by multiple signals and explicit limitations rather than a single vanity metric.
- Connect metric interpretation to Chapter 14: Coverage, Test Execution Data, Duplication, and External Analyzer Reports.
1. Checkpoint scenario and success criteria
Use the Chapter 13 fixture—or an equivalent disposable project—to produce a metric interpretation dossier for two controlled revisions. Success is not “all metrics improve.” Success is proving what changed, why it changed, which signals did not change, what the values mean, and what they do not prove.
The dossier must distinguish source state, measurement inputs, server processing, metric definitions, policy context, and human interpretation.
2. Assumptions and preflight
| Assumption | Mandatory evidence |
|---|---|
| Community Build / scanner |
Server version output + sonar-scanner --version
|
| Instance mode | MQR or Standard Experience recorded from administration/configuration evidence |
| Source | Full Git history and exact baseline/change SHA |
| Project identity | academy-sq-ch13, stable project key/name |
| Profile/policy | Assigned quality profile, New Code definition, gate assignment |
| Coverage | pytest/pytest-cov versions and preserved XML for both revisions |
| Plugins/integrations | None required; record any installed custom analyzer/plugin that could alter metrics |
| Credentials | Project-analysis token + read-only Browse user token; optional project-admin token for baseline setup; values never stored |
git --version
python --version
sonar-scanner --version
curl -fsS "$SONAR_HOST_URL/api/server/version"
git status --short
git rev-parse HEAD
git rev-parse --is-shallow-repository
3. Write predictions before the second revision
At minimum predict two state changes independently. A strong packet predicts several:
| State | Prediction | Independent verification |
|---|---|---|
| Revision | SHA changes once | git rev-parse HEAD |
| Indexed LOC | Increases | ncloc + scanner indexing log |
| Complexity | Increases |
complexity / cognitive_complexity
|
| Overall coverage | Decreases | coverage XML + Sonar coverage |
| New coverage | Low for changed code |
Specific-analysis baseline + new_coverage
|
| Reliability/security rating | No deliberate degradation | Mode-appropriate rating keys |
| Maintainability | Issue/effort may change; grade may remain same | Mode-appropriate issue/effort/ratio/rating keys |
4. Execute the two-revision workflow
- Run the baseline tests/coverage and commit the baseline.
- Run baseline Sonar analysis and wait for CE success.
-
Discover available metric keys and capture
baseline-metrics.json. - Pin the baseline analysis as the New Code reference using the Chapter 12 method.
- Add the deliberately complex, untested routing function and commit once.
- Regenerate coverage evidence without adding tests for the new function.
- Run the second Sonar analysis and wait for CE success.
-
Capture
change-metrics.jsonusing the same helper and same metric-key discovery logic. - Do not modify profile, gate, scope, baseline, instance mode, or scanner/server versions between the two analyses.
5. Required evidence packet
evidence/
assumptions.md
instance-mode.txt
server-version.txt
scanner-version.txt
python-tooling.txt
baseline-revision.txt
change-revision.txt
baseline-coverage.xml
change-coverage.xml
baseline-report-task.txt
change-report-task.txt
baseline-ce-task.json
change-ce-task.json
metric-catalog.json
baseline-metrics.json
change-metrics.json
metric-delta.txt
new-code-definition.json
project-profile-and-gate.md
scope-and-indexing.md
interpretation.md
limitations.md
recommendation.md
cleanup.txt
Optional evidence includes screenshots, but the packet must remain interpretable without them. Do not store bearer tokens, passwords, cookies, or unrelated source.
6. Interpretation worksheet
| Metric | Before | After | Expected? | Causal explanation | Does not prove |
|---|---|---|---|---|---|
ncloc |
… | … | Yes/No | Changed indexed source size | Productivity/value |
complexity |
… | … | Yes/No | Added decision branches | Runtime speed |
cognitive_complexity |
… | … | Yes/No | Nested/compound flow | Production reliability by itself |
coverage |
… | … | Yes/No | Uncovered executable new code | Correctness of covered behavior |
new_coverage |
… | … | Yes/No | Specific-analysis New Code population | Overall legacy test health |
| Maintainability rating | … | … | Explain | Debt ratio band | Absence of difficult design |
| Reliability/security rating | … | … | Explain | Mode-specific issue impact/severity | Absence of all defects/vulnerabilities |
7. Write a recommendation that uses multiple signals
A strong recommendation resembles:
Revision B added a new routing function that increased LOC and both structural complexity measures while decreasing New Code coverage. Reliability and security ratings did not degrade under the current MQR/Standard profile, which is consistent with the change being primarily structural. Maintainability issue/effort evidence should be reviewed at the rule level, but the letter rating alone is not sufficient to dismiss the complexity increase. Before merging a comparable production change, add focused tests for routing branches and consider refactoring the nested decision structure if the active rule/profile flags it or if change/test difficulty warrants it. No claim about runtime performance is supported by these metrics.
The recommendation uses source, complexity, coverage, ratings, and limitations together. It does not claim “score went down, therefore bad engineer.”
8. Mandatory limitations note
- Static analysis cannot prove runtime correctness or complete security.
- Rule/analyzer updates can change issue-derived metrics across time.
- Coverage depends on external test evidence and does not prove test quality.
- Complexity is language/analyzer-specific and is not runtime performance.
- LOC is scope/language-dependent and is not productivity.
- Debt effort is an estimate derived from rule remediation functions.
- New Code values depend on the recorded baseline/reference.
- Cross-project comparison requires a separate comparability contract.
9. Guarded cleanup and credential revocation
-
Copy evidence out of
.scannerworkand verify both CE tasks are preserved. - Revoke the project-analysis token, Browse API token, and temporary project-admin token if created.
- Delete the disposable project only after preserving the metric dossier and only if the key is uniquely the lab project.
- Delete the local fixture/virtual environment when reproduction is no longer needed.
- Do not delete shared SonarQube database/search state, global profiles/gates, unrelated projects, scanner caches, or infrastructure volumes as checkpoint cleanup.
10. What Chapter 13 adds to a governed SonarQube operating model
The organization can now state metrics as reproducible evidence: metric key, definition, population, source revision, mode/profile, analysis task, raw value, derived rating, trend window, and decision context. This prevents dashboards from becoming unexamined scoreboards and makes metric changes diagnosable.
Chapter 14—Coverage, Test Execution Data, Duplication, and External Analyzer Reports—takes four important evidence families that appeared here and teaches their import formats, provenance, path mapping, detection rules, and failure modes in depth.
Knowledge check
What makes the two-revision experiment causal rather than merely correlational?
The source change is controlled while scanner/server/profile/mode/scope/baseline remain fixed, and each analysis is tied to its own revision and CE task.
Why must unchanged metrics be discussed?
They constrain the causal story. For example, unchanged security/reliability ratings support the conclusion that the intended change was structural rather than a detected security/reliability regression.
Which value in the dossier is safe to use as an individual developer productivity score?
None. These are code-analysis measures, not individual performance metrics.
Why retain both New Code and Overall Code measures?
New Code governs current changes; Overall Code preserves legacy/system context and modernization trends.
The maintainability rating stays A but cognitive complexity rises. What is the proper response?
Inspect the raw complexity and any relevant maintainability rule/effort evidence. The A rating is a debt-ratio band, not a veto on diagnostic structural concerns.
What topic comes next?
Coverage, test execution data, duplication detection, and external analyzer reports—the provenance-heavy inputs behind several metrics used in this dossier.
Official references and version notes
- Understanding measures and metrics — current metric definitions and metric keys for software qualities, maintainability, coverage, duplication, size, complexity, and issues.
- Code metrics introduction — how metrics participate in rules, quality gates, monitoring, and mode-dependent UI behavior.
- Changing instance modes — MQR versus Standard Experience classifications, severities, and metric-family implications.
- Instance mode overview — current MQR/Standard model and the default for new Community Build instances.
-
Web API
— bearer authentication,
/api/measures,/api/metrics, and Web API V2 migration guidance. - Understanding quality gates — how selected measures become enforceable policy rather than descriptive dashboards.
-
New Code
— population semantics used by
new_*measures. - Analysis overview — scanner/report/Compute Engine lifecycle behind persisted measures.
- Test coverage overview — current external-coverage model; Chapter 14 expands this subject.
- SonarQube downloads — current Community Build release identity.
- SonarScanner CLI 8.1.0.6389 — scanner baseline used by the local lab.
Rechecked 2026-09-07. Mandatory executable examples target
SonarQube Community Build 26.9.0.129388 and
SonarScanner CLI 8.1.0.6389. New Community Build
instances use MQR Mode by default, but upgraded instances can
remain in Standard Experience; therefore examples discover
available metric keys and record the actual instance mode rather
than hard-coding one issue/rating vocabulary. Core structural keys
used in both modes include ncloc, lines,
complexity, cognitive_complexity,
coverage, and duplicated_lines_density.
MQR and Standard issue/rating families are treated as distinct.
Re-check primary documentation and the instance’s built-in Web API
before automating against another release.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.