Metrics: Reliability, Maintainability, Security, Complexity, and Size: Configuration, Design Patterns, and Trade-Offs
Choose raw counts, ratios, ratings, new-code views, trends, gate metrics, and diagnostic metrics appropriately while avoiding invalid aggregation and cross-project comparisons.
Learning objectives
- Choose between raw counts, ratios, ratings, and trends according to the engineering question rather than dashboard convenience.
- Prefer New Code metrics for change control while retaining Overall Code measures for modernization and risk context.
- Aggregate percentages from their underlying numerators and denominators instead of summing or naively averaging displayed percentages.
- Separate metrics appropriate for quality-gate enforcement from metrics better used for diagnosis, capacity planning, or architectural discussion.
- Recognize when cross-project comparisons are invalid because language, scope, profile, mode, baseline, analyzer version, or repository shape differs.
- Design a metric review record that preserves definition, metric key, population, revision, trend window, and decision rationale.
1. Configuration is really a measurement-design problem
SonarQube exposes many values, but more numbers do not automatically create better governance. A metric review begins by stating the question: “Did this change add maintainability risk?”, “Is the legacy backlog shrinking?”, “Is test evidence sufficient for new code?”, or “Did this release introduce duplication?” The question determines the correct population, metric family, time window, and evidence.
2. Raw counts versus ratios and ratings
| Representation | Good use | Failure mode | Evidence to retain |
|---|---|---|---|
| Raw count | Queue size, issue population, number of duplicated blocks | Bigger project always looks worse | Scope/size plus issue/rule context |
| Density/ratio | Normalize a quantity against a denominator | Small denominator creates volatile percentages | Numerator + denominator, not only percentage |
| Rating | Communicate banded policy/health state | Hides movement inside a band | Underlying raw/ratio values and thresholds |
| Trend | Direction within the same governed project | Scope/profile/baseline changed mid-series | Change history and configuration snapshots |
Do not choose a ratio merely because it “normalizes” data. The denominator must represent a meaningful exposure. Debt ratio uses development-cost assumptions and LOC; duplicated-lines density uses analyzed lines; coverage uses executable lines/conditions. Each ratio answers a different question.
3. New Code versus Overall Code: two governance horizons
New Code is strongest for release/pull-request discipline because teams can control what they change now. Overall Code preserves legacy and architectural context. A production operating model normally needs both: New Code for enforcement and Overall Code/trends for modernization.
Delivery decision
Prefer current New Code measures tied to a recorded baseline/reference and gate.
Modernization strategy
Use Overall Code plus trends and a legacy backlog; do not reset New Code to make history disappear.
Incident diagnosis
Inspect both populations plus exact revision/task because a new-code failure can be invisible in a stable overall percentage.
4. Trend versus point-in-time
A point-in-time metric describes one analysis. A trend can reveal direction, but only if the measurement system is sufficiently stable. Profile changes, analyzer upgrades, source exclusions, mode switches, New Code baseline changes, coverage-tool changes, and repository restructuring can all create discontinuities.
When one of those events occurs, annotate the metric history. “Complexity dropped on September 7” is incomplete if the same change excluded generated files or split a monorepo into separate projects.
5. Gate metrics versus diagnostic metrics
A quality gate should enforce clear, actionable release policy. Diagnostic metrics can be valuable without being suitable gates. For example, project-level cognitive complexity can help architecture review, but a raw overall complexity threshold often punishes project size and inherited design rather than current change. A New Code issue/coverage condition is usually easier to connect to a remediation decision.
| Metric | Potential gate use | Diagnostic use | Caution |
|---|---|---|---|
| New-code issue/quality rating | Strong when policy is explicit | Triage severity/impact | Mode/profile-dependent |
| New coverage | Common delivery control | Find untested changed areas | Requires valid external coverage evidence; small-change semantics apply |
| Duplication density on new code | Common delivery control | Locate copied blocks | Language-specific detection thresholds |
| Overall cognitive complexity | Usually weak as a single release gate | Architecture/refactoring signal | Scales with codebase and language structure |
| LOC | Almost never a quality gate | Scope/normalization context | Not productivity/value |
6. Never sum percentages; aggregate the underlying evidence
Suppose module A reports 90% line coverage over 10 coverable lines and module B reports 50% over 1,000 coverable lines. The arithmetic mean, 70%, is not the combined coverage. Aggregate covered/uncovered lines (and conditions for the full Sonar coverage formula) or use the SonarQube project-level aggregate.
(90% + 50%) / 2Better: combine covered/executable counts, then calculate the ratio—or query the project measure SonarQube already computed.
The same principle applies to duplication density and debt ratio. Percentages are outputs of denominators, not additive quantities.
7. Cross-project comparison requires a comparability contract
Comparing two projects by “A rating, 5k LOC, complexity 300” is weak unless their measurement systems align. At minimum record:
- SonarQube edition/version and instance mode.
- Languages/analyzer/plugin versions.
- Source/test scopes and exclusions.
- Quality profile and rule parameters.
- New Code definition/reference.
- Coverage/report tooling and import configuration.
- Project architecture (monolith, generated code, framework-heavy, library, application).
- Revision/time window.
Even with those aligned, a within-project trend is often more actionable than league-table ranking across unlike systems.
8. Centrally managed metric policy versus project-specific interpretation
Central standards improve consistency: shared gates, stable metric definitions, documented New Code policy, common dashboards, and version-change controls. But teams still need local interpretation for domain risk. A safety-critical embedded controller and a static marketing site can share metric definitions while making different engineering decisions from the same raw value.
Use exceptions sparingly and record owner, rationale, scope, expiry/review, and evidence. Do not create project-specific thresholds after seeing a failing result.
9. Worked decision scenario
A 12-year-old service has an A maintainability rating, cognitive complexity of 4,200, overall coverage of 44%, new coverage of 92%, and no new reliability/security issues. Which statement is defensible?
| Claim | Decision | Reason |
|---|---|---|
| “The service is excellent because maintainability is A.” | Reject | A is one banded debt-ratio signal, not a correctness/security/architecture proof. |
| “The current change demonstrates strong test discipline on New Code.” | Conditionally accept | Verify coverage provenance, New Code baseline, task, and changed lines. |
| “Complexity 4,200 proves the service is slow.” | Reject | Complexity is structural, not a runtime performance metric. |
| “Legacy testability/structure deserve separate modernization work.” | Accept | Overall coverage/complexity can justify diagnostic review without invalidating the new-code gate. |
Knowledge check
Why is a within-project trend often stronger than a cross-project ranking?
More of the measurement system—language, scope, profile, architecture, baseline—can remain stable, reducing confounding differences.
Can two A-rated projects have materially different remediation effort?
Yes. The rating is a banded ratio; raw effort, LOC, and position inside the band can differ substantially.
What is wrong with summing duplicated-lines densities across modules?
Densities are ratios with denominators. Combine underlying duplicated/total line evidence or use the aggregate measure.
Why might an overall complexity threshold be a weak release gate?
It scales with project size and legacy structure and may not isolate the current change or provide an actionable remediation boundary.
What should accompany a metric trend when the analyzer version changes?
A change annotation/version record and cautious interpretation, because analyzer changes can alter findings or measurements independently of source behavior.
Official references and version notes
- Understanding measures and metrics — current metric definitions and metric keys for software qualities, maintainability, coverage, duplication, size, complexity, and issues.
- Code metrics introduction — how metrics participate in rules, quality gates, monitoring, and mode-dependent UI behavior.
- Changing instance modes — MQR versus Standard Experience classifications, severities, and metric-family implications.
- Instance mode overview — current MQR/Standard model and the default for new Community Build instances.
-
Web API
— bearer authentication,
/api/measures,/api/metrics, and Web API V2 migration guidance. - Understanding quality gates — how selected measures become enforceable policy rather than descriptive dashboards.
-
New Code
— population semantics used by
new_*measures. - Analysis overview — scanner/report/Compute Engine lifecycle behind persisted measures.
- Test coverage overview — current external-coverage model; Chapter 14 expands this subject.
- SonarQube downloads — current Community Build release identity.
- SonarScanner CLI 8.1.0.6389 — scanner baseline used by the local lab.
Rechecked 2026-09-07. Mandatory executable examples target
SonarQube Community Build 26.9.0.129388 and
SonarScanner CLI 8.1.0.6389. New Community Build
instances use MQR Mode by default, but upgraded instances can
remain in Standard Experience; therefore examples discover
available metric keys and record the actual instance mode rather
than hard-coding one issue/rating vocabulary. Core structural keys
used in both modes include ncloc, lines,
complexity, cognitive_complexity,
coverage, and duplicated_lines_density.
MQR and Standard issue/rating families are treated as distinct.
Re-check primary documentation and the instance’s built-in Web API
before automating against another release.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.