Chapter 13Lesson 03~115 minutes

Metrics: Reliability, Maintainability, Security, Complexity, and Size: Configuration, Design Patterns, and Trade-Offs

Choose raw counts, ratios, ratings, new-code views, trends, gate metrics, and diagnostic metrics appropriately while avoiding invalid aggregation and cross-project comparisons.

RatiosTrendsAggregationTrade-offsGovernance

Learning objectives

  • Choose between raw counts, ratios, ratings, and trends according to the engineering question rather than dashboard convenience.
  • Prefer New Code metrics for change control while retaining Overall Code measures for modernization and risk context.
  • Aggregate percentages from their underlying numerators and denominators instead of summing or naively averaging displayed percentages.
  • Separate metrics appropriate for quality-gate enforcement from metrics better used for diagnosis, capacity planning, or architectural discussion.
  • Recognize when cross-project comparisons are invalid because language, scope, profile, mode, baseline, analyzer version, or repository shape differs.
  • Design a metric review record that preserves definition, metric key, population, revision, trend window, and decision rationale.

1. Configuration is really a measurement-design problem

SonarQube exposes many values, but more numbers do not automatically create better governance. A metric review begins by stating the question: “Did this change add maintainability risk?”, “Is the legacy backlog shrinking?”, “Is test evidence sufficient for new code?”, or “Did this release introduce duplication?” The question determines the correct population, metric family, time window, and evidence.

2. Raw counts versus ratios and ratings

Representation Good use Failure mode Evidence to retain
Raw count Queue size, issue population, number of duplicated blocks Bigger project always looks worse Scope/size plus issue/rule context
Density/ratio Normalize a quantity against a denominator Small denominator creates volatile percentages Numerator + denominator, not only percentage
Rating Communicate banded policy/health state Hides movement inside a band Underlying raw/ratio values and thresholds
Trend Direction within the same governed project Scope/profile/baseline changed mid-series Change history and configuration snapshots

Do not choose a ratio merely because it “normalizes” data. The denominator must represent a meaningful exposure. Debt ratio uses development-cost assumptions and LOC; duplicated-lines density uses analyzed lines; coverage uses executable lines/conditions. Each ratio answers a different question.

3. New Code versus Overall Code: two governance horizons

New Code is strongest for release/pull-request discipline because teams can control what they change now. Overall Code preserves legacy and architectural context. A production operating model normally needs both: New Code for enforcement and Overall Code/trends for modernization.

Delivery decision

Prefer current New Code measures tied to a recorded baseline/reference and gate.

Modernization strategy

Use Overall Code plus trends and a legacy backlog; do not reset New Code to make history disappear.

Incident diagnosis

Inspect both populations plus exact revision/task because a new-code failure can be invisible in a stable overall percentage.

4. Trend versus point-in-time

A point-in-time metric describes one analysis. A trend can reveal direction, but only if the measurement system is sufficiently stable. Profile changes, analyzer upgrades, source exclusions, mode switches, New Code baseline changes, coverage-tool changes, and repository restructuring can all create discontinuities.

When one of those events occurs, annotate the metric history. “Complexity dropped on September 7” is incomplete if the same change excluded generated files or split a monorepo into separate projects.

5. Gate metrics versus diagnostic metrics

A quality gate should enforce clear, actionable release policy. Diagnostic metrics can be valuable without being suitable gates. For example, project-level cognitive complexity can help architecture review, but a raw overall complexity threshold often punishes project size and inherited design rather than current change. A New Code issue/coverage condition is usually easier to connect to a remediation decision.

Metric Potential gate use Diagnostic use Caution
New-code issue/quality rating Strong when policy is explicit Triage severity/impact Mode/profile-dependent
New coverage Common delivery control Find untested changed areas Requires valid external coverage evidence; small-change semantics apply
Duplication density on new code Common delivery control Locate copied blocks Language-specific detection thresholds
Overall cognitive complexity Usually weak as a single release gate Architecture/refactoring signal Scales with codebase and language structure
LOC Almost never a quality gate Scope/normalization context Not productivity/value

6. Never sum percentages; aggregate the underlying evidence

Suppose module A reports 90% line coverage over 10 coverable lines and module B reports 50% over 1,000 coverable lines. The arithmetic mean, 70%, is not the combined coverage. Aggregate covered/uncovered lines (and conditions for the full Sonar coverage formula) or use the SonarQube project-level aggregate.

Bad: (90% + 50%) / 2
Better: combine covered/executable counts, then calculate the ratio—or query the project measure SonarQube already computed.

The same principle applies to duplication density and debt ratio. Percentages are outputs of denominators, not additive quantities.

7. Cross-project comparison requires a comparability contract

Comparing two projects by “A rating, 5k LOC, complexity 300” is weak unless their measurement systems align. At minimum record:

  • SonarQube edition/version and instance mode.
  • Languages/analyzer/plugin versions.
  • Source/test scopes and exclusions.
  • Quality profile and rule parameters.
  • New Code definition/reference.
  • Coverage/report tooling and import configuration.
  • Project architecture (monolith, generated code, framework-heavy, library, application).
  • Revision/time window.

Even with those aligned, a within-project trend is often more actionable than league-table ranking across unlike systems.

8. Centrally managed metric policy versus project-specific interpretation

Central standards improve consistency: shared gates, stable metric definitions, documented New Code policy, common dashboards, and version-change controls. But teams still need local interpretation for domain risk. A safety-critical embedded controller and a static marketing site can share metric definitions while making different engineering decisions from the same raw value.

Use exceptions sparingly and record owner, rationale, scope, expiry/review, and evidence. Do not create project-specific thresholds after seeing a failing result.

9. Worked decision scenario

A 12-year-old service has an A maintainability rating, cognitive complexity of 4,200, overall coverage of 44%, new coverage of 92%, and no new reliability/security issues. Which statement is defensible?

Claim Decision Reason
“The service is excellent because maintainability is A.” Reject A is one banded debt-ratio signal, not a correctness/security/architecture proof.
“The current change demonstrates strong test discipline on New Code.” Conditionally accept Verify coverage provenance, New Code baseline, task, and changed lines.
“Complexity 4,200 proves the service is slow.” Reject Complexity is structural, not a runtime performance metric.
“Legacy testability/structure deserve separate modernization work.” Accept Overall coverage/complexity can justify diagnostic review without invalidating the new-code gate.

Knowledge check

Why is a within-project trend often stronger than a cross-project ranking?

Can two A-rated projects have materially different remediation effort?

What is wrong with summing duplicated-lines densities across modules?

Why might an overall complexity threshold be a weak release gate?

What should accompany a metric trend when the analyzer version changes?

Next lesson

Diagnose bad metric conclusions before changing code or policy

Lesson 4 treats metric misuse as an operational failure mode and applies the same first-evidence discipline used throughout the course.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-07. Mandatory executable examples target SonarQube Community Build 26.9.0.129388 and SonarScanner CLI 8.1.0.6389. New Community Build instances use MQR Mode by default, but upgraded instances can remain in Standard Experience; therefore examples discover available metric keys and record the actual instance mode rather than hard-coding one issue/rating vocabulary. Core structural keys used in both modes include ncloc, lines, complexity, cognitive_complexity, coverage, and duplicated_lines_density. MQR and Standard issue/rating families are treated as distinct. Re-check primary documentation and the instance’s built-in Web API before automating against another release.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.