Chapter 13Lesson 04~125 minutes

Metrics: Reliability, Maintainability, Security, Complexity, and Size: Diagnostics, Failure Modes, and Production Practices

Diagnose metric misuse such as LOC-as-productivity, complexity-as-performance, percentage summation, A-rating overclaims, scope drift, and incomparable project baselines without changing policy to manufacture a desired story.

DiagnosticsMetric misuseScope driftCausalityProduction

Learning objectives

  • Detect and reject LOC or issue-count productivity rankings and explain the incentive and measurement failures they create.
  • Explain why cyclomatic/cognitive complexity does not measure runtime speed, latency, throughput, or resource consumption.
  • Diagnose incorrect percentage aggregation and repair it with underlying counts or SonarQube project-level measures.
  • Treat an A rating as bounded evidence about one metric family rather than proof of correctness, absence of exploitable defects, or delivery readiness.
  • Preserve source-scope, profile, New Code definition, analyzer version, mode, and task evidence when a metric changes unexpectedly.
  • Follow an evidence-first diagnostic sequence that separates scanner/indexing, analyzer/rule, coverage/report, policy, server/Compute Engine, and metric-interpretation causes.

1. Metric incidents are often interpretation incidents

When a dashboard “looks bad,” teams often change code, exclusions, rules, or thresholds before proving what changed. That destroys the evidence needed to distinguish source behavior from measurement-system drift. Chapter 13 applies the same incident discipline as scanner/server troubleshooting: preserve first evidence, classify the layer, make the least destructive correction, and rerun the smallest equivalent scenario.

2. Evidence-first diagnostic sequence

  1. Preserve the metric snapshot, dashboard/export, API response, scanner log, report-task.txt, CE task, revision, and timestamp.
  2. Confirm SonarQube edition/version and instance mode.
  3. Confirm scanner/analyzer/plugin/runtime versions and source revision.
  4. Confirm base directory, indexed files, source/test scope, inclusions/exclusions, generated/vendor treatment, and SCM state.
  5. Confirm profile/rule parameters, issue workflow state, New Code definition/reference, and gate assignment.
  6. Confirm external report provenance when coverage/test/external analyzer measures are involved.
  7. Inspect the raw metric definition/key and underlying numerator/denominator or issue population.
  8. Inspect CE success and analysis history to ensure the measure belongs to the intended analysis.
  9. Classify the problem as measurement, interpretation, policy, source, or product defect.
  10. Apply the least destructive correction and rerun the same source/inputs where possible.

3. Failure mode: LOC or issue count as productivity

A ranking such as “Developer A wrote 1,200 LOC and Developer B wrote 500 LOC, therefore A was 2.4× more productive” rewards verbosity, generated code, copy/paste, and risky changes. Likewise, “fewest Sonar issues wins” encourages suppressions, exclusions, avoiding difficult code, or splitting work strategically.

Repair: keep Sonar metrics at the code/system/change level. Use delivery outcomes, peer review, reliability, customer impact, operational quality, and context-sensitive engineering evidence for organizational decisions—not an individual metric leaderboard.

4. Failure mode: complexity interpreted as runtime performance

Cyclomatic complexity counts decision paths; cognitive complexity models understandability. A jump from 20 to 35 says nothing quantitative about response time. If performance matters, preserve a representative workload, latency/throughput distribution, CPU/memory/I/O telemetry, environment versions, and benchmark methodology.

Complexity can still be operationally relevant: highly branched code may be harder to test and change safely. That is a maintainability/testing argument—not a performance benchmark.

5. Intentionally broken example: averaging percentages without weights

A release dashboard receives two module summaries and calculates:

module_a_coverage = 90.0  # 9 of 10 executable lines covered
module_b_coverage = 50.0  # 500 of 1000 executable lines covered
reported = (module_a_coverage + module_b_coverage) / 2
print(reported)  # 70.0 -- misleading combined result

Preserve the original dashboard calculation and module evidence. The corrected line-only illustration uses counts:

covered = 9 + 500
executable = 10 + 1000
combined_line_coverage = 100.0 * covered / executable
print(round(combined_line_coverage, 2))  # about 50.4

For SonarQube’s full coverage metric, conditions are also part of the formula; prefer the project-level aggregate from SonarQube or combine the documented underlying counts correctly. Chapter 14 develops coverage semantics in detail.

6. Failure mode: “A means correct and secure”

An A maintainability rating means the measured debt ratio is in the A band under the current grid. An A reliability/security rating means the relevant issue population does not cross lower-grade severity thresholds for that mode. None of those statements proves absence of runtime defects, undiscovered vulnerabilities, unsafe architecture, wrong requirements, dependency risk, or operational failure.

Static analysis is one evidence source. Security Hotspots, dependency/SCA analysis, dynamic testing, penetration testing, observability, tests, reviews, threat modeling, and production behavior remain separate evidence layers.

7. Failure mode: a beautiful trend caused by scope drift

Imagine complexity drops 30%, LOC drops 35%, and duplication improves on one analysis. Before celebrating, inspect scanner logs and settings. If vendor/** or a generated subtree was newly excluded, the trend discontinuity may be legitimate—but it is a scope change, not a refactoring result.

Preserve: revision + effective scope + indexed file count + exclusions + analyzer version + metric snapshot
Then classify: code improvement, measurement correction, or both.

8. Failure mode: incomparable projects

A Python API and a generated Java SDK may have different analyzer behavior, code-generation patterns, test tooling, LOC distributions, and complexity characteristics. Ranking them by “issues per KLOC” can still be misleading because rule profiles, severity distributions, and domain risk differ.

If leadership needs a portfolio view, use governed portfolio/application features where available and preserve each metric’s semantics. Community Build learners can simulate the governance record with a table of project context rather than invent a universal score.

9. Causal matrix for unexpected metric movement

Observed change First layers to inspect Do not do first
LOC suddenly falls Indexing scope, exclusions, language recognition, branch/revision Declare productivity improvement
Maintainability debt rises New issues, rule/profile/analyzer changes, remediation functions, scope Disable the rule or lower the gate
Coverage collapses Coverage report generation/path, executable-line changes, New Code population Exclude uncovered code
Reliability/security rating changes Instance mode, issue severity/impact, profile/analyzer version, workflow state Switch mode or accept issues blindly
Complexity falls but source seems unchanged Revision, indexed files, analyzer version, language parsing Infer runtime optimization

10. Production practices for metric governance

  • Version and review dashboards/reports like operational code.
  • Use stable metric keys and validate them against /api/metrics during upgrades/mode changes.
  • Annotate analyzer, profile, scope, mode, and baseline changes in trend reports.
  • Separate decision metrics from diagnostic metrics and state which is which.
  • Retain raw API evidence for release/audit decisions; screenshots alone are weak provenance.
  • Prefer read-only Browse tokens for metrics extraction; never use administrator tokens for routine dashboards.
  • Never edit SonarQube database/search state directly to “repair” a metric.

Knowledge check

A project’s LOC falls 40% with no known refactor. What do you inspect first?

Why is a metric screenshot insufficient first-failure evidence?

What is wrong with fixing a bad trend by excluding the affected directory?

Does an A security rating prove no exploitable vulnerability exists?

When a rating changes after an instance mode switch, is that automatically source regression?

Next lesson

Produce a defensible metric dossier

Lesson 5 turns the chapter into a repeatable checkpoint with two revisions, explicit predictions, raw evidence, limitations, and a multi-signal recommendation.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-07. Mandatory executable examples target SonarQube Community Build 26.9.0.129388 and SonarScanner CLI 8.1.0.6389. New Community Build instances use MQR Mode by default, but upgraded instances can remain in Standard Experience; therefore examples discover available metric keys and record the actual instance mode rather than hard-coding one issue/rating vocabulary. Core structural keys used in both modes include ncloc, lines, complexity, cognitive_complexity, coverage, and duplicated_lines_density. MQR and Standard issue/rating families are treated as distinct. Re-check primary documentation and the instance’s built-in Web API before automating against another release.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.