Metrics: Reliability, Maintainability, Security, Complexity, and Size: Diagnostics, Failure Modes, and Production Practices
Diagnose metric misuse such as LOC-as-productivity, complexity-as-performance, percentage summation, A-rating overclaims, scope drift, and incomparable project baselines without changing policy to manufacture a desired story.
Learning objectives
- Detect and reject LOC or issue-count productivity rankings and explain the incentive and measurement failures they create.
- Explain why cyclomatic/cognitive complexity does not measure runtime speed, latency, throughput, or resource consumption.
- Diagnose incorrect percentage aggregation and repair it with underlying counts or SonarQube project-level measures.
- Treat an A rating as bounded evidence about one metric family rather than proof of correctness, absence of exploitable defects, or delivery readiness.
- Preserve source-scope, profile, New Code definition, analyzer version, mode, and task evidence when a metric changes unexpectedly.
- Follow an evidence-first diagnostic sequence that separates scanner/indexing, analyzer/rule, coverage/report, policy, server/Compute Engine, and metric-interpretation causes.
1. Metric incidents are often interpretation incidents
When a dashboard “looks bad,” teams often change code, exclusions, rules, or thresholds before proving what changed. That destroys the evidence needed to distinguish source behavior from measurement-system drift. Chapter 13 applies the same incident discipline as scanner/server troubleshooting: preserve first evidence, classify the layer, make the least destructive correction, and rerun the smallest equivalent scenario.
2. Evidence-first diagnostic sequence
-
Preserve the metric snapshot, dashboard/export, API response,
scanner log,
report-task.txt, CE task, revision, and timestamp. - Confirm SonarQube edition/version and instance mode.
- Confirm scanner/analyzer/plugin/runtime versions and source revision.
- Confirm base directory, indexed files, source/test scope, inclusions/exclusions, generated/vendor treatment, and SCM state.
- Confirm profile/rule parameters, issue workflow state, New Code definition/reference, and gate assignment.
- Confirm external report provenance when coverage/test/external analyzer measures are involved.
- Inspect the raw metric definition/key and underlying numerator/denominator or issue population.
- Inspect CE success and analysis history to ensure the measure belongs to the intended analysis.
- Classify the problem as measurement, interpretation, policy, source, or product defect.
- Apply the least destructive correction and rerun the same source/inputs where possible.
3. Failure mode: LOC or issue count as productivity
A ranking such as “Developer A wrote 1,200 LOC and Developer B wrote 500 LOC, therefore A was 2.4× more productive” rewards verbosity, generated code, copy/paste, and risky changes. Likewise, “fewest Sonar issues wins” encourages suppressions, exclusions, avoiding difficult code, or splitting work strategically.
4. Failure mode: complexity interpreted as runtime performance
Cyclomatic complexity counts decision paths; cognitive complexity models understandability. A jump from 20 to 35 says nothing quantitative about response time. If performance matters, preserve a representative workload, latency/throughput distribution, CPU/memory/I/O telemetry, environment versions, and benchmark methodology.
Complexity can still be operationally relevant: highly branched code may be harder to test and change safely. That is a maintainability/testing argument—not a performance benchmark.
5. Intentionally broken example: averaging percentages without weights
A release dashboard receives two module summaries and calculates:
module_a_coverage = 90.0 # 9 of 10 executable lines covered
module_b_coverage = 50.0 # 500 of 1000 executable lines covered
reported = (module_a_coverage + module_b_coverage) / 2
print(reported) # 70.0 -- misleading combined result
Preserve the original dashboard calculation and module evidence. The corrected line-only illustration uses counts:
covered = 9 + 500
executable = 10 + 1000
combined_line_coverage = 100.0 * covered / executable
print(round(combined_line_coverage, 2)) # about 50.4
For SonarQube’s full coverage metric, conditions are
also part of the formula; prefer the project-level aggregate from
SonarQube or combine the documented underlying counts correctly.
Chapter 14 develops coverage semantics in detail.
6. Failure mode: “A means correct and secure”
An A maintainability rating means the measured debt ratio is in the A band under the current grid. An A reliability/security rating means the relevant issue population does not cross lower-grade severity thresholds for that mode. None of those statements proves absence of runtime defects, undiscovered vulnerabilities, unsafe architecture, wrong requirements, dependency risk, or operational failure.
Static analysis is one evidence source. Security Hotspots, dependency/SCA analysis, dynamic testing, penetration testing, observability, tests, reviews, threat modeling, and production behavior remain separate evidence layers.
7. Failure mode: a beautiful trend caused by scope drift
Imagine complexity drops 30%, LOC drops 35%, and duplication
improves on one analysis. Before celebrating, inspect scanner logs
and settings. If vendor/** or a generated subtree was
newly excluded, the trend discontinuity may be legitimate—but it is
a scope change, not a refactoring result.
revision + effective scope + indexed file count + exclusions +
analyzer version + metric snapshotThen classify: code improvement, measurement correction, or both.
8. Failure mode: incomparable projects
A Python API and a generated Java SDK may have different analyzer behavior, code-generation patterns, test tooling, LOC distributions, and complexity characteristics. Ranking them by “issues per KLOC” can still be misleading because rule profiles, severity distributions, and domain risk differ.
If leadership needs a portfolio view, use governed portfolio/application features where available and preserve each metric’s semantics. Community Build learners can simulate the governance record with a table of project context rather than invent a universal score.
9. Causal matrix for unexpected metric movement
| Observed change | First layers to inspect | Do not do first |
|---|---|---|
| LOC suddenly falls | Indexing scope, exclusions, language recognition, branch/revision | Declare productivity improvement |
| Maintainability debt rises | New issues, rule/profile/analyzer changes, remediation functions, scope | Disable the rule or lower the gate |
| Coverage collapses | Coverage report generation/path, executable-line changes, New Code population | Exclude uncovered code |
| Reliability/security rating changes | Instance mode, issue severity/impact, profile/analyzer version, workflow state | Switch mode or accept issues blindly |
| Complexity falls but source seems unchanged | Revision, indexed files, analyzer version, language parsing | Infer runtime optimization |
10. Production practices for metric governance
- Version and review dashboards/reports like operational code.
-
Use stable metric keys and validate them against
/api/metricsduring upgrades/mode changes. - Annotate analyzer, profile, scope, mode, and baseline changes in trend reports.
- Separate decision metrics from diagnostic metrics and state which is which.
- Retain raw API evidence for release/audit decisions; screenshots alone are weak provenance.
- Prefer read-only Browse tokens for metrics extraction; never use administrator tokens for routine dashboards.
- Never edit SonarQube database/search state directly to “repair” a metric.
Knowledge check
A project’s LOC falls 40% with no known refactor. What do you inspect first?
Revision, scanner indexing/scope, exclusions, language recognition, and analyzer/configuration changes before drawing any engineering conclusion.
Why is a metric screenshot insufficient first-failure evidence?
It may omit metric key, population, revision, CE task, mode, profile, scope, and raw values needed to reproduce the result.
What is wrong with fixing a bad trend by excluding the affected directory?
It changes the measurement population and can hide risk. Scope changes require independent technical justification and explicit change history.
Does an A security rating prove no exploitable vulnerability exists?
No. It describes SonarQube’s detected issue population under current analyzers/rules/mode, not the complete security state of the system.
When a rating changes after an instance mode switch, is that automatically source regression?
No. Mode changes alter classifications/metric families. Preserve the mode-change event and compare like with like.
Official references and version notes
- Understanding measures and metrics — current metric definitions and metric keys for software qualities, maintainability, coverage, duplication, size, complexity, and issues.
- Code metrics introduction — how metrics participate in rules, quality gates, monitoring, and mode-dependent UI behavior.
- Changing instance modes — MQR versus Standard Experience classifications, severities, and metric-family implications.
- Instance mode overview — current MQR/Standard model and the default for new Community Build instances.
-
Web API
— bearer authentication,
/api/measures,/api/metrics, and Web API V2 migration guidance. - Understanding quality gates — how selected measures become enforceable policy rather than descriptive dashboards.
-
New Code
— population semantics used by
new_*measures. - Analysis overview — scanner/report/Compute Engine lifecycle behind persisted measures.
- Test coverage overview — current external-coverage model; Chapter 14 expands this subject.
- SonarQube downloads — current Community Build release identity.
- SonarScanner CLI 8.1.0.6389 — scanner baseline used by the local lab.
Rechecked 2026-09-07. Mandatory executable examples target
SonarQube Community Build 26.9.0.129388 and
SonarScanner CLI 8.1.0.6389. New Community Build
instances use MQR Mode by default, but upgraded instances can
remain in Standard Experience; therefore examples discover
available metric keys and record the actual instance mode rather
than hard-coding one issue/rating vocabulary. Core structural keys
used in both modes include ncloc, lines,
complexity, cognitive_complexity,
coverage, and duplicated_lines_density.
MQR and Standard issue/rating families are treated as distinct.
Re-check primary documentation and the instance’s built-in Web API
before automating against another release.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.