Chapter 13Lesson 01~110 minutes

Metrics: Reliability, Maintainability, Security, Complexity, and Size: Core Concepts and Mental Model

Learn how SonarQube turns indexed code, findings, structural analysis, and imported test evidence into measures, ratios, ratings, new-code views, and trends—and why none of those values is a context-free performance score.

MetricsMeasuresRatingsComplexityEvidence

Learning objectives

  • Distinguish a metric definition, a measured value, an issue count, a remediation estimate, a ratio, a rating, and a quality-gate result.
  • Explain how indexed source, analyzer findings, structural analysis, SCM/new-code state, and imported test evidence feed different metric families.
  • Interpret reliability, security, and maintainability ratings according to the instance mode instead of mixing MQR and Standard Experience keys.
  • Separate cyclomatic complexity, cognitive complexity, size/LOC, duplication, coverage, and remediation effort from runtime performance or developer productivity.
  • Compare New Code and Overall Code measures without assuming that one is a substitute for the other.
  • Use read-only API and UI inspection to record metric keys, revision, analysis task, and measurement context before making an engineering claim.

1. Current baseline and the practical problem

Chapter 12 made the New Code population explicit. Chapter 13 now asks a harder question: when SonarQube says complexity = 38, maintainability rating = A, coverage = 72.5%, or technical debt = 46 minutes, what exactly does that value mean—and what does it not mean?

A metric becomes useful engineering evidence only when its definition, population, revision, analyzer/profile context, and measurement inputs are known. A number without those facts is easy to turn into a vanity score. The chapter therefore treats measures as causal observations, not employee rankings or universal product scores.

Governance boundary: never use LOC, issue counts, complexity, debt estimates, or letter ratings as a direct measure of developer productivity, individual performance, or business value. These measures describe selected properties of analyzed code under a particular configuration.

2. Mental model: from analyzed state to engineering interpretation

Before reading the diagram, name each state. Indexed code is the source/test file population admitted by scanner scope. Issues are rule findings after analyzer execution. Structural analysis calculates properties such as LOC and complexity. Imported reports provide evidence such as coverage that SonarQube does not generate by running tests itself. The Compute Engine persists results against the project/revision, after which the UI/API expose measures and derived ratings.

Metric causality, not a score pipeline
flowchart TD
A[Indexed code + revision] --> B[Analyzers and structural analysis]
C[Coverage / external evidence] --> B
B --> D[Issues + raw measures]
D --> E[Effort / ratios / ratings]
D --> F[New-code and overall populations]
E --> G[Trend + gate + diagnosis]
F --> G
G --> H[Engineering decision with context]

The arrows matter. A rating is not measured directly from source. For example, maintainability rating is derived from a technical-debt ratio; the ratio depends on remediation effort and estimated development cost, which in turn depends on analyzed LOC. Reliability/security ratings depend on issue impact/severity semantics and the instance mode. Complexity is a separate structural measure and does not automatically lower a rating unless a rule raises an issue because of that complexity.

Evidence chain
revision → indexed files → profile/mode/analyzers → report/task → raw measures + issues → ratios/ratings → new/overall view → gate/trend → engineering interpretation

3. Six terms that must stay separate

Term Meaning Example Common misuse
Metric A defined quantity and key. ncloc, complexity, coverage Calling the definition itself a project result.
Measure The value of a metric for a component/analysis/population. ncloc=412 Comparing values without matching scope/revision.
Issue count Population of rule findings under workflow state. Reliability-impacting issues Equating count with developer quality.
Effort / technical debt Estimated remediation effort derived from rule remediation functions. Maintainability remediation effort Treating minutes as an accounting liability or delivery estimate.
Ratio / density A normalized relationship between underlying quantities. Debt ratio, duplicated-lines density Summing percentages across modules.
Rating A categorical grade produced from defined thresholds or issue severity. A–E maintainability/reliability/security rating Treating A as proof of correctness/security.

4. Record the instance mode before interpreting software-quality metrics

Community Build supports MQR Mode and Standard Experience. New instances default to MQR, but upgraded instances can retain Standard Experience. The two modes deliberately use different issue/severity and metric families.

Area MQR example keys Standard Experience example keys
Reliability software_quality_reliability_issues, software_quality_reliability_rating bugs, reliability_rating
Security software_quality_security_issues, software_quality_security_rating vulnerabilities, security_rating
Maintainability software_quality_maintainability_issues, software_quality_maintainability_remediation_effort, software_quality_maintainability_debt_ratio, software_quality_maintainability_rating code_smells, sqale_index, sqale_debt_ratio, sqale_rating

MQR grades reliability/security according to the worst software-quality impact level (Info/Low/Medium/High/Blocker). Standard Experience uses the older Bug/Vulnerability/Code Smell model with Info/Minor/Major/Critical/Blocker severities. Do not combine one mode’s count with the other mode’s rating in the same report.

5. Maintainability: effort, debt ratio, and rating

Technical debt is SonarQube’s estimated remediation effort for maintainability issues. It is not a financial liability and not a promise that a developer will need exactly that many minutes. SonarQube displays an 8-hour day when effort is shown in days.

The technical debt ratio compares remediation effort with an estimated development cost. Current documentation expresses it conceptually as:

debt ratio = remediation effort / (cost to develop one LOC × analyzed LOC)

The default maintainability rating grid is A at or below 5%, B from 5% to under 10%, C from 10% to under 20%, D from 20% to under 50%, and E at/above 50%. Because the rating is a band, remediation effort can increase while the letter remains A. Conversely, changing analyzed LOC or the configured rating grid can affect the ratio/grade without a comparable change in code behavior.

6. Complexity: structural difficulty, not runtime performance

Cyclomatic complexity counts control-flow paths according to language-specific analyzer rules. Cognitive complexity estimates how difficult control flow is to understand. They answer questions about structure and comprehension; neither measures CPU time, latency, memory, throughput, database load, or scalability.

A function can become more complex yet run faster after an optimization, or become simpler yet run slower because of an expensive I/O call. Runtime claims require performance evidence, not a Sonar complexity value.

7. Size, duplication, and coverage need denominators

ncloc counts physical non-comment lines under SonarQube’s language rules; lines counts physical lines. Neither indicates features delivered, value produced, developer effort, or quality. Duplication is expressed through duplicated lines/blocks and densities. Coverage combines executable-line and condition evidence imported from test tooling.

Ratios depend on denominators. A 90% coverage measure over 10 lines and 90% over 100,000 lines are not the same evidence volume. Likewise, 2% duplication in a generated/vendor-heavy scope means something different from 2% after those files are correctly excluded.

8. New Code versus Overall Code

Many metrics have new_* counterparts. Chapter 12 established that the New Code population is policy-defined. Therefore new_coverage, new_violations, new_software_quality_maintainability_issues, or new_duplicated_lines_density cannot be interpreted without the New Code definition/reference.

Use New Code measures to govern current changes and Overall Code measures to understand accumulated state. A green New Code view can coexist with substantial legacy debt; an ugly Overall Code view can coexist with disciplined recent changes.

9. Read-only metric discovery before analysis changes

export SONAR_HOST_URL="http://localhost:9000"
export PROJECT_KEY="academy-sq-ch13"

# Server identity is safe and unauthenticated on a normal local lab.
curl -fsS "$SONAR_HOST_URL/api/server/version"

# Discover metric definitions/keys from the running instance.
curl -fsS -H "Authorization: Bearer $SONAR_API_TOKEN" --get \
  --data-urlencode "ps=500" \
  "$SONAR_HOST_URL/api/metrics/search" > evidence/metric-catalog.json

# Mode-neutral structural measures.
curl -fsS -H "Authorization: Bearer $SONAR_API_TOKEN" --get \
  --data-urlencode "component=$PROJECT_KEY" \
  --data-urlencode "metricKeys=ncloc,lines,complexity,cognitive_complexity,coverage,duplicated_lines_density" \
  "$SONAR_HOST_URL/api/measures/component" > evidence/structural-measures.json

Then record the instance mode from Administration → Configuration → General Settings → Mode (or from the approved configuration evidence for your deployment) and select the matching issue/rating keys. Do not switch the mode merely to make a report easier to compare.

10. DevOps connection: a metric claim needs a reproducibility envelope

A defensible statement is not “complexity improved.” It is closer to: “For revision X, with scanner Y, profile Z, MQR mode, New Code baseline B, and Compute Engine task T, project-level cognitive complexity decreased from A to B while LOC and coverage stayed materially stable.” That wording reveals what was measured and what competing explanations were controlled.

Knowledge check

Can a project’s complexity increase while its maintainability rating remains A?

Why must the instance mode be recorded with reliability/security metrics?

Does ncloc=20,000 mean a team was twice as productive as a project with 10,000 LOC?

Why can remediation effort rise without the maintainability letter changing?

Which evidence would you need before comparing new_coverage across two projects?

Next lesson

Turn the model into a controlled two-revision experiment

Lesson 2 captures a baseline, changes one structural property deliberately, and proves why some measures move while others correctly do not.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-07. Mandatory executable examples target SonarQube Community Build 26.9.0.129388 and SonarScanner CLI 8.1.0.6389. New Community Build instances use MQR Mode by default, but upgraded instances can remain in Standard Experience; therefore examples discover available metric keys and record the actual instance mode rather than hard-coding one issue/rating vocabulary. Core structural keys used in both modes include ncloc, lines, complexity, cognitive_complexity, coverage, and duplicated_lines_density. MQR and Standard issue/rating families are treated as distinct. Re-check primary documentation and the instance’s built-in Web API before automating against another release.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.