Metrics: Reliability, Maintainability, Security, Complexity, and Size: Core Concepts and Mental Model
Learn how SonarQube turns indexed code, findings, structural analysis, and imported test evidence into measures, ratios, ratings, new-code views, and trends—and why none of those values is a context-free performance score.
Learning objectives
- Distinguish a metric definition, a measured value, an issue count, a remediation estimate, a ratio, a rating, and a quality-gate result.
- Explain how indexed source, analyzer findings, structural analysis, SCM/new-code state, and imported test evidence feed different metric families.
- Interpret reliability, security, and maintainability ratings according to the instance mode instead of mixing MQR and Standard Experience keys.
- Separate cyclomatic complexity, cognitive complexity, size/LOC, duplication, coverage, and remediation effort from runtime performance or developer productivity.
- Compare New Code and Overall Code measures without assuming that one is a substitute for the other.
- Use read-only API and UI inspection to record metric keys, revision, analysis task, and measurement context before making an engineering claim.
1. Current baseline and the practical problem
Chapter 12 made the New Code population explicit. Chapter 13 now asks a harder question: when SonarQube says complexity = 38, maintainability rating = A, coverage = 72.5%, or technical debt = 46 minutes, what exactly does that value mean—and what does it not mean?
A metric becomes useful engineering evidence only when its definition, population, revision, analyzer/profile context, and measurement inputs are known. A number without those facts is easy to turn into a vanity score. The chapter therefore treats measures as causal observations, not employee rankings or universal product scores.
2. Mental model: from analyzed state to engineering interpretation
Before reading the diagram, name each state. Indexed code is the source/test file population admitted by scanner scope. Issues are rule findings after analyzer execution. Structural analysis calculates properties such as LOC and complexity. Imported reports provide evidence such as coverage that SonarQube does not generate by running tests itself. The Compute Engine persists results against the project/revision, after which the UI/API expose measures and derived ratings.
flowchart TD A[Indexed code + revision] --> B[Analyzers and structural analysis] C[Coverage / external evidence] --> B B --> D[Issues + raw measures] D --> E[Effort / ratios / ratings] D --> F[New-code and overall populations] E --> G[Trend + gate + diagnosis] F --> G G --> H[Engineering decision with context]
The arrows matter. A rating is not measured directly from source. For example, maintainability rating is derived from a technical-debt ratio; the ratio depends on remediation effort and estimated development cost, which in turn depends on analyzed LOC. Reliability/security ratings depend on issue impact/severity semantics and the instance mode. Complexity is a separate structural measure and does not automatically lower a rating unless a rule raises an issue because of that complexity.
revision → indexed files → profile/mode/analyzers → report/task →
raw measures + issues → ratios/ratings → new/overall view →
gate/trend → engineering interpretation
3. Six terms that must stay separate
| Term | Meaning | Example | Common misuse |
|---|---|---|---|
| Metric | A defined quantity and key. |
ncloc, complexity,
coverage
|
Calling the definition itself a project result. |
| Measure | The value of a metric for a component/analysis/population. | ncloc=412 |
Comparing values without matching scope/revision. |
| Issue count | Population of rule findings under workflow state. | Reliability-impacting issues | Equating count with developer quality. |
| Effort / technical debt | Estimated remediation effort derived from rule remediation functions. | Maintainability remediation effort | Treating minutes as an accounting liability or delivery estimate. |
| Ratio / density | A normalized relationship between underlying quantities. | Debt ratio, duplicated-lines density | Summing percentages across modules. |
| Rating | A categorical grade produced from defined thresholds or issue severity. | A–E maintainability/reliability/security rating | Treating A as proof of correctness/security. |
4. Record the instance mode before interpreting software-quality metrics
Community Build supports MQR Mode and Standard Experience. New instances default to MQR, but upgraded instances can retain Standard Experience. The two modes deliberately use different issue/severity and metric families.
| Area | MQR example keys | Standard Experience example keys |
|---|---|---|
| Reliability |
software_quality_reliability_issues,
software_quality_reliability_rating
|
bugs, reliability_rating |
| Security |
software_quality_security_issues,
software_quality_security_rating
|
vulnerabilities, security_rating
|
| Maintainability |
software_quality_maintainability_issues,
software_quality_maintainability_remediation_effort, software_quality_maintainability_debt_ratio,
software_quality_maintainability_rating
|
code_smells, sqale_index,
sqale_debt_ratio, sqale_rating
|
MQR grades reliability/security according to the worst software-quality impact level (Info/Low/Medium/High/Blocker). Standard Experience uses the older Bug/Vulnerability/Code Smell model with Info/Minor/Major/Critical/Blocker severities. Do not combine one mode’s count with the other mode’s rating in the same report.
5. Maintainability: effort, debt ratio, and rating
Technical debt is SonarQube’s estimated remediation effort for maintainability issues. It is not a financial liability and not a promise that a developer will need exactly that many minutes. SonarQube displays an 8-hour day when effort is shown in days.
The technical debt ratio compares remediation effort with an estimated development cost. Current documentation expresses it conceptually as:
debt ratio = remediation effort / (cost to develop one LOC ×
analyzed LOC)
The default maintainability rating grid is A at or below 5%, B from 5% to under 10%, C from 10% to under 20%, D from 20% to under 50%, and E at/above 50%. Because the rating is a band, remediation effort can increase while the letter remains A. Conversely, changing analyzed LOC or the configured rating grid can affect the ratio/grade without a comparable change in code behavior.
6. Complexity: structural difficulty, not runtime performance
Cyclomatic complexity counts control-flow paths according to language-specific analyzer rules. Cognitive complexity estimates how difficult control flow is to understand. They answer questions about structure and comprehension; neither measures CPU time, latency, memory, throughput, database load, or scalability.
A function can become more complex yet run faster after an optimization, or become simpler yet run slower because of an expensive I/O call. Runtime claims require performance evidence, not a Sonar complexity value.
7. Size, duplication, and coverage need denominators
ncloc counts physical non-comment lines under
SonarQube’s language rules; lines counts physical
lines. Neither indicates features delivered, value produced,
developer effort, or quality. Duplication is expressed through
duplicated lines/blocks and densities. Coverage combines
executable-line and condition evidence imported from test tooling.
Ratios depend on denominators. A 90% coverage measure over 10 lines and 90% over 100,000 lines are not the same evidence volume. Likewise, 2% duplication in a generated/vendor-heavy scope means something different from 2% after those files are correctly excluded.
8. New Code versus Overall Code
Many metrics have new_* counterparts. Chapter 12
established that the New Code population is policy-defined.
Therefore new_coverage, new_violations,
new_software_quality_maintainability_issues, or
new_duplicated_lines_density cannot be interpreted
without the New Code definition/reference.
Use New Code measures to govern current changes and Overall Code measures to understand accumulated state. A green New Code view can coexist with substantial legacy debt; an ugly Overall Code view can coexist with disciplined recent changes.
9. Read-only metric discovery before analysis changes
export SONAR_HOST_URL="http://localhost:9000"
export PROJECT_KEY="academy-sq-ch13"
# Server identity is safe and unauthenticated on a normal local lab.
curl -fsS "$SONAR_HOST_URL/api/server/version"
# Discover metric definitions/keys from the running instance.
curl -fsS -H "Authorization: Bearer $SONAR_API_TOKEN" --get \
--data-urlencode "ps=500" \
"$SONAR_HOST_URL/api/metrics/search" > evidence/metric-catalog.json
# Mode-neutral structural measures.
curl -fsS -H "Authorization: Bearer $SONAR_API_TOKEN" --get \
--data-urlencode "component=$PROJECT_KEY" \
--data-urlencode "metricKeys=ncloc,lines,complexity,cognitive_complexity,coverage,duplicated_lines_density" \
"$SONAR_HOST_URL/api/measures/component" > evidence/structural-measures.json
Then record the instance mode from Administration → Configuration → General Settings → Mode (or from the approved configuration evidence for your deployment) and select the matching issue/rating keys. Do not switch the mode merely to make a report easier to compare.
10. DevOps connection: a metric claim needs a reproducibility envelope
A defensible statement is not “complexity improved.” It is closer to: “For revision X, with scanner Y, profile Z, MQR mode, New Code baseline B, and Compute Engine task T, project-level cognitive complexity decreased from A to B while LOC and coverage stayed materially stable.” That wording reveals what was measured and what competing explanations were controlled.
Knowledge check
Can a project’s complexity increase while its maintainability rating remains A?
Yes. Complexity is a structural measure. Maintainability rating is derived from maintainability remediation effort/debt ratio; complexity affects it only indirectly if active rules raise maintainability issues.
Why must the instance mode be recorded with reliability/security metrics?
MQR and Standard Experience use different issue classifications, severities, and metric keys, so mixing them produces an invalid comparison.
Does ncloc=20,000 mean a team was twice as
productive as a project with 10,000 LOC?
No. LOC measures analyzed code size under language rules, not delivered value, productivity, effort, or quality.
Why can remediation effort rise without the maintainability letter changing?
The rating is a band derived from the debt ratio. A changed value can stay inside the same A–E band.
Which evidence would you need before comparing
new_coverage across two projects?
At minimum the coverage inputs, New Code definitions/references, source scope, revisions, analyzer/scanner context, and ideally comparable language/project structure.
Official references and version notes
- Understanding measures and metrics — current metric definitions and metric keys for software qualities, maintainability, coverage, duplication, size, complexity, and issues.
- Code metrics introduction — how metrics participate in rules, quality gates, monitoring, and mode-dependent UI behavior.
- Changing instance modes — MQR versus Standard Experience classifications, severities, and metric-family implications.
- Instance mode overview — current MQR/Standard model and the default for new Community Build instances.
-
Web API
— bearer authentication,
/api/measures,/api/metrics, and Web API V2 migration guidance. - Understanding quality gates — how selected measures become enforceable policy rather than descriptive dashboards.
-
New Code
— population semantics used by
new_*measures. - Analysis overview — scanner/report/Compute Engine lifecycle behind persisted measures.
- Test coverage overview — current external-coverage model; Chapter 14 expands this subject.
- SonarQube downloads — current Community Build release identity.
- SonarScanner CLI 8.1.0.6389 — scanner baseline used by the local lab.
Rechecked 2026-09-07. Mandatory executable examples target
SonarQube Community Build 26.9.0.129388 and
SonarScanner CLI 8.1.0.6389. New Community Build
instances use MQR Mode by default, but upgraded instances can
remain in Standard Experience; therefore examples discover
available metric keys and record the actual instance mode rather
than hard-coding one issue/rating vocabulary. Core structural keys
used in both modes include ncloc, lines,
complexity, cognitive_complexity,
coverage, and duplicated_lines_density.
MQR and Standard issue/rating families are treated as distinct.
Re-check primary documentation and the instance’s built-in Web API
before automating against another release.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.