Chapter 05Lesson 04~125 minutes

Issues, Labels, Milestones, Templates, Forms, Discussions, and Triage: Diagnostics, Failure Modes, Security, and Performance

Diagnose broken triage systems from evidence: label explosion, low-value forms, lost duplicate context, incident work stranded in Discussions, and over-permissive automation—then repair the system with the least destructive change.

DiagnosticsDuplicatesModerationAutomation safety

Learning objectives

  • Apply an evidence-preserving diagnostic sequence before changing issue metadata or automation.
  • Detect label explosion by measuring semantic overlap and query inconsistency rather than by counting labels alone.
  • Recognize issue forms that collect data without improving reproduction, ownership, or prioritization.
  • Repair duplicate handling while preserving canonical links and original reporter context.
  • Route incident-like work out of Discussions into explicit ownership/escalation records when necessary.
  • Constrain triage automation to least-privilege, reversible actions with audit evidence and human recovery paths.
Availability: All failure examples can be simulated in a disposable GitHub Free public repository or with local fixtures. No organization/enterprise feature is required. Moderation and transfer examples are explanatory unless the learner already has appropriate disposable resources and permissions.

1. Diagnostic sequence: preserve evidence before “cleaning up” the queue

Evidence-first issue diagnosis
flowchart TD
  E[Preserve issue/timeline/query evidence] --> S[Identify repo + issue + actor + permission scope]
  S --> I[Inspect labels, milestone, assignee, state, form config, rules/automation]
  I --> C[Classify cause: taxonomy / intake / permission / routing / automation]
  C --> F[Choose least destructive correction]
  F --> V[Verify query + timeline + canonical link]

Do not begin by deleting labels, closing issues, or rewriting templates. First capture the original issue numbers, metadata, timeline events, form commit OID, automation run/log if relevant, and the exact query that produced the confusing result. The broken state is evidence.

2. Failure mode: label explosion makes filtering meaningless

Suppose the repository has bug, type:bug, kind:defect, severity:bug, urgent, and priority:high. The problem is not merely “too many labels.” The problem is overlapping semantics: different triagers can classify the same work differently, so queries cease to be reproducible.

gh label list --repo OWNER/REPO --limit 200   --json name,description,isDefault --jq '.[] | {name,description,isDefault}'

gh issue list --repo OWNER/REPO --state all --limit 200   --json number,title,labels   --jq '.[] | {number,title,labels:[.labels[].name]}'

Repair by defining dimensions first, mapping old labels to canonical ones, applying the new labels, verifying saved queries, and only then deleting obsolete label definitions. Label deletion is a repository metadata mutation; capture affected issue counts before doing it.

3. Failure mode: a form collects data but does not produce an actionable report

A form with required fields for “department,” “favorite browser,” “screenshot,” and “urgency” may look structured while omitting the only evidence that matters: observed behavior, expected behavior, smallest reproduction, affected version, and impact. Required inputs can therefore increase form completion while decreasing diagnosis quality.

Field Ask: does this change a triage decision? Action
Steps to reproduce Yes—determines reproducibility Keep, required for bug class
Affected version Usually—routing/regression scope Keep
“Urgency: low/medium/high” chosen by reporter Weak—does not prove impact Replace with observable impact questions
Screenshot always required Often no; may expose sensitive data Optional; ask only when visually diagnostic
Contact details in public issue Potential privacy risk Avoid unless genuinely needed; prefer GitHub identity/thread

Repair the form as a normal Git change on a branch, review the diff, merge to the default branch, then create a new test issue to prove the rendered behavior. Old issues remain valid historical records; do not rewrite their bodies to pretend they came from the new form.

4. Failure mode: “duplicate” closure loses the canonical path

Broken behavior: a maintainer closes issue #42 with the comment “duplicate” and no link. Six months later a search finds #42 but there is no way to know which issue survived or whether the underlying bug was fixed.

# Read the issue before changing it.
gh issue view 42 --repo OWNER/REPO --json number,title,state,comments,url --jq .

# Correct pattern in a disposable lab when #17 is the verified canonical record.
gh issue close 42 --repo OWNER/REPO --duplicate-of 17

gh issue view 42 --repo OWNER/REPO --json number,state,stateReason,comments,url --jq .

GitHub’s current CLI supports duplicate-aware close. The repair does not need to delete #42. Preserve reporter-specific context in the duplicate, and move any unique reproduction evidence to the canonical issue with attribution/link rather than silently copying it.

5. Failure mode: a production incident is handled only as a Discussion

A Discussion can contain valuable diagnosis, but it is a poor sole control for an active incident if the team needs an owner, severity/priority, escalation deadline, or release/remediation linkage. The repair is not “delete the Discussion.” Preserve it as context, then create/link the authoritative incident/task record in the approved incident system or a tracked Issue, assign an owner, and define next action.

Do not turn GitHub Issues into an emergency paging system merely because they are trackable. Real incidents often require a dedicated incident platform/on-call channel. The GitHub issue can hold repository remediation/follow-up work.

6. Failure mode: over-permissive automation closes or rewrites work incorrectly

Imagine a workflow with broad repository write permissions that closes any issue containing the word “duplicate.” A malicious or accidental phrase can cause the bot to mutate work. The root cause is not only a bad regex; it is a trust design that lets untrusted issue text directly drive a privileged mutation.

Evidence to inspect Question
Workflow run event + actor What event triggered the decision?
Exact workflow commit SHA Which automation code executed?
Job token permissions Could the job write issues/contents/actions?
Input/title/body used by logic Was untrusted text interpreted safely?
Timeline close/comment actor Can we distinguish bot action from human action?
Reopen/override path Can a human recover without admin escalation?

Repair with least privilege, deterministic criteria, a human-confirmation boundary for semantic decisions, and a visible comment/audit trail. Later Actions chapters will implement this securely; here the important lesson is to diagnose the policy + permission boundary before editing the regex.

7. Intentionally broken example: permission denied while triaging

A collaborator with Triage permission runs a command that tries to create a new label:

gh label create "priority:p0" --repo ORG/REPO   --description "Immediate" --color B60205
# Example failure: API/CLI reports insufficient permission / resource not accessible.

Interpretation: authentication may be valid; the action itself requires stronger repository permission. Under the standard organization role table, Triage can apply/dismiss existing labels but cannot create/edit/delete label definitions. The least-destructive correction is not “give the triager Write forever.” Ask a Write/Maintain/Admin actor to create the approved label definition, or reconsider whether the taxonomy needs it. Then the triager can apply it.

Authorization diagnosis: do not solve a policy/role failure by generating a broader token. First confirm the role/action matrix and whether the operation is actually part of the triager’s job.

8. API and performance failures: large queues need pagination and efficient queries

A script that calls the REST API once and assumes it saw every issue can silently misclassify an old queue. Collections paginate; GitHub also rate-limits API clients. Prefer gh issue list with explicit limits for operator queries and documented pagination for API automation. Avoid fetching every comment/timeline for every issue unless the decision requires it.

gh api -H "X-GitHub-Api-Version: 2026-03-10"   "/repos/OWNER/REPO/issues?state=open&per_page=100" --paginate   --jq '.[] | select(.pull_request == null) | {number,title,updated_at}'

For real automation, inspect rate-limit response headers and design retries/conditional requests appropriately. Do not blind-retry a mutation such as closing or transferring an issue without confirming current state.

9. Security-sensitive and destructive issue operations

  • Deleting an issue: admin-only and permanent; not part of normal triage. Close instead when historical context matters.
  • Transferring an issue: changes repository ownership context and can affect label/milestone mapping; inventory permissions/target first.
  • Locking conversation: moderation control that restricts participation; record reason and use only when appropriate.
  • Automation token/workflow permission changes: security-sensitive; use a disposable repository and least privilege.
  • Repository deletion: destroys the whole hosted work system; only used for explicit lab cleanup after evidence capture.

Branch/tag force updates, secrets, runner registration, package deletion, and history rewrite are not needed for issue triage. If a proposed “fix” requires those operations, you are probably debugging the wrong layer.

10. Evidence-first diagnostic runbook

  1. Record repository owner/name, issue number/URL, current state/reason, labels, assignees, milestone, and relevant form commit OID.
  2. Identify the actor and required role for the attempted action.
  3. Inspect the exact query/filter and whether REST results include pull requests.
  4. Inspect timeline/comments or automation run only as deep as needed to explain the mutation.
  5. Classify root cause: intake design, taxonomy, routing, authorization, automation logic, or API/query handling.
  6. Choose a reversible correction: add/remove metadata, reopen, link canonical issue, revise form in Git, or narrow automation permissions.
  7. Re-run the original query and independently verify the expected state.

11. Lesson summary

A broken triage queue should be debugged like an operational system: preserve state, isolate the layer, correct minimally, and verify. Do not erase duplicates, grant permissions, delete labels, or close work merely to make dashboards look clean.

Knowledge check

Why is “we have 80 labels” not enough evidence of label explosion?

A Triage-role user cannot create a label. Is the token necessarily invalid?

Why should a Discussion from an incident not simply be deleted after creating a tracked issue?

What is dangerous about auto-closing issues based on title/body text?

Why is issue deletion normally worse than closing?

What should you verify after a taxonomy repair?

Next lesson

Operate the full triage checkpoint

Lesson 05 builds a fresh issue taxonomy/form, opens multiple cases, handles a duplicate, verifies the queue with structured queries, and produces a reusable triage runbook.

Authoritative references

 Repository roles
 Managing labels
 Syntax for issue forms
 Marking duplicates
 Moderating discussions
 Locking conversations
 Transferring an issue
 gh issue close
 REST API endpoints for issues
 REST API timeline events
 REST API versions

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.