Chapter 18Lesson 04~275 minutes

Manual Jobs, Delayed Jobs, Scheduled Pipelines, Rollbacks, Freeze Windows, and Release Controls: Diagnostics, Failure Modes, Security, and Performance

Preserve release evidence first, then diagnose permission, schedule-owner, stale-deployment, freeze-rule, and rollback-identity failures without weakening protections.

DiagnosticsPermissionsSchedule ownerOutdated deploymentFreezeArtifact identity

Learning objectives

  • Use an evidence-first sequence for release-control failures.
  • Diagnose a manual job that is runnable by a broader population than policy intended.
  • Diagnose schedule inactivity caused by owner/ref permission changes.
  • Prevent a delayed stale deployment from overwriting newer state.
  • Repair freeze and rollback misunderstandings without deleting failure history.
Availability baseline (verified 2026-08-22 against GitLab 19.3 current documentation). Manual jobs, manual_confirmation, delayed jobs with start_in, pipeline schedules, deployment freeze periods / CI_DEPLOY_FREEZE, environment/deployment history, the project setting that prevents outdated deployment jobs, and rollback/retry controls are available on GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. Protected environments and deployment approvals—the fine-grained “only these people may deploy” layer—remain Premium/Ultimate. The mandatory chapter path uses only a disposable project, synthetic artifacts/environments, and tiny jobs. No production endpoint, cloud account, paid tier, runner registration, real credential, or release publication is required.

1. Diagnostic sequence: preserve → scope → identify → correct → verify

Release incidents are tempting moments to click Retry, Play, Cancel, or edit YAML immediately. Resist that first impulse. Preserve the state that explains why GitLab behaved as it did:

  1. Record instance/offering, project, pipeline ID/source/SHA, job ID/status/user, environment/deployment ID, schedule/freeze IDs, artifact checksum, and runner where relevant.
  2. Identify the control type: permission, manual/blocking semantics, delay timer, schedule owner/ref, freeze rule, stale-deployment setting, resource-group queue, or rollback artifact.
  3. Inspect API/UI evidence before editing.
  4. Choose the least destructive correction.
  5. Create a new verification event; do not rewrite/delete the original failure.

2. Failure mode: manual job is triggerable by broader roles than policy intended

Symptom: the team believed “only release managers can click deploy,” but a Developer with merge permission can run the ordinary manual job.

Cause: when: manual is a human-start condition. Current GitLab requires permission to merge to the assigned branch. It is not automatically an environment-specific allowlist.

Evidence: inspect branch protection, job ref, actor role, environment protection state, and whether the project is Premium/Ultimate.

Correction: on Premium/Ultimate, use a protected environment with Allowed to deploy / deployment approvals. On Free, narrow protected-ref merge rights where appropriate, separate deployment permissions/project, use target IAM, and document that a manual gate is not independent approval.

Do not “fix” this by pasting a shared deploy token into the job. That converts an authorization-design problem into a credential-sharing problem.

3. Failure mode: schedule owner loses access and pipelines stop

Symptom: a previously healthy schedule becomes Inactive, or protected-ref scheduling fails after membership/protection changes.

Likely cause: the schedule owner was blocked, removed, or no longer has permission to target the protected branch/tag.

glab api "projects/$PROJECT_ID/pipeline_schedules" --paginate \
  --jq '.[] | {id,description,ref,active,owner:.owner.username,next_run_at,last_pipeline:(.last_pipeline.id // null)}'

Correction: a Maintainer/Owner can take ownership according to current GitLab workflow, then revalidate target-ref permissions. Do not create a new schedule blindly while leaving the inactive one; duplicate cron objects can later produce duplicate pipelines.

4. Intentionally broken example: old delayed deployment wakes up after a newer release

Pipeline A creates a delayed deployment. Before its timer fires, Pipeline B deploys a newer commit. If Pipeline A later deploys, it can regress the environment.

Evidence timeline (synthetic)
10:00 Pipeline A sha=aaaa created; deploy_A delayed
10:05 Pipeline B sha=bbbb created
10:08 deploy_B succeeds; environment latest=bbbb
10:10 deploy_A becomes runnable
10:10 EXPECTATION: deploy_A must not overwrite bbbb

If Prevent outdated deployment jobs is enabled, GitLab can fail/disable the stale deployment and show an explicit outdated-deployment message. If it is disabled, serialize the environment with a resource_group and add an explicit freshness check/policy before external mutation.

Preserve the stale job’s failure/disabled state. Do not retry it to make the pipeline green.

5. Nuance: deployment “age” is not simply commit age

Current GitLab documentation notes that stale-deployment ordering is based on job start ordering. Two manual deployment jobs can therefore produce counterintuitive outcomes if an older-commit job is started later. This is another reason to model release identity with job/deployment timestamps and explicit artifact SHA—not a branch-name story.

6. Failure mode: freeze window is treated as authorization

Symptom: a team believes “nobody can deploy during the freeze,” yet an API user or manually constructed deployment path can still mutate the target.

Cause: a freeze period makes CI_DEPLOY_FREEZE available to pipelines created during the freeze. Your CI config decides how deployment jobs react. The freeze does not revoke cloud credentials or replace protected environments.

Repair: verify the freeze cron/time zone, inspect whether the pipeline actually had CI_DEPLOY_FREEZE at creation, review the deployment job’s rules, then align exception authorization with branch/environment/target IAM.

7. Broken rule ordering can make freeze policy unreachable

# BROKEN: broad default-branch rule matches before freeze rule.
deploy:
  script: echo "synthetic"
  rules:
    - if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
      when: on_success
    - if: '$CI_DEPLOY_FREEZE'
      when: manual

# FIXED: exceptional state first.
deploy_fixed:
  script: echo "synthetic"
  rules:
    - if: '$CI_DEPLOY_FREEZE'
      when: manual
      allow_failure: true
    - if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
      when: on_success
    - when: never

Rules stop at the first match. A correct freeze period can still be defeated by incorrect rule order.

8. Failure mode: “rollback” rebuilds different bytes

Symptom: operator selects an old commit, build passes, deployment succeeds, but the resulting package/image checksum differs from the original release.

Cause: mutable dependencies, moving container tags, deleted caches, or toolchain drift caused a rebuild rather than a true redeploy.

Repair: locate the original artifact digest and build pipeline, confirm retention/integrity, then deploy that known artifact through a new controlled recovery path. If the artifact is gone, call the action a rebuild/recovery—not an identical rollback—and increase evidence retention for future releases.

9. Retry is operationally sensitive for deployment jobs

Current deployment-safety guidance warns that retries can permit rollback deployments that bypass normal freshness behavior. Sensitive projects should disable deployment job retries and use a new pipeline to a previous commit when rollback is required. This produces a clearer new pipeline/job/deployment audit trail.

10. Reliability/performance failure: cron herd

Dozens of schedules at 0 * * * * can create runner spikes. Symptoms include pending jobs and delayed maintenance. The fix is not bigger runners by default: distribute schedules, measure runner capacity, use bounded retries, and remove duplicate schedules. On GitLab.com, hosted compute may be quota/billing constrained; on Self-Managed, you own the fleet capacity.

11. Security failure: sensitive value entered as manual job variable

Current GitLab documentation warns that manually specified job variables can be visible to project members who can run/retry the manual job, depending on project visibility. They override same-named variables and are not a secure secret-input channel. If a real credential was entered, revoke/rotate it first, then clean logs/history/access—not the other way around.

12. Compact release-control incident runbook

1. STOP external mutation if unsafe; cancel only the specific stale path after recording IDs.
2. RECORD pipeline/job/deployment/schedule/freeze IDs, SHA, artifact digest, actor, timestamps.
3. CLASSIFY: permission | manual semantics | timer | schedule owner | freeze policy | stale deploy | rollback identity.
4. INSPECT branch/environment permissions, YAML/rules, schedule owner/ref, freeze cron/timezone, project stale-deploy settings.
5. CORRECT the smallest control; do not loosen unrelated protections.
6. VERIFY with a new disposable/synthetic event and independent API/UI evidence.
7. RECOVER known artifact/state if needed.
8. DOCUMENT exception and CLEAN only disposable objects by exact ID.

13. Security-sensitive operations in this chapter

Creating/deleting schedules and freeze periods, changing outdated-deployment settings, canceling/running deployment jobs, changing protected refs/environments, and retrying rollback deployments all change release behavior. Use disposable resources, show before-state, record IDs, and verify cleanup. Project deletion, history rewrite, protection bypass, token creation, real secret entry, runner registration, registry deletion, or production deployment are outside the mandatory path.

Knowledge check

A Developer can run a manual deploy but policy says only release managers should. What is the root problem?

Why can a schedule become inactive even though its YAML is valid?

What evidence distinguishes a stale delayed deployment from a legitimate rollback?

Why might a correct freeze window still fail to block auto-deploy?

What should happen first if a real secret was pasted into a manual job variable?

Summary

Release-control diagnosis is evidence-first: preserve actor/time/SHA/artifact/environment state, identify the exact policy surface, correct the smallest control, and verify without erasing the failure. The most dangerous mistakes come from assuming manual means authorized, delayed means fresh, freeze means secure, or retry means rollback.

Official references

Next lesson

Checkpoint: operate a complete controlled release state machine

Lesson 5 combines the chapter into one disposable evidence-driven lab with manual and timed gates, schedule ownership, freeze policy, a stale path, supersession, and known-artifact recovery.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.