Manual Jobs, Delayed Jobs, Scheduled Pipelines, Rollbacks, Freeze Windows, and Release Controls: Diagnostics, Failure Modes, Security, and Performance
Preserve release evidence first, then diagnose permission, schedule-owner, stale-deployment, freeze-rule, and rollback-identity failures without weakening protections.
Learning objectives
- Use an evidence-first sequence for release-control failures.
- Diagnose a manual job that is runnable by a broader population than policy intended.
- Diagnose schedule inactivity caused by owner/ref permission changes.
- Prevent a delayed stale deployment from overwriting newer state.
- Repair freeze and rollback misunderstandings without deleting failure history.
manual_confirmation, delayed jobs with
start_in, pipeline schedules, deployment freeze periods /
CI_DEPLOY_FREEZE, environment/deployment history, the
project setting that prevents outdated deployment jobs, and
rollback/retry controls are available on GitLab Free/Premium/Ultimate
across GitLab.com, Self-Managed, and Dedicated. Protected environments
and deployment approvals—the fine-grained “only these people may
deploy” layer—remain Premium/Ultimate. The mandatory chapter path uses
only a disposable project, synthetic artifacts/environments, and tiny
jobs. No production endpoint, cloud account, paid tier, runner
registration, real credential, or release publication is required.
1. Diagnostic sequence: preserve → scope → identify → correct → verify
Release incidents are tempting moments to click Retry, Play, Cancel, or edit YAML immediately. Resist that first impulse. Preserve the state that explains why GitLab behaved as it did:
- Record instance/offering, project, pipeline ID/source/SHA, job ID/status/user, environment/deployment ID, schedule/freeze IDs, artifact checksum, and runner where relevant.
- Identify the control type: permission, manual/blocking semantics, delay timer, schedule owner/ref, freeze rule, stale-deployment setting, resource-group queue, or rollback artifact.
- Inspect API/UI evidence before editing.
- Choose the least destructive correction.
- Create a new verification event; do not rewrite/delete the original failure.
2. Failure mode: manual job is triggerable by broader roles than policy intended
Symptom: the team believed “only release managers can click deploy,” but a Developer with merge permission can run the ordinary manual job.
Cause: when: manual is a human-start
condition. Current GitLab requires permission to merge to the
assigned branch. It is not automatically an environment-specific
allowlist.
Evidence: inspect branch protection, job ref, actor role, environment protection state, and whether the project is Premium/Ultimate.
Correction: on Premium/Ultimate, use a protected environment with Allowed to deploy / deployment approvals. On Free, narrow protected-ref merge rights where appropriate, separate deployment permissions/project, use target IAM, and document that a manual gate is not independent approval.
3. Failure mode: schedule owner loses access and pipelines stop
Symptom: a previously healthy schedule becomes
Inactive, or protected-ref scheduling fails after
membership/protection changes.
Likely cause: the schedule owner was blocked, removed, or no longer has permission to target the protected branch/tag.
glab api "projects/$PROJECT_ID/pipeline_schedules" --paginate \
--jq '.[] | {id,description,ref,active,owner:.owner.username,next_run_at,last_pipeline:(.last_pipeline.id // null)}'
Correction: a Maintainer/Owner can take ownership according to current GitLab workflow, then revalidate target-ref permissions. Do not create a new schedule blindly while leaving the inactive one; duplicate cron objects can later produce duplicate pipelines.
4. Intentionally broken example: old delayed deployment wakes up after a newer release
Pipeline A creates a delayed deployment. Before its timer fires, Pipeline B deploys a newer commit. If Pipeline A later deploys, it can regress the environment.
Evidence timeline (synthetic)
10:00 Pipeline A sha=aaaa created; deploy_A delayed
10:05 Pipeline B sha=bbbb created
10:08 deploy_B succeeds; environment latest=bbbb
10:10 deploy_A becomes runnable
10:10 EXPECTATION: deploy_A must not overwrite bbbb
If Prevent outdated deployment jobs is enabled, GitLab can
fail/disable the stale deployment and show an explicit
outdated-deployment message. If it is disabled, serialize the
environment with a resource_group and add an explicit
freshness check/policy before external mutation.
Preserve the stale job’s failure/disabled state. Do not retry it to make the pipeline green.
5. Nuance: deployment “age” is not simply commit age
Current GitLab documentation notes that stale-deployment ordering is based on job start ordering. Two manual deployment jobs can therefore produce counterintuitive outcomes if an older-commit job is started later. This is another reason to model release identity with job/deployment timestamps and explicit artifact SHA—not a branch-name story.
6. Failure mode: freeze window is treated as authorization
Symptom: a team believes “nobody can deploy during the freeze,” yet an API user or manually constructed deployment path can still mutate the target.
Cause: a freeze period makes
CI_DEPLOY_FREEZE available to pipelines created during
the freeze. Your CI config decides how deployment jobs react. The
freeze does not revoke cloud credentials or replace protected
environments.
Repair: verify the freeze cron/time zone, inspect
whether the pipeline actually had CI_DEPLOY_FREEZE at
creation, review the deployment job’s rules, then align exception
authorization with branch/environment/target IAM.
7. Broken rule ordering can make freeze policy unreachable
# BROKEN: broad default-branch rule matches before freeze rule.
deploy:
script: echo "synthetic"
rules:
- if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
when: on_success
- if: '$CI_DEPLOY_FREEZE'
when: manual
# FIXED: exceptional state first.
deploy_fixed:
script: echo "synthetic"
rules:
- if: '$CI_DEPLOY_FREEZE'
when: manual
allow_failure: true
- if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
when: on_success
- when: never
Rules stop at the first match. A correct freeze period can still be defeated by incorrect rule order.
8. Failure mode: “rollback” rebuilds different bytes
Symptom: operator selects an old commit, build passes, deployment succeeds, but the resulting package/image checksum differs from the original release.
Cause: mutable dependencies, moving container tags, deleted caches, or toolchain drift caused a rebuild rather than a true redeploy.
Repair: locate the original artifact digest and build pipeline, confirm retention/integrity, then deploy that known artifact through a new controlled recovery path. If the artifact is gone, call the action a rebuild/recovery—not an identical rollback—and increase evidence retention for future releases.
9. Retry is operationally sensitive for deployment jobs
Current deployment-safety guidance warns that retries can permit rollback deployments that bypass normal freshness behavior. Sensitive projects should disable deployment job retries and use a new pipeline to a previous commit when rollback is required. This produces a clearer new pipeline/job/deployment audit trail.
10. Reliability/performance failure: cron herd
Dozens of schedules at 0 * * * * can create runner
spikes. Symptoms include pending jobs and delayed maintenance. The
fix is not bigger runners by default: distribute schedules, measure
runner capacity, use bounded retries, and remove duplicate
schedules. On GitLab.com, hosted compute may be quota/billing
constrained; on Self-Managed, you own the fleet capacity.
11. Security failure: sensitive value entered as manual job variable
Current GitLab documentation warns that manually specified job variables can be visible to project members who can run/retry the manual job, depending on project visibility. They override same-named variables and are not a secure secret-input channel. If a real credential was entered, revoke/rotate it first, then clean logs/history/access—not the other way around.
12. Compact release-control incident runbook
1. STOP external mutation if unsafe; cancel only the specific stale path after recording IDs.
2. RECORD pipeline/job/deployment/schedule/freeze IDs, SHA, artifact digest, actor, timestamps.
3. CLASSIFY: permission | manual semantics | timer | schedule owner | freeze policy | stale deploy | rollback identity.
4. INSPECT branch/environment permissions, YAML/rules, schedule owner/ref, freeze cron/timezone, project stale-deploy settings.
5. CORRECT the smallest control; do not loosen unrelated protections.
6. VERIFY with a new disposable/synthetic event and independent API/UI evidence.
7. RECOVER known artifact/state if needed.
8. DOCUMENT exception and CLEAN only disposable objects by exact ID.
13. Security-sensitive operations in this chapter
Creating/deleting schedules and freeze periods, changing outdated-deployment settings, canceling/running deployment jobs, changing protected refs/environments, and retrying rollback deployments all change release behavior. Use disposable resources, show before-state, record IDs, and verify cleanup. Project deletion, history rewrite, protection bypass, token creation, real secret entry, runner registration, registry deletion, or production deployment are outside the mandatory path.
Knowledge check
A Developer can run a manual deploy but policy says only release managers should. What is the root problem?
The manual job’s authorization boundary is broader than the policy. Use protected-environment authorization where available or redesign Free-tier permissions/target IAM; do not treat manual as fine-grained approval.
Why can a schedule become inactive even though its YAML is valid?
Schedule execution depends on the schedule owner and target-ref permissions. If ownership/access changes, the schedule can stop.
What evidence distinguishes a stale delayed deployment from a legitimate rollback?
Pipeline/job/deployment ordering plus an explicit recovery intent and known previous artifact/commit identity. Stale work is accidental old execution; rollback is deliberate new recovery.
Why might a correct freeze window still fail to block auto-deploy?
The job rules may match a broader rule before the freeze condition, or the deployment path may not use the freeze-aware policy at all.
What should happen first if a real secret was pasted into a manual job variable?
Revoke/rotate the credential, then investigate logs/visibility and clean up. Do not rely on deletion as the first containment step.
Summary
Release-control diagnosis is evidence-first: preserve actor/time/SHA/artifact/environment state, identify the exact policy surface, correct the smallest control, and verify without erasing the failure. The most dangerous mistakes come from assuming manual means authorized, delayed means fresh, freeze means secure, or retry means rollback.
Official references
- GitLab Docs — Control how jobs run
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Scheduled pipelines
- GitLab Docs — Pipeline schedules API
- GitLab Docs — Deployment safety
- GitLab Docs — Deployments
- GitLab Docs — Environments
- GitLab Docs — Freeze Periods API
- GitLab Docs — Predefined CI/CD variables
- GitLab Docs — Pipeline settings
- GitLab Docs — Pipelines
- GitLab Docs — Job troubleshooting
- GitLab Docs — Resource groups
- GitLab Docs — Protected environments
- GitLab Docs — Deployment approvals
- GitLab Docs — Job token
- GitLab Docs — Roles and permissions
- GitLab 19.3 — What’s new
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.