Bounded Acceptance Scheduler Runbook
Purpose and limits
Restart read-only coverage without replaying the historic queue. This pilot does
not enable signatures, donations, cancellations, emails, cleanup or production
monitoring. The existing non-marketing signature check already excludes email
delivery from its signing gate (PR #54); passing that check is not proof of
marketing email delivery. Record an unconfigured acceptance email path as
not_validated, and retain failures for an explicitly required delivery check.
Prepare
- Obtain owner approval of the three checks, 15-minute cadence, two-minute deadline, one worker, zero retries, current runtime identity and report recipients. During the initial pilot the operator posts sanitized reports to the linked Asana task; no inherited email recipient is used.
- Keep
lti-dev-cron-triggerpaused. Save restricted-access JSON exports of the trigger andlti-dev-cronjob, current image digest and runtime settings. Configuration exports can contain sensitive values. Never commit them. - Deploy the reviewed feature image to acceptance only. Do not merge merely to deploy if main automatically deploys production. Retain image SHA/digest.
- Set
LTI_HEALTH_PILOT_DB_NAMEandLTI_HEALTH_PILOT_DB_HOSTto the verified acceptance LTI database and Cloud SQL socket, not the backend database. - Run
python manage.py run_acceptance_health_cycle --inventoryin the job image. Retain the aggregate inventory and its SHA-256. Every old row remains excluded with unverified synthetic ownership. No deletion or rescheduling.
Activate only after approval
Set LTI_SCHEDULER_READ_ONLY_PILOT=true, the exact inventory hash in
LTI_HEALTH_PILOT_BACKLOG_SHA256, and a timezone-aware approval expiry in
LTI_HEALTH_PILOT_EXPIRES_AT (maximum 24 hours). Change the job entrypoint to
python manage.py run_acceptance_health_cycle; do not wrap it in runcrons or
the old session-lock wrapper. Set task count and parallelism to one, timeout
120 seconds, max retries zero. Preserve credentials, network and attachments.
Execute once manually. Check every health result, unchanged backlog hash and
absence of external mutations. An unhealthy FRAPI queue is a real failed check,
not an email-environment exclusion. On success change the trigger to
*/15 * * * *, resume, and observe the next genuinely scheduled execution.
Retain its Cloud Scheduler and Cloud Run evidence separately from the manual
run. Deliver both JSON summaries to the approved recipient in Asana.
The reports count only the three read-only checks. They explicitly show production as not run, mutable journeys paused, and email delivery unvalidated. The old dashboard history is not rewritten or represented as current coverage.
Stop and rollback
Pause the trigger immediately on authentication/query failure, backlog drift, environment mismatch, unexpected mutation or timeout. The runtime fails closed but has no new IAM power to pause the Scheduler itself. Approval expiration also stops further checks; extend only with renewed operator approval.
Rollback by pausing first, waiting for any bounded execution to finish, then restoring the saved job image/entrypoint/configuration. Keep the trigger paused. Restoring the legacy runner is not approval to execute its historic backlog. No queue rows or schemas are modified by the health command, so no data rollback or database migration is required.