Skip to content

Powered by Grav

Ægir task failures

Ægir task failures

The BOA task queue is cron-driven, not daemon-driven: there is no queue daemon to restart. runner.sh scans /var/xdrago/ for run-<USER> dispatchers and runs each — under a load gate (_O_LOAD < _CPU_TASK_RATIO × 100, default ratio 3.1) — which sus to the Octopus user, runs the instance's aegir.sh (driving drush @hostmaster hosting-dispatch), and touches /var/log/boa/last-run-<USER>.

Two rules keep that queue honest. A request for a task of a kind the node already has queued (a verify queued by a re-save, say) is folded into the waiting one instead of queueing a twin — a task that carries arguments (which backup to restore, where to clone to) is always its own request, and a task already running does not count, so a change made while it runs still gets its turn. And a site or platform Delete cancels the node's remaining queued tasks (status Warning, log line Task was cancelled: the site was deleted), so nothing is dispatched later against a context that no longer exists. Before this a busy box carried eighteen identical queued verifies per platform and dozens of queued verifies on deleted platforms, all draining at one task a minute.

A "stuck" task is therefore almost always one of three things:

  • a dispatch that is not firing (load hold, maintenance pause),
  • a task whose backend Drush run aborted, or
  • stale on-disk state Ægir cannot reconcile.

The recovery paths below act on those three classes directly.

Task spinning — status "Processing", nothing happening

Diagnosis

SH
# Is cron firing the dispatcher? (expect a fresh marker within ~1 min)
ls -l /var/log/boa/last-run-<USER>
pgrep -af 'runner.sh'

# Is processing paused for maintenance, or load-throttled?
ls -l /etc/boa/.pause_tasks_maint.cnf /run/max_load.pid /run/critical_load.pid

# Per-instance Aegir log/state dir (task markers + run state)
ls -lt /data/disk/<USER>/log/ | head

There is nothing to "restart". If a maintenance pause marker is present, remove it to resume dispatch:

SH
rm -f /etc/boa/.pause_tasks_maint.cnf

If /run/max_load.pid or /run/critical_load.pid is present, the box is in a load-control hold (see Load control) and task dispatch is deliberately deferred until load drops — do not force a run into a saturated box; let it clear. One exception: a genuine in-progress Octopus install/upgrade runs the queue despite the hold, honoured only while /run/octopus_install_run.pid is under 15 minutes old or a live /opt/local/bin/octopus process exists; a stale marker is deleted, so a crashed run cannot permanently bypass load protection.

What now heals itself

Several formerly manual recoveries are self-clearing on current code — check the BOA version before reaching for the older workarounds:

  • Dispatcher died holding the queue semaphore. The dispatch lock times out after 15 minutes (HOSTING_QUEUE_LOCK_TIMEOUT, 900 s), not the old hour. The manual semaphore clear shown on the task queue engine page is only needed inside that window.

  • Orphan _tmp_ build directories. A _tmp_ working dir under /data/disk/<USER>/.tmp counts as an active platform build only while a live drush.php process for that instance user exists. The check is re-evaluated on every runner pass and deletes nothing, so a _tmp_ dir left behind by a crashed build no longer wedges the instance queue until an unrelated cache clear — dispatch resumes on the next pass by design.

  • Dispatcher disabled after an upgrade or hostmaster reinstall. On older versions, a failed hosting-setup self-test could leave hosting_dispatch_enabled off, piling tasks up as Queued forever; the manual check is drush8 @hostmaster vget hosting_dispatch_enabled and, if disabled, drush8 @hostmaster vset hosting_dispatch_enabled 1. Current code self-repairs this: hosting-setup preserves a previously-enabled dispatcher when its self-test fails, and BOA checks-and-repairs the variable on Barracuda master and Octopus instance install/upgrade finalise (_ensure_hosting_dispatch_enabled, logging "task queue dispatcher re-enabled" only when it corrected something).

  • A task left Processing by a crashed runner. A hosting task's final status is written only by its runner's own PHP shutdown handler, so a runner killed without reaching it — a host reboot, an interrupted upgrade swapping the code trees under it, a stray signal — leaves the row at Processing, and every such corpse is subtracted from the dispatcher's concurrency budget for the whole 8-hour running-items window. Two independent reapers clear it, and they coexist through a conditional per-row update, so whichever gets there first wins and the other side no-ops. The frontend dispatcher reaps at the start of every pass of an enabled tasks queue, before it counts running tasks, and again inside hosting-pause's wait loop — the wait every hostmaster upgrade runs, which would otherwise spin on the corpse. It takes current-revision Processing rows at least two minutes old (a floor on the row's own age, so a task that has only just started is never judged) and checks the runner PID recorded on the row against /proc; a row whose runner is verifiably gone is marked failed on that same pass, however old it is, and gains the task log line "Task runner process … is gone without reporting a final status; marked failed by the queue reaper." A row carrying no PID — written before the database update that added the column — is reaped once it is past the 8-hour window instead. Independently, BOA's box-side task_guard.sh watchdog reaps the same rows from outside on a roughly five-minute cadence, on every hostmaster generation and with no Drush bootstrap, so the healing still works on an instance whose queue is disabled, whose frontend has not taken that database update, or whose Drush tree the interrupted upgrade broke. It judges on process evidence rather than the PID: a current-revision Processing row older than _TASK_GUARD_GRACE_MINS (10 minutes, floored at 3) is reaped only when the instance's own system user has no live task runner and no live provision backend, so one live process holds the entire instance for that pass. _USE_TASK_GUARD=NO opts a box out and _TASK_GUARD_DETECT_ONLY=YES keeps its detection and alerts without any writes; it also stands down on a replication standby and on a finalised migration proxy. See Auto-healing watchdogs.

What does not heal: nothing re-runs a crashed task. Both reapers mark the row failed and the queue dispatches again on its own — re-firing a half-finished migrate or clone is exactly what a recovery path must not do, so the retry stays your call. A row whose runner is still alive is left alone, as are superseded task revisions, which stay Processing by design and gate nothing.

Force a manual dispatch

SH
# As root. run-<USER> requires root and su's to the instance user itself.
bash /var/xdrago/run-<USER>

This dispatches every queued task for that instance in one pass. The dispatcher serialises itself per instance with a non-blocking flock on /run/run-xdrago-<USER>.lock (held on fd 9, released automatically on exit, failing open if /run or flock is unavailable).

If a runner pass is already in flight for that instance, the command exits cleanly and silently (flock -n 9 || exit 0) — that is de-duplication, not a failure; the in-flight pass is already dispatching the queue.

Delete and rename: what the task removes, and the store it sets aside

A site Delete removes the vhost, the database, the site directory and the site's two links into its per-account static store. Two leftovers used to survive it: a site deleted before its first Verify kept its directory on the platform (a Drupal 11 install left a fresh site directory read-only, 0555, and every Deleting ... failed warning in the task log was one unlink inside it), and the per-account store static/files/<uri> stayed at its live name until the nightly sweep. The task now makes the tree owner-writable before deleting it, and asks the root-side wrapper to set the store aside under static/files/.archived/<stamp>/<uri>/ before the site's alias goes: DELETE/STORE/ARCHIVED on success, DELETE/STORE/NONE when the site had no store, DELETE/STORE/KEPT when the operator's orphan-archiving switch (/data/conf/disable_orphan_store_archiving.cnf) holds it, and a warning DELETE/STORE/LEFT when the wrapper declined (a store named by a share control file, which another site reads; or a store some registered site still reads through its files/private link, such as a clone whose unshare was refused for disk space) or failed, with the reason in /var/log/boa/autosymlink.log. Nothing is ever deleted from the store.

The same read-only site directory used to break the other direction on Drupal 11: Provision records a site as installed by rewriting sites/<uri>/drushrc.php at the end of the Install task, and that rewrite needs a writable site directory. The D11 install runs core's own installer through the site-local Drush, which hardens the directory to 0555 on its way out, so the flag was never written until the site's first Verify set 0755 and rewrote the file. In that window every backup-first task on the site (Backup, Clone, Migrate) aborted with Site was not found — on a busy queue, for as long as the Verify stayed behind it. The Install task now restores the directory's owner write bit right after core's installer returns (D11/INSTALL/SITEDIR restored ... in the task log), so a fresh D11 site can be cloned at once; Drupal 6–10 installs run in-process and were never affected.

A Migrate under a new name (the rename path) does the same for the old name's store once the renamed site has been re-homed into its own copy, and before the old alias goes: the identical lines under a RENAME/ prefix (RENAME/STORE/ARCHIVED, NONE, KEPT, LEFT). If the re-home itself was refused (the [ALERT] … unshare … not completed line, usually disk space), the renamed site still reads the old store through its links, so the wrapper leaves that store where it is (RENAME/STORE/LEFT) until the unshare is re-run by hand — see site cloning for the recovery command and orphans & archiving for the archive.

"Could not delete this site" / "Could not delete this platform"

The Delete task fails when files were removed manually from the platform's sites/ directory before the Ægir Delete task ran — Ægir tries to back up or verify a tree that is already gone.

Recovery — clean the Ægir record by hand

{hosting_site} carries no title column; the human-readable name is the {node} title, and the site-status column is status (not site_statussite_status is only the loaded-node PHP property).

So the record cleanup must join against {node} on title, exactly as the companion task-delete already does:

SH
# Mark the site node deleted (HOSTING_SITE_DELETED = -2), keyed via {node}.title
drush @hostmaster sqlq "UPDATE hosting_site SET status = -2 \
  WHERE vid = (SELECT vid FROM node WHERE title = 'foo.example.com');"

# Drop the orphaned tasks for that node
drush @hostmaster sqlq "DELETE FROM hosting_task \
  WHERE rid = (SELECT nid FROM node WHERE title = 'foo.example.com');"

-2 is the correct value (define('HOSTING_SITE_DELETED', -2)). Confirm the join target first if the same title exists across revisions:

SH
drush @hostmaster sqlq "SELECT nid, vid, status FROM node n \
  JOIN hosting_site h USING (vid) WHERE n.title = 'foo.example.com';"

Then re-run Delete on the now-clean record.

Grants on a site's own database: two stored spellings

A site's grant on its own database is issued with _ and % escaped, so it names one database rather than a LIKE pattern of them. That is why SHOW GRANTS can read back either spelling, and why a by-hand REVOKE has to use the spelling MySQL stored — see Migration source for the wildcard mechanics and the site\_0 read-back. Only plain identifiers ([A-Za-z0-9_]+) take the escaped path; any other name keeps the unescaped form, so both spellings can legitimately exist on one server.

A site created before that escaping still holds the old unescaped grant. Wherever the site's grant is re-issued — site Install, Deploy (the path Restore, Migrate and Clone take) and the change-site-password task — BOA adds the exact grant first and only then drops the old wildcard one, logging at notice level:

TXT
Replaced the wildcard grant on <database> with an exact one for <user>

A plain site Verify does not re-issue the grant, so it does not converge this. If the REVOKE fails, the task logs instead, at warning level:

TXT
Could not drop the wildcard grant on <database> for <user>; the exact grant is in place alongside it

That shape means the exact grant is in place and the wildcard one is still beside it — nothing is broken, but the leftover is worth clearing by hand once SHOW GRANTS confirms it:

SH
mysql -e "SHOW GRANTS FOR '<user>'@'<host>';"
mysql -e "REVOKE ALL PRIVILEGES ON \`<db>\`.* FROM '<user>'@'<host>';"

Delete/Restore warning: object-level grants that could not be revoked

A Delete — or a Restore, Deploy or Migrate, which rotate the site's database credentials and retire the superseded MySQL account — can end Warning with this line in the task log:

TXT
REVOKE/1: sql user <user> holds table, column or routine level grants on <database> that could not be revoked; leaving the account and those grants in place

(REVOKE/2 is the same finding at the site's regular 127.0.0.1/localhost host entries; REVOKE/1 covers every other host row the account has, % included.)

When a site's database is destroyed, BOA revokes every grant the account holds on that one database and then drops the account once nothing beyond the bare global USAGE row remains. That sweep covers both the database-level grants (ONdb.*, in either stored spelling) and the grants one level below them — GRANT ... ON db.table, column lists, ON PROCEDURE/FUNCTION — each re-issued per object exactly as SHOW GRANTS reports it, so nothing outside the destroyed database is ever touched. BOA never creates grants below database level itself, so they are operator- or import-made shapes; clearing them is routine and logs a notice rather than a warning:

TXT
REVOKE/1: revoked table, column or routine level grants held by sql user <user> on <database>

The warning appears only when such a row resists: its stored form does not parse, or the server refuses the revoke. The task then keeps both the account and those grants rather than guess at them — nothing is dropped on partial or unknown state. The hazard it points at: freed database names get reused, so a leftover table-level grant would silently cover a future site's identically named tables.

What to do — review the named account, then clear what should not stay:

SH
mysql -e "SHOW GRANTS FOR '<user>'@'<host>';"
# if those grants should go:
mysql -e "REVOKE ALL PRIVILEGES ON \`<db>\`.\`<table>\` FROM '<user>'@'<host>';"
mysql -e "DROP USER '<user>'@'<host>';"   # only once nothing but USAGE remains

Related, and not a warning: a MySQL account that also holds grants on other databases survives the site's deletion by design — imported sites may share one account across several databases, and the destroy removes only the destroyed database's grants. Such an account disappears on its own once the last database using it goes.

sites/<domain>.restore left after a Clone/Migrate failure

First, the failure class worth ruling out before anything else: a Migrate whose final updatedb step fails is very often a site that entered the move with database updates already pending on the source. The pre-flight is drush @site-alias updbst on the source site — run any pending updates there first, then migrate. When the update run does fail mid-task, the task log names the failing update_N() hook, which is the lead worth chasing before re-running anything.

A failed Clone or Migrate leaves a sites/<domain>.restore directory — the site's pre-task state, kept so the task can be rolled back. The real platform tree lives under /data/disk/<USER>/static/<platform>; the chroot-visible copy is the SFTP symlink under /home/<USER>.ftp/static/<platform>.

A platform Verify skips these revert artefacts (*.restore, *.original) in its automatic site import — no more phantom example.com.restore site records — and records each one it sees, per platform. That record drives the panel path below; it self-heals, so a leftover removed by hand simply disappears from it at the next Verify.

The directory is only half of what the failed task strands — the same abort can leave a <domain>.restore nginx vhost and its matching Drush alias behind as well, and neither path below removes those. That half is covered at the end of this section.

The panel path (works for clients too)

  1. On the platform node, run Import leftover and pick the recorded leftover. It appears as a clearly badged, disabled Leftover artifact pseudo-site owned by the same client as the original site — visible, but serving nothing and offering exactly one task.
  2. On that record, run Purge leftover: the stranded directory is removed, and the copy database it points at is dropped only when it is provably orphaned — the name parsed from the leftover's own settings.php must still match, no site record anywhere may own that database, and no other sites/*/settings.php on the platform may reference it. Anything short of full agreement degrades to a directory-only purge with the reason in the task log.

Clients hold the purge permission (not the platform-level import), so the usual flow is: you expose, they press the button — or you do both.

Source site intact (shell path)

SH
# Remove the stray .restore directory (real platform path)
rm -rf /data/disk/<USER>/static/<platform>/sites/<domain>.restore

Then Verify the site to reconcile Ægir's records. Any orphan copy database (<db>_N) from the failed task is yours to confirm and drop by hand on this path — the panel purge is what automates that check.

Site half-migrated (exists broken on both platforms)

  1. Determine the good copy (compare timestamps and DB row counts).
  2. Disable the broken copy.
  3. Verify the good copy.
  4. Delete the broken copy.

See Cross-host migration for the full migrate mechanics.

The leftover vhost and alias go separately

Removing the directory — by either path above — leaves two more artefacts of the same failed task in place:

  • /data/disk/<USER>/config/server_master/nginx/vhost.d/<domain>.restore
  • /data/disk/<USER>/.drush/<domain>.restore.alias.drushrc.php

Purge leftover does not touch them. The backend command removes the stranded directory and, when its orphan test agrees, the copy database — nothing else. The pseudo-site it retires never had a vhost of its own either, since the purge is the only task that record ever accepts. The rm -rf path leaves the pair just the same.

They come off through a different mechanism: the nightly ghost-vhost check, which moves the vhost and its alias into /data/disk/<USER>/undo/. It acts only once the vhost has looked like a ghost on 2 consecutive nightly runs, and skips any vhost written in the last 24 hours — a restore still in flight.

The catch is the flag. _GHOST_VHOSTS_CLEANUP defaults to NO, which detects and logs only — so on a self-hosted box nothing moves these two files for you. BOA seeds it to YES on omega8.cc-hosted servers alone (hostname ending .aegir.cc). Either set it in /root/.<account>.octopus.cnf or /root/.barracuda.cnf, or move the pair aside by hand once the site is confirmed gone from the platform:

SH
mkdir -p /data/disk/<USER>/undo
mv /data/disk/<USER>/config/server_master/nginx/vhost.d/<domain>.restore \
  /data/disk/<USER>/undo/
mv /data/disk/<USER>/.drush/<domain>.restore.alias.drushrc.php \
  /data/disk/<USER>/undo/

Neither the reaper nor a manual move reloads nginx — the server block stays in the running configuration until nginx is next reloaded. Run nginx -t && service nginx reload if you want it out of the live config now.

Full behaviour, flags and dry-run output: Ghost & empty-artefact cleanup.

CiviCRM backend tasks blocked by the *.drush.inc filter

Symptom. Ægir backend tasks (Verify, Migrate, Clone) on a CiviCRM site start failing after the 5.10.1 upgrade, while the site keeps serving traffic normally under PHP-FPM.

Cause. BOA's *.drush.inc command-file filter — shipped inert in 5.9.5, active from 5.10.1 — denies Drush command files located under tenant-writable site paths after realpath() resolution.

CiviCRM legitimately ships its own command files (civicrm.drush.inc, cv.drush.inc, civicrm_drush.drush.inc) inside the site tree, so the backend Drush run that drives those Ægir tasks refuses to load them. This affects Drush command-file loading during backend tasks only; normal web requests never go through Drush.

Fix — per-Octopus opt-in allowlist

SH
# As root. The Octopus user is the one owning the backend task's $HOME under /data/disk/.
touch /data/conf/<octopus-user>_civicrm.txt

When the file exists, the three CiviCRM command-file basenames are loaded during backend tasks; otherwise they stay blocked (off by default). Re-run the failed task once the file is in place. The filter mechanism is documented in lshell & limited users.

Verify fails after permission drift

Symptom. Verify fails with "Could not write to …" or similar ownership errors after Composer or a manual file operation left files with the wrong owner/mode.

Fix — BOA permission/ownership scripts

These scripts accept only the --flag=PATH equals form. The space-separated form (--root /path) hits the invalid-argument branch and exits 1 before doing any work. Run them as the SFTP user oN.ftp, against the real platform path under /data/disk/<USER> (or, inside the chroot, the /home/<USER>.ftp symlink):

SH
# Whole platform
fix-drupal-platform-permissions.sh --root=/data/disk/<USER>/static/<platform>
fix-drupal-platform-ownership.sh   --root=/data/disk/<USER>/static/<platform>

# Single site (--site-path must contain settings.php or the script exits 1)
fix-drupal-site-permissions.sh --site-path=/data/disk/<USER>/static/<platform>/sites/<domain>
fix-drupal-site-ownership.sh   --site-path=/data/disk/<USER>/static/<platform>/sites/<domain>

The ownership script chowns what exists and says nothing about optional paths (libraries/, core/, misc/ differ by Drupal generation). A tree without a sites/ directory is refused outright for every generation with Error: Please provide a valid Drupal root directory (a Drupal 7 tree without sites/ used to slip through that check and be chowned silently).

Then re-run Verify. See lshell & limited users for the full fix-drupal-* script set and the oN.ftp limited shell.

Every Verify fails with "Access denied for user ''@'localhost'"

Symptom. Every Verify (and any other backend task) on every site fails with Access denied for user ''@'localhost' (using password: NO).

Cause. A botched hostmaster Migrate left the {hosting_db_server} credentials row empty or deleted, so the master_db credential URL rendered as mysql://:@host and every provision-verify failed. Previously the condition was self-perpetuating, with no UI recovery.

Post-5.10.3 Hosting carries three reinforcing guards:

  • Credentials are written via an atomic upsert instead of delete-then-insert, so the row is never momentarily absent.
  • Blank credentials are refused, with the last known-good db_user/db_passwd recovered from the current or, failing that, the latest revision of the DB-server node.
  • A degenerate mysql://:@host alias can no longer blank stored values on context import (the server presave hook carries db_user forward alongside db_passwd).

On current code the condition is designed to self-recover on the next task that saves server credentials (e.g. Verify), as long as any revision of the DB-server node still holds the values — which is the botched-migrate case.

Operator action on current code: none beyond re-running the failed task. If the symptom appears at all, confirm the box is on a post-5.10.3 build before attempting deeper surgery. See Database for credential storage and Cross-host migration for the trigger scenario.

D10/D11: updatedb, cache rebuild, fixed container-poison fatals

Ægir's own Drush 8 must not compile a D10/D11 site's container, and a locked D10/D11 platform runs no site-local Drush at all under PHP 8 (the lock de-types Symfony Console and overlays psr/log v1 for Ægir's Drush 8, and a modern Drush fatals at class load against both). So the backend splits the work: the container and cache rebuild is done by Drupal core itself, in a Drush-free subprocess (REBUILD/CORE/... lines), on any lock state; the deploy pipeline's updatedb runs through the platform's site-local Drush inside a real unlock/re-lock window. Behaviours to know when reading a task log:

REBUILD/CORE/OK is the rebuild. Every site Verify, Flush all caches, Rebuild registry, install and deploy of a D10+ site logs REBUILD/CORE/CMD, a REBUILD/CORE/DONE line naming the Drupal and PHP versions, and REBUILD/CORE/OK. A rebuild that did not happen is a warning: REBUILD/CORE/FAIL followed by the subprocess output (a PHP fatal is reported as REBUILD/CORE/FATAL), or REBUILD/CORE/SKIP when the docroot or the script could not be found. Before this (BOA 5.88.8 up to the 2026-09-08 backend) the rebuild ran drush cr through the site-local Drush inside a chmod-only window; on a locked platform under PHP 8 that command died at class load and the task log showed only the window's own REBUILD/LOCAL/DRUSH/EXEC, CMD and RELOCK lines, so no D10+ Verify ever rebuilt anything. If you see those three lines with nothing between them, the box runs the old backend.

DEPLOY2/LOCAL/DRUSH/UNLOCK ... DEPLOY10/LOCAL/DRUSH/LOCK around updatedb. A clone, migrate, rename or restore of a D10+ site unlocks the platform for real (CORE/PATCH/REVERT, CONSOLE/PATCH/REVERT), prints the site-local Drush version on the DEPLOY7/LOCAL/DRUSH/VERSION line, runs updatedb and the update_fetch_task clear (UPDFETCH/LOCAL/DRUSH/CMD), re-locks (CORE/PATCH/APPLY, CONSOLE/PATCH/APPLY), and only then rebuilds the container by core, so the container updatedb compiled in the stock state never stays the one the patched web serves. updatedb is addressed by --root/--uri (the site bootstraps from its own settings), not by the site's Drush 10+ alias, and the alias store refresh just before the window (DRUSH/TEN/ALIASES/UPDATE) runs on the backend's own interpreter and php.ini. Two warnings name the failures that used to pass as notices: DRUSH/TEN/ALIASES/FAIL (the alias conversion refused or failed, typically "restricted PHP modes ... open_basedir" from a Drush 10+ that was handed the account's Drush 8 ini) and DEPLOY9/LOCAL/DRUSH/FAIL (updatedb did not run to completion; the lines above it say why, e.g. an alias that "could not be found" on an older backend, and the pending updates must be applied by hand).

The same window serves the Run db updates command. provision-update (command line only; the front-end task was retired long ago) applies a Drupal 10+ site's pending updates through exactly this unlock, site-local updatedb, re-lock and core rebuild sequence, with UPDATE2 ... UPDATE10 markers in place of DEPLOY...; its Site-local Drush updatedb task completed line is printed only when updatedb ran to completion, and a window that did not run (no executable site-local Drush) or did not complete fails the task, which restores the pre-update backup the task took first (the same net the Drush 8 updatedb had). Ægir's own Drush 8 no longer runs updatedb on a Drupal 10+ site anywhere (before this it did here, and the Drush 8-compiled container it left was the one the web served until the next Verify). Drupal 6-9 keep Ægir's Drush 8 for both.

DEPLOY/BOOTSTRAP/FAIL ends the deploy WARNING. After the window the deploy finishes in-process under Ægir's Drush 8 (the site's packages map for the front end, the URL-rename engine of a clone or rename, the node access rebuild). When that bootstrap fails the task no longer ends SUCCESSFUL over a plain log line: the warning names what was skipped; run a Verify on the site once it bootstraps. An empty DEPLOY7 version line and no updatedb output after DEPLOY9 is the old backend's chmod-only window: the schema updates did not run on that deploy. A platform found unlocked (a developer's Drush window, a manual Unlock) is left unlocked; the trailing site Verify re-locks it as before.

updatedb skipped on a degraded platform. When a D10+ platform has no executable site-local Drush, provision-deploy skips updatedb with a single actionable warning — updatedb skipped: site-local Drush not executable at <path> and Aegir Drush 8 cannot bootstrap D10+; apply the DB update manually — instead of a doomed Drush 8 attempt that only failed noisily.

Recovery: run the platform's Unlock Local Drush task, then vdrush @alias updb (or repair the site-local Drush; vdrush is the pre-existing BOA wrapper).

On this degraded path the automatic update_fetch_task clear also logs-and-skips, so the manual DELETE FROM key_value WHERE collection = 'update_fetch_task' workaround applies only here — every normal Clone, Migrate, platform Migrate, Rename, and Restore path clears it automatically after updatedb.

The fixed container-poison fatal family (post-5.10.3). Older builds fully unlocked vendor/drush for the cache-rebuild/updatedb windows, reverting the Ægir patches; the D8+ service container then compiled in stock state and baked Drush's DrushLog logger into the persisted (Valkey) container, which the patched web runtime could not resolve — ServiceCircularReferenceException on D11, Class Drush\Log\DrushLog not found on D10 — breaking every site on the platform after routine tasks (disable cron, Clone, Verify, deploy/updatedb) until a manual non-Drush cache clear.

Current code never lets a Drush-compiled container be the last one written: the rebuild that follows every Drush-driven step is core's own, with no Drush loaded, so the container the web serves carries no Drush service; a failed step re-locks before it returns, so a failed task never leaves the shared platform vendor/drush unlocked.

The old fatal presented as a 500-class WSOD, not a 502 — see 502 Bad Gateway for separating the two.

could not bootstrap drupal after updatedb in a DEPLOY log. On D10/D11 this line no longer means stale caches were shipped: the container and cache rebuild is core's own, run after the re-lock (non-fatal by design, never silent). D8/D9 keep the Ægir Drush 8 cache-rebuild via backend invoke.

Fresh D11+ installs fail visibly. A failed site-install now aborts the install task (PROVISION_DRUPAL_INSTALL_FAILED) instead of reporting a broken site as installed. On success, the one-time /user/reset/ login link is derived through the site-local Drush, and automatic_updates / package_manager are uninstalled right after install (Ægir never lets web UI tools overwrite the codebase).

The install task log also says why a D11+ site-install died. Before the subprocess starts, D11/INSTALL/MEM prints the account's CLI memory_limit and the limit the child actually runs with: the child is floored at a quarter of the installed RAM, capped at 1024 MB (a Varbase install needs about 600 MB; a 4 GB box with the default 2 GB reserve derives only 466 MB), and ~/static/control/install-memory.info holding a number of MB pins it instead. On a non-zero exit, D11/INSTALL/RC logs the exit status with a reading (255 = PHP fatal, 137 = SIGKILL, "terminated abnormally" = a fatal inside modern Drush), and D11/INSTALL/PHPLOG logs the tail of the child's own PHP error log, kept at ~/log/php-site-install-<uri>.log. The child logs PHP errors to that file and a prepended shutdown hook records the last fatal there and to stderr even when the runtime turned log_errors off, so an install that stops after its last [info] ... module installed line now names the fatal (typically Allowed memory size ... exhausted) instead of failing silently.

Ægir frontend not loading after an upgrade completed

If barracuda up-* reports a finished run but the Ægir control panel is dead afterwards, the first suspect is an aborted hostmaster-migrate — the frontend-upgrade step inside the barracuda pass. A failed migrate is one of the failure classes that deliberately does not raise the status-AegirUpgrade-FAIL outcome marker (the script still stamps status-AegirUpgrade-OK), so the run can end looking clean while the frontend never finished migrating. The classic trigger is Drush briefly unable to talk to Percona while it restarts mid-upgrade (the pass guards this with _check_sql_running, but an abort mid-migrate still leaves the frontend half-moved).

What to do: check the barracuda run log for the migrate step, and if the failure reproduces, re-run with verbose debug mode enabled — a debug run executes the migrate with -d and without output suppression, which is usually the difference between a mystery and a fixable error. The migrate internals (state bus, abort markers, zombie sweeps) are on Hostmaster upgrade in the maintainer area.

Task log won't load in the UI

A task log too large for Drupal to render is still on disk. The most recent .log files under the instance log dir are the task logs:

SH
ls -lt /data/disk/<USER>/log/ | head
less /data/disk/<USER>/log/<task-id>.log

Reading "Drush command terminated abnormally" in task logs

Since the Drush 8.5.3 bundle shipped with BOA 5.10.3, drush_shutdown() reports a cause for an abnormal exit only when the last recorded PHP error is fatal-class (E_ERROR | E_PARSE | E_CORE_ERROR | E_COMPILE_ERROR | E_USER_ERROR | E_RECOVERABLE_ERROR).

Previously any lingering recorded error — typically a PHP 8.x implicitly-nullable deprecation emitted at autoload by the old vendored psy/psysh and correctly suppressed by error_reporting — was printed as Drush command terminated abnormally due to an unrecoverable error: ... Implicitly marking parameter $config as nullable, mislabeling a harmless deprecation and hiding the real cause.

The common trigger was exiting the psysh REPL (drush php / core-cli), which calls exit() before drush_main() completes.

Operator takeaway: on current code the cause line is trustworthy — if "terminated abnormally due to an unrecoverable error" names an error, it is a genuine fatal. In older task logs, an ...unrecoverable error: ... deprecated/nullable... line was noise; look for the real failure elsewhere in the log.

© 2026 BOA Documentation. All rights reserved.