dpkg calls the OLD package's prerm on an UPGRADE as well as on a removal (rpm passes a
remaining-instance count), and preremove.sh ran `systemctl disable --now monky-deployd.timer`
unconditionally. Upgrading env-dev-01 and env-dev-08 to 0.1.8 today stopped and disabled both
agents.
The failure is silent, which is the dangerous part: the box stays reachable, the containers keep
running, and nothing reports that check-ins have ceased — the backend just stops converging. A
fleet upgrade would have taken every agent offline at once and looked like a success.
preremove.sh now returns early for every upgrade shape (upgrade, failed-upgrade, deconfigure,
rpm's 1) and only disables on a real removal. postinstall.sh try-restarts the long-lived proxy
unit so it picks up the new code; the timer needs nothing, since each tick is a fresh process.
Tests drive the script with a fake systemctl on PATH and assert an upgrade touches no units.
OPERATIONS.md warns that a box coming FROM 0.1.8 or earlier still needs its timer re-enabled by
hand, because the old prerm has already run by then.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLB7jieMNRkTsJ2epr4Ds1
Three faults from one agent-managed onboarding (env-dev-08, 2026-09-09), each of which pointed
the diagnosis away from the actual fault.
1. install.sh granted the agent's read on the ziti identity with a POSIX ACL. ziti-edge-tunnel
rewrites that file on a controller config update and the rewrite drops the ACL: the agent
applied cleanly at 01:21 and was failing every tick by 01:32. Group membership survives the
rewrite (the file stays ziti:ziti 0640), so install.sh and the package postinstall now add
monky-deployd to the `ziti` group, and a default ACL on the identity directory carries the
grant onto a newly created file. The explicit ACLs stay for the boxes that need them.
2. openziti.load() accepts an unreadable or malformed identity: the C SDK logs "configuration is
invalid" and returns a context that only fails at dial, as a bare TypeError, which the
transport reported as a missing intercept or a policy gap. The SDK transport now reads and
parses the identity itself and names the real fault first.
3. The disk pre-flight ran only when the bundle declared disk_need_bytes, so a bundle without one
died mid-pull with containerd's "no space left on device" — which reads as a registry fault.
A bundle that declares no size now has to clear the headroom floor, and the pre-flight measures
containerd's root as well as the docker data-root: docker 29 keeps image layers in the
containerd image store, and on env-dev-08 those sat on different filesystems (93 GiB free where
the agent looked, 2.8 GiB where the pull wrote).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLB7jieMNRkTsJ2epr4Ds1