mirror of
https://scm.tikali.ai/tikali/applications/monky/monky-deployd.git
synced 2026-09-18 04:36:15 +00:00
fix(packaging): an upgrade must not stop and disable the agent — 0.1.9
dpkg calls the OLD package's prerm on an UPGRADE as well as on a removal (rpm passes a remaining-instance count), and preremove.sh ran `systemctl disable --now monky-deployd.timer` unconditionally. Upgrading env-dev-01 and env-dev-08 to 0.1.8 today stopped and disabled both agents. The failure is silent, which is the dangerous part: the box stays reachable, the containers keep running, and nothing reports that check-ins have ceased — the backend just stops converging. A fleet upgrade would have taken every agent offline at once and looked like a success. preremove.sh now returns early for every upgrade shape (upgrade, failed-upgrade, deconfigure, rpm's 1) and only disables on a real removal. postinstall.sh try-restarts the long-lived proxy unit so it picks up the new code; the timer needs nothing, since each tick is a fresh process. Tests drive the script with a fake systemctl on PATH and assert an upgrade touches no units. OPERATIONS.md warns that a box coming FROM 0.1.8 or earlier still needs its timer re-enabled by hand, because the old prerm has already run by then. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KLB7jieMNRkTsJ2epr4Ds1
This commit is contained in:
@@ -1,6 +1,18 @@
|
||||
<!-- xlate:verbatim-fences -->
|
||||
# Changelog
|
||||
|
||||
## 0.1.9 — an upgrade no longer stops the agent (2026-09-09)
|
||||
|
||||
- **`dpkg -i` over a running agent disabled it.** dpkg calls the OLD package's `prerm` on an
|
||||
**upgrade** as well as on a removal (rpm passes a remaining-instance count), and `preremove.sh`
|
||||
ran `systemctl disable --now monky-deployd.timer` unconditionally. Upgrading env-dev-01 and
|
||||
env-dev-08 from 0.1.6/0.1.7 to 0.1.8 stopped and **disabled** both agents. It is silent: the box
|
||||
stays up, the containers keep running, and nothing reports that check-ins have ceased — the
|
||||
backend simply stops converging. `preremove.sh` now returns early for every upgrade shape
|
||||
(`upgrade`, `failed-upgrade`, `deconfigure`, rpm's `1`), and `postinstall.sh` `try-restart`s the
|
||||
long-lived proxy unit so it picks up the new code. A fleet upgrade would have taken every agent
|
||||
offline at once.
|
||||
|
||||
## 0.1.8 — onboarding: keep the identity readable, refuse a full disk (2026-09-09)
|
||||
|
||||
Three faults from one onboarding (env-dev-08, agent-managed, 2026-09-09), each of which sent the
|
||||
|
||||
Reference in New Issue
Block a user