fix(onboarding): identity read that survives a rewrite, disk refused before the pull — 0.1.8

Three faults from one agent-managed onboarding (env-dev-08, 2026-09-09), each of which pointed
the diagnosis away from the actual fault.

1. install.sh granted the agent's read on the ziti identity with a POSIX ACL. ziti-edge-tunnel
   rewrites that file on a controller config update and the rewrite drops the ACL: the agent
   applied cleanly at 01:21 and was failing every tick by 01:32. Group membership survives the
   rewrite (the file stays ziti:ziti 0640), so install.sh and the package postinstall now add
   monky-deployd to the `ziti` group, and a default ACL on the identity directory carries the
   grant onto a newly created file. The explicit ACLs stay for the boxes that need them.

2. openziti.load() accepts an unreadable or malformed identity: the C SDK logs "configuration is
   invalid" and returns a context that only fails at dial, as a bare TypeError, which the
   transport reported as a missing intercept or a policy gap. The SDK transport now reads and
   parses the identity itself and names the real fault first.

3. The disk pre-flight ran only when the bundle declared disk_need_bytes, so a bundle without one
   died mid-pull with containerd's "no space left on device" — which reads as a registry fault.
   A bundle that declares no size now has to clear the headroom floor, and the pre-flight measures
   containerd's root as well as the docker data-root: docker 29 keeps image layers in the
   containerd image store, and on env-dev-08 those sat on different filesystems (93 GiB free where
   the agent looked, 2.8 GiB where the pull wrote).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLB7jieMNRkTsJ2epr4Ds1
This commit is contained in:
2026-09-09 01:48:49 +00:00
parent cc9dcebd7e
commit 43615a6fda
14 changed files with 261 additions and 15 deletions
+27 -2
View File
@@ -12,6 +12,10 @@ import time
from dataclasses import dataclass
from pathlib import Path
# containerd's default root: docker 29's image store lives here, often on another filesystem
# than DockerRootDir. Both are checked before a pull (see Docker.storage_paths).
CONTAINERD_ROOTS = ("/var/lib/containerd",)
log = logging.getLogger("monky-deployd.compose")
@@ -100,8 +104,23 @@ class Docker:
root = ""
return root or "/var/lib/docker"
def free_bytes(self, path: str | None = None) -> int | None:
p = path or self.data_root()
def storage_paths(self) -> list[str]:
"""Every filesystem a `compose pull` can fill.
docker 29 keeps IMAGE layers in the containerd image store (containerd's own root,
/var/lib/containerd by default), NOT under DockerRootDir. On a box where those two sit
on different filesystems, measuring only the data-root reports plenty of room while the
pull dies with "no space left on device" (env-dev-08, 2026-09-09: 93 GiB free on the
data-root, 2.8 GiB on the root filesystem that held containerd).
"""
paths = [self.data_root()]
for extra in CONTAINERD_ROOTS:
if os.path.isdir(extra):
paths.append(extra)
return paths
def _free_at(self, path: str) -> int | None:
p = path
while p and not os.path.exists(p):
p = os.path.dirname(p)
try:
@@ -110,6 +129,12 @@ class Docker:
return None
return st.f_bavail * st.f_frsize
def free_bytes(self, path: str | None = None) -> int | None:
"""Free bytes on `path`, or the TIGHTEST of the image-storage filesystems."""
paths = [path] if path else self.storage_paths()
seen = [v for v in (self._free_at(p) for p in paths) if v is not None]
return min(seen) if seen else None
def image_prune(self) -> None:
try:
self.run(["image", "prune", "-f"], timeout=300)