Make backup_previous_artifact (src/apply.rs) link-based. It is a cp -fp today, and an interrupted copy leaves a truncated .previous that a later revert succeeds onto — a broken binary installed by a rollback that reported success.
This is the code half of a decision recorded in aicers/roxyd RFC 0002 §4/§8 and aicers/bootler RFC 0004 §8. Both consumers keep their call sites unchanged; both gain the guarantee, and bootler backups stop being truncatable too. roxyd self-update already required link semantics and described this primitive as too weak to reuse — after this change the two describe one operation the same way.
What changes
backup_previous_artifact (src/apply.rs) is a cp -fp today. It becomes
link-based, with this exact sequence:
link(2) the destination to a temporary sibling name in the same directory
rename(2) that temporary over <dest>.previous
fsync the containing directory
Why each step is what it is
- A link rather than a copy. An interrupted
cp leaves a truncated
.previous, so a later revert succeeds onto a broken binary. An interrupted
link leaves no .previous, so the revert fails where anyone looking can
see it. The link also subsumes cp -p's mode and timestamp preservation:
sharing an inode makes them identical rather than copied.
- Not
ln -f. link(2) fails with EEXIST, so ln -f unlinks the
existing .previous and then links — a window in which no backup exists at
all. Link-to-temp then rename is atomic and never exposes an absent backup.
This mirrors what put_file already does for the artifact itself, so it is
the crate's existing idiom rather than a new one.
- Not a bare
link. A directory entry is not durable until its directory
is fsynced, and the backup is precisely the thing that must survive a power
loss.
This is an executor/API-level change, not a substitution of one shell verb
for another: the Executor gains whatever primitive it lacks to express the
sequence.
That primitive is Executor::hard_link_over, and every transport runs the same sequence through it: InDaemonExecutor as direct syscalls, being root already, and the shell transports as a single elevated sh -c script, so nothing can be interleaved between the link and the rename by a second elevation. Both report every on-host failure the same way, and backup_previous_artifact folds that into the subject-labelled CoreError::Command the cp raised, so neither consumer's call site nor its rendering changes.
Constraints that already hold
- Same filesystem.
.previous is the artifact's sibling by construction.
- A regular file, not a directory. The subject is one path placed by
place_file → put_file; bootler refuses a ComposeBundle before reaching
it and container images go through docker_load_image.
- In-place writes would destroy a linked backup, since the link shares the
inode. They do not happen: put_file writes a temporary and renames over the
destination, so the swap replaces the directory entry and the old inode
survives behind the link.
- The destination always exists. The function returns early unless
test -e succeeds, so the link is never asked to stand in for a missing file.
What it refuses
The applier is specified for a regular file. A symlink at the
destination is refused, not linked: cp -fp follows it and writes a regular
file, while a link without -L would capture the symlink itself and leave a
.previous pointing wherever the operator pointed it. A directory or any
other non-regular file is refused too.
Those guards are about the artifact. The backup's own name is the crate's rather than the caller's, so it is not guarded the same way: the publish is a rename, which replaces the entry it is given without resolving it. A symlink planted at .previous is therefore displaced rather than written through, and whatever it pointed at keeps its inode and its bytes — where the copy this replaces would have opened it and landed the artifact on the far end while leaving no backup at all. A directory at .previous is a failure instead: rename(2) cannot move a file into one, and the shell transports refuse it before the link, since mv would take the temporary inside it and exit 0, reporting a backup taken under a path nobody named.
What it does not fix
backup_previous_artifact is still not idempotent. A resumed apply that
runs it a second time backs up the half-applied bytes and destroys the rollback
point — a property of when the backup runs, not of copy versus link. The
caller's apply journal remains what guards it, and the backup step stays
conditional on that journal's backup-taken record.
Acceptance criteria
Two are about what the applier accepts: a symlink destination is refused,
and so is a directory or other non-regular file.
The rest are about crashes, and they split twice — on whether a .previous
already existed, and on which fault model is being asserted:
- Process interruption — the applier is killed between steps; the
filesystem's state is whatever the last completed call left.
- Power loss — a completed call whose directory entry has not been fsynced
may or may not survive, so nothing above the fsync line is asserted portably.
Replacing an existing .previous is a process-interruption guarantee: a
fault injected between the link and the rename, and again between the
rename and the directory fsync, leaves the old .previous or the
new one, never a partial or absent file. It holds because the old entry is
never unlinked and rename is atomic. It is not claimed against power loss.
Creating the first .previous has nothing to preserve, so absence is the
correct intermediate state and the criterion is recovery rather than the
file, in either fault model: a fault anywhere in the sequence leaves the
caller's journal without a backup-taken record, so the resumed apply re-takes
it. That is sound because backup_previous_artifact runs before place_file,
so what gets linked is still the live, un-replaced binary.
Under power loss both cases fall back to the same answer: the journal
decides. No portable durability claim is available between the rename and
the directory fsync, so none is made.
Consumers
Both, unchanged at the call site: bootler (core/src/update.rs) and roxyd's
generic apply path. Every consumer gains the guarantee, and bootler's backups
stop being truncatable too. roxyd's self-update already required link
semantics and described backup_previous_artifact as too weak to reuse; with
this change the two describe one operation the same way.
Make
backup_previous_artifact(src/apply.rs) link-based. It is acp -fptoday, and an interrupted copy leaves a truncated.previousthat a later revert succeeds onto — a broken binary installed by a rollback that reported success.This is the code half of a decision recorded in
aicers/roxydRFC 0002 §4/§8 andaicers/bootlerRFC 0004 §8. Both consumers keep their call sites unchanged; both gain the guarantee, and bootler backups stop being truncatable too. roxyd self-update already requiredlinksemantics and described this primitive as too weak to reuse — after this change the two describe one operation the same way.What changes
backup_previous_artifact(src/apply.rs) is acp -fptoday. It becomeslink-based, with this exact sequence:
link(2)the destination to a temporary sibling name in the same directoryrename(2)that temporary over<dest>.previousfsyncthe containing directoryWhy each step is what it is
cpleaves a truncated.previous, so a later revert succeeds onto a broken binary. An interruptedlink leaves no
.previous, so the revert fails where anyone looking cansee it. The link also subsumes
cp -p's mode and timestamp preservation:sharing an inode makes them identical rather than copied.
ln -f.link(2)fails withEEXIST, soln -funlinks theexisting
.previousand then links — a window in which no backup exists atall. Link-to-temp then
renameis atomic and never exposes an absent backup.This mirrors what
put_filealready does for the artifact itself, so it isthe crate's existing idiom rather than a new one.
link. A directory entry is not durable until its directoryis fsynced, and the backup is precisely the thing that must survive a power
loss.
This is an executor/API-level change, not a substitution of one shell verb
for another: the
Executorgains whatever primitive it lacks to express thesequence.
That primitive is
Executor::hard_link_over, and every transport runs the same sequence through it:InDaemonExecutoras direct syscalls, being root already, and the shell transports as a single elevatedsh -cscript, so nothing can be interleaved between the link and the rename by a second elevation. Both report every on-host failure the same way, andbackup_previous_artifactfolds that into the subject-labelledCoreError::Commandthecpraised, so neither consumer's call site nor its rendering changes.Constraints that already hold
.previousis the artifact's sibling by construction.place_file→put_file; bootler refuses aComposeBundlebefore reachingit and container images go through
docker_load_image.inode. They do not happen:
put_filewrites a temporary and renames over thedestination, so the swap replaces the directory entry and the old inode
survives behind the link.
test -esucceeds, so the link is never asked to stand in for a missing file.What it refuses
The applier is specified for a regular file. A symlink at the
destination is refused, not linked:
cp -fpfollows it and writes a regularfile, while a link without
-Lwould capture the symlink itself and leave a.previouspointing wherever the operator pointed it. A directory or anyother non-regular file is refused too.
Those guards are about the artifact. The backup's own name is the crate's rather than the caller's, so it is not guarded the same way: the publish is a
rename, which replaces the entry it is given without resolving it. A symlink planted at.previousis therefore displaced rather than written through, and whatever it pointed at keeps its inode and its bytes — where the copy this replaces would have opened it and landed the artifact on the far end while leaving no backup at all. A directory at.previousis a failure instead:rename(2)cannot move a file into one, and the shell transports refuse it before the link, sincemvwould take the temporary inside it and exit0, reporting a backup taken under a path nobody named.What it does not fix
backup_previous_artifactis still not idempotent. A resumed apply thatruns it a second time backs up the half-applied bytes and destroys the rollback
point — a property of when the backup runs, not of copy versus link. The
caller's apply journal remains what guards it, and the backup step stays
conditional on that journal's backup-taken record.
Acceptance criteria
Two are about what the applier accepts: a symlink destination is refused,
and so is a directory or other non-regular file.
The rest are about crashes, and they split twice — on whether a
.previousalready existed, and on which fault model is being asserted:
filesystem's state is whatever the last completed call left.
may or may not survive, so nothing above the fsync line is asserted portably.
Replacing an existing
.previousis a process-interruption guarantee: afault injected between the
linkand therename, and again between therenameand the directoryfsync, leaves the old.previousor thenew one, never a partial or absent file. It holds because the old entry is
never unlinked and
renameis atomic. It is not claimed against power loss.Creating the first
.previoushas nothing to preserve, so absence is thecorrect intermediate state and the criterion is recovery rather than the
file, in either fault model: a fault anywhere in the sequence leaves the
caller's journal without a backup-taken record, so the resumed apply re-takes
it. That is sound because
backup_previous_artifactruns beforeplace_file,so what gets linked is still the live, un-replaced binary.
Under power loss both cases fall back to the same answer: the journal
decides. No portable durability claim is available between the
renameandthe directory
fsync, so none is made.Consumers
Both, unchanged at the call site: bootler (
core/src/update.rs) and roxyd'sgeneric apply path. Every consumer gains the guarantee, and bootler's backups
stop being truncatable too. roxyd's self-update already required
linksemantics and described
backup_previous_artifactas too weak to reuse; withthis change the two describe one operation the same way.