The #105 T3 change (PR #156) tried to keep the resident metadata cache
monotonic by guarding peer-reload publication on lastUpdate(). But
lastUpdate() is the max of per-config timestamps and cannot order whole
records: a node caching {policy@20, CORS@10} that receives a newer CORS@15
still has lastUpdate()==20, so the guard rejects the legitimately-newer
record and the periodic refresh (same comparator) cannot repair it. A
paused reload could also resurrect deleted resident state.
Per the maintainer decision, revert the reload publication to its original
unconditional (acceptable-until-refresh) behavior:
- remove setReloaded and restore the plain Set plus notification/target
registry updates in LoadBucketMetadataHandler;
- restore refreshBucketsMetadataLoop's own lastUpdate() staleness check and
globalEventNotifier.set / globalBucketTargetSys.set publication;
- restore the unconditional GetConfig cache-miss publication;
- document the known freshness limitation at the reload site (the periodic
refresh is best-effort and cannot repair an equal-maximum-timestamp
divergence).
The T1 lifecycle merge-under-lock (UpdateExpiryLCConfig) and both T2 fixes
(DeleteBucket takes metadata.lock before deleting; saveMetadata and
loadBucketMetadataParseUnderLock recheck physical bucket existence) are
kept fully intact.
Tests:
- drop the T3 reproductions (overlapping-reload resident-cache test and the
peer-reload-preserves-current-targets publication test);
- add lockBucketMetadataAcquireHook, a nil-in-production atomic test hook in
the shared metadata.lock path, so tests can deterministically observe a
caller (notably DeleteBucket, whose lock is taken through its
erasureServerPools receiver and is invisible to an injected object layer)
reaching the lock;
- rewrite the T2 delete-race ghost test to hold metadata.lock MID-SAVE (past
saveMetadata's existence recheck) and synchronize on the delete's actual
lock attempt via the hook, so it isolates the lock-before-delete fix:
removing only DeleteBucket's metadata.lock (recheck kept) now fails it;
- rewrite the cancellation test to observe the delete's actual lock attempt,
then cancel and await its error while still holding the lock, so a
scheduling-delayed delete stopped by the canceled context can no longer
pass on a broken tree.
Refs #105. Follows #156.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>