Compare commits

...

28 Commits

Author SHA1 Message Date
Feng Ruohang 4620be394b test: synchronize conditional PUT disk fixtures
Replace unsynchronized getDisks swaps with backing disk-list updates under
erasureDisksMu, matching the existing GetDisks reader lock. Apply the same
helper to capacity and read-fault adapters while preserving nested restore
ordering.

Add a regression that overlaps fixture changes with the real IAM Walk
reader, and run the conditional PUT suite under the race detector in CI.
The regression reproduces the old fixture race; ten fixed race iterations
pass without warnings. Production conditional PUT behavior is unchanged.

Refs #199

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 10:53:24 +08:00
Feng Ruohang 5e7d603083 fix(pools): evaluate cross-pool PUT conditions against the current object
Multi-pool PUT selected a destination by capacity and evaluated
If-Match/If-None-Match only against that destination's local object
state. An empty or stale destination could accept a stale ETag or
If-None-Match:* while another pool held the current object, replacing
it; a current ETag could instead be rejected with 412 or 404.

Under PUT's existing pools-layer object lock, resolve the comparison
object with objectPoolInfos (including draining pools), treat a latest
delete marker as absence, fail closed on unreadable pool metadata, and
clear an accepted callback before destination dispatch. Replica and
data-movement callbacks keep their addressed-version semantics and
metadata reconciliation.

Reproduced on 40220bd836 and RELEASE.2026-09-03T13-18-01Z with six
signed HTTP scenarios: four defect cases failed, two controls passed.

Refs #199

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:21:21 +08:00
Feng Ruohang 40220bd836 Merge pull request #196 from pgsty/codex/merge-r4-r8
Complete R5, R6 and R8 on the main baseline containing R4 and R7.

Preserve the individual signed repairs, independent Opus 5 Max review, and integration validation evidence.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 01:14:06 +08:00
Feng Ruohang df0dfa0a34 docs: record R4-R8 integration review and validation
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:58:31 +08:00
Feng Ruohang 80684fed59 chore: align integrated repair tests with contribution checks
Use the actual author and AGPL-3.0-or-later notices for new R6/R8 tests, preserving all test bodies and the Linux build tag. Record the existing replication ARN prefix used by the new R6 fixture in the compatibility inventory; no wire behavior changes.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:44:34 +08:00
Feng Ruohang 055030ea53 fix(http): honor configured request header deadlines
Signed-off-by: Feng Ruohang <rh@vonng.com>
(cherry picked from commit 0d48d32d7e038ae1ea5966f3d7e0cb86780a6311)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:56 +08:00
Feng Ruohang aea3882c95 docs: record R6 integration verification
(cherry picked from commit d38edb2c46182d3a8fa96e040604493d20a4b478)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:56 +08:00
Feng Ruohang 0c61128d23 fix(replication): retry marker purges through persisted MRF
(cherry picked from commit cf381a7151ef25fc95ace5fedcd767fa19410de2)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:55 +08:00
Feng Ruohang 680eac66e4 fix(replication): preserve ordered tag deletions
Persist, send and reconcile empty tag states together with their revision
across COPY, PUT and multipart replication. Advance local tag mutations
under the existing locks and preserve current tags during replication ACK.

Cover signed HTTP, persistent single/multiple pool state, KMS, SSE-C key
rotation, ordering, retry and duplicate requests. Record real Opus plan
consensus, implementation review and local verification evidence.

Signed-off-by: Feng Ruohang <rh@vonng.com>
(cherry picked from commit 115fe8b12329d147adbaf817faa1737392ecbf9b)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:55 +08:00
Feng Ruohang 9f3037e941 Merge pull request #194 from pgsty/codex/r7-replication-content-encoding
fix(replication): preserve normalized replica metadata
2026-09-16 00:31:03 +08:00
Feng Ruohang 5f00f6762f Merge main after R4 validation into R7 candidate
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:21:21 +08:00
Feng Ruohang af2b1794d3 Merge pull request #193 from pgsty/codex/r4-kms-tag-timestamp
fix(replication): preserve SSE-KMS tag timestamps
2026-09-16 00:20:10 +08:00
Feng Ruohang d371f77dcc docs: record final Opus 5 Max R7 implementation review
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:10:31 +08:00
Feng Ruohang 022722a7a7 chore: align R4 contribution notices and record final review
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:08:09 +08:00
Feng Ruohang 03027727d1 fix(replication): preserve SSE-KMS tag timestamps
Carry the already-parsed source tagging timestamp through the KMS options
constructor so replica COPY can apply newer tag updates on explicitly or
automatically encrypted destinations.

Cover all option encryption modes and signed COPY persistence for newer,
stale, duplicate and timestamp-less updates, including bucket defaults.
Preserve the real Opus 5.0/max plan review, consensus and local validation.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:03:52 +08:00
Feng Ruohang 4fcdf37ce6 fix(replication): preserve normalized replica metadata
Restore only the six replication-specific metadata fields after trust
validation, so streaming uploads retain their actual content encoding and
Snowball entries do not inherit ordinary metadata from the outer archive.

Include helper, authenticated PUT/COPY/multipart and Snowball regressions,
plus the R7 investigation, actual Opus 5 consensus and local verification.
The production change is based on PR #187 by Mikhail Khadarenka.

Co-authored-by: Mikhail Khadarenka <chodorenko@gmail.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:00:17 +08:00
Feng Ruohang 9ebe81c1b3 Merge pull request #192 from pgsty/codex/iam-revision-tombstones
fix(iam): retain revocation versions through replay and recovery
2026-09-15 23:14:22 +08:00
Feng Ruohang e5f5c9e7f6 Merge pull request #191 from pgsty/codex/iam-peer-delete-reload
fix(iam): reload committed state on peer deletion notifications
2026-09-15 23:10:38 +08:00
Feng Ruohang 7b4cacc392 chore: align IAM file notices with contribution policy
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang a0dd7dae9b fix(iam): resume site healing after leadership changes
Keep one healing loop per process across replication configuration reloads. Reacquire leadership after a lease is canceled and allow shutdown while waiting, so temporary quorum loss cannot permanently stop revocation propagation.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang 709d50a916 fix(iam): persist revocations across site replay and recovery
Retain source-ordered tombstones and parent grant boundaries across both IAM backends, cache reloads, and deliberate identity recreation. Reconcile deletions through a versioned, bounded replication protocol with restart-aware acknowledgements.

Cover inherited group grants, STS retention, same-key service recreation, absolute expiration, and failures after the durable commit. Document coordinated upgrades and the remaining consistency boundaries.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang dff81f293b chore: align IAM test notice with contribution policy
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:58:52 +08:00
Feng Ruohang f653a6ea03 fix(iam): reload committed state on peer deletion notifications
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:58:52 +08:00
Feng Ruohang 47d239f84f Merge pull request #190 from pgsty/codex/fix-pool-multipart-preconditions
fix(storage): evaluate multipart preconditions across pools
2026-09-15 21:57:26 +08:00
Feng Ruohang e069fe9d92 fix(storage): evaluate multipart preconditions across pools
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 21:13:37 +08:00
Feng Ruohang d848fb52b5 Merge pull request #189 from pgsty/codex/fix-pool-tag-reconciliation
fix(storage): preserve tags during pool reconciliation
2026-09-15 19:38:19 +08:00
Feng Ruohang 3ce8319251 fix(storage): preserve tags during pool reconciliation
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 19:26:58 +08:00
Feng Ruohang 9df0f4abaf Merge pull request #188 from pgsty/codex/remove-access-tiering-final
Remove access-frequency pool tiering and preserve independent multi-pool correctness fixes. Reconcile ordinary addressed-version DELETE across pools, retaining existing quorum and compatibility boundaries.

Verified delivery head: 4d0693cb8c. All 11 final CI checks and three qualified Linux upgrade runs passed. Introduction, retirement decisions and historical unresolved observations are documented.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 15:09:48 +08:00
234 changed files with 20340 additions and 562 deletions
+3
View File
@@ -90,6 +90,9 @@ jobs:
- name: Run S3 Select tests under race detector
run: go test -race ./internal/s3select/... -count=1
- name: Run conditional PUT tests under race detector
run: go test -race ./cmd -run '^Test(PoolsConditionalPut|SinglePoolConditionalPutHTTP)' -count=1 -timeout=5m
crosscompile:
name: Cross Compile
runs-on: ubuntu-latest
+14 -3
View File
@@ -1,11 +1,11 @@
# Changelog
## Unreleased — main as of 2026-09-13
## Unreleased
The coordinated source is merged through `5d955b5b7444f8a3ab550ce92713607998f89c0d`.
The entries below describe source changes on main since the latest published Server.
**The latest published Server remains 20260903.** These changes are not in its
binaries, packages or images. See the [component matrix](https://silo.pgsty.com/compatibility/versions/)
and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-09-03T13-18-01Z...5d955b5b7444f8a3ab550ce92713607998f89c0d).
and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-09-03T13-18-01Z...main).
### Authorization and security
@@ -22,6 +22,17 @@ and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-0
### Object storage and replication
- Evaluate conditional multipart completion against the logical current object
across all pools while holding the existing object lock. A stale `If-Match`
can no longer replace newer data in another pool, and the current ETag is no
longer rejected because the upload resides next to an older copy. Conditions
are evaluated once; a current delete marker counts as an absent object.
**Availability change:** if metadata cannot be read from any pool, conditional
completion fails even when another pool can still serve GET/HEAD. This also
applies when the unreadable pool may not hold the object: absence cannot be
verified. Retry after the pool recovers. Unconditional completion and the
single-pool path retain their existing behavior.
- Reconcile ordinary single-object version DELETE across all pools, including
null versions, delete markers and unqualified directory-marker DELETE. This
applies the deletion to every resolved pool copy under existing quorum
@@ -829,6 +829,7 @@
"/site-replication/peer/bucket-ops",
"/site-replication/peer/edit",
"/site-replication/peer/iam-item",
"/site-replication/peer/iam-revisions",
"/site-replication/peer/idp-settings",
"/site-replication/peer/join",
"/site-replication/peer/remove",
@@ -877,6 +878,7 @@
"/v2/metrics/cluster",
"/v2/metrics/node",
"/v2/metrics/resource",
"/v3/site-replication/peer/iam-revisions",
"/var/vcap/bosh",
"/verifybinary",
"/version",
@@ -932,6 +934,7 @@
"arn:minio:kms:::this-is-disregarded",
"arn:minio:kms:::xyz-test-key",
"arn:minio:replication:",
"arn:minio:replication::",
"arn:minio:replication::8320b6d18f9032b4700f1f03b50d8d1853de8f22cab86931ee794e12f190852c:destinationbucket",
"arn:minio:replication:::",
"arn:minio:replication:::dest-bucket",
+1
View File
@@ -388,6 +388,7 @@ func registerAdminRouter(router *mux.Router, enableConfigOps bool) {
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/join").HandlerFunc(adminMiddleware(adminAPI.SRPeerJoin))
adminRouter.Methods(http.MethodPut).Path(adminVersion+"/site-replication/peer/bucket-ops").HandlerFunc(adminMiddleware(adminAPI.SRPeerBucketOps)).Queries("bucket", "{bucket:.*}").Queries("operation", "{operation:.*}")
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/iam-item").HandlerFunc(adminMiddleware(adminAPI.SRPeerReplicateIAMItem))
adminRouter.Methods(http.MethodGet, http.MethodPut).Path(adminVersion + "/site-replication/peer/iam-revisions").HandlerFunc(adminMiddleware(adminAPI.SRPeerIAMRevisions))
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/bucket-meta").HandlerFunc(adminMiddleware(adminAPI.SRPeerReplicateBucketItem))
adminRouter.Methods(http.MethodGet).Path(adminVersion + "/site-replication/peer/idp-settings").HandlerFunc(adminMiddleware(adminAPI.SRPeerGetIDPSettings))
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/edit").HandlerFunc(adminMiddleware(adminAPI.SiteReplicationEdit))
+3
View File
@@ -418,6 +418,9 @@ func getReplicationState(rinfos replicatedInfos, prevState ReplicationState, vID
for _, rinfo := range rinfos.Targets {
if rinfo.ResyncTimestamp != "" {
if rs.ResetStatusesMap == nil {
rs.ResetStatusesMap = make(map[string]string)
}
rs.ResetStatusesMap[targetResetHeader(rinfo.Arn)] = rinfo.ResyncTimestamp
}
}
+111 -63
View File
@@ -425,6 +425,7 @@ func checkReplicateDelete(ctx context.Context, bucket string, dobj ObjectToDelet
// the mere presence or absence of the target version.
func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, objectAPI ObjectLayer) replicatedInfos {
var replicationStatus replication.StatusType
isPurge := dobj.isVersionPurge()
bucket := dobj.Bucket
versionID := dobj.DeleteMarkerVersionID
if versionID == "" {
@@ -484,6 +485,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
lk := objectAPI.NewNSLock(bucket, "/[replicate]/"+dobj.ObjectName)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
dobj.RetryCount++
globalReplicationPool.Get().queueMRFSave(dobj.ToMRFEntry())
sendEvent(eventArgs{
BucketName: bucket,
@@ -546,25 +548,38 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
replicationStatus = rinfos.ReplicationStatus()
prevStatus := dobj.DeleteMarkerReplicationStatus()
if dobj.VersionID != "" {
prevStatus = replication.StatusType(dobj.VersionPurgeStatus())
replicationStatus = replication.StatusType(rinfos.VersionPurgeStatus())
if isPurge {
prevStatus = purgeReplicationStatus(dobj.VersionPurgeStatus())
replicationStatus = purgeReplicationStatus(rinfos.VersionPurgeStatus())
}
// to decrement pending count later.
for _, rinfo := range rinfos.Targets {
if rinfo.ReplicationStatus != rinfo.PrevReplicationStatus {
globalReplicationStats.Load().Update(dobj.Bucket, rinfo, replicationStatus,
prevStatus)
status, previous := rinfo.ReplicationStatus, rinfo.PrevReplicationStatus
if isPurge {
status = purgeReplicationStatus(rinfo.VersionPurgeStatus)
previous = purgeReplicationStatus(dobj.ReplicationState.PurgeTargets[rinfo.Arn])
}
if status != previous {
globalReplicationStats.Load().Update(dobj.Bucket, rinfo, status, previous)
}
}
eventName := event.ObjectReplicationComplete
if replicationStatus == replication.Failed {
eventName = event.ObjectReplicationFailed
dobj.RetryCount++
globalReplicationPool.Get().queueMRFSave(dobj.ToMRFEntry())
}
drs := getReplicationState(rinfos, dobj.ReplicationState, dobj.VersionID)
if isPurge {
// A purge must not rewrite the marker's creation/replica metadata.
// Multiple serialized empty target statuses can parse as nonempty,
// so explicitly send an empty creation update to the metadata writer.
drs.ReplicationStatusInternal = ""
drs.Targets = nil
drs.ReplicaStatus = ""
}
if replicationStatus != prevStatus {
drs.ReplicationTimeStamp = UTCNow()
}
@@ -606,6 +621,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
}
func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationInfo, tgt *TargetClient) (rinfo replicatedTargetInfo) {
isPurge := dobj.isVersionPurge()
versionID := dobj.DeleteMarkerVersionID
if versionID == "" {
versionID = dobj.VersionID
@@ -615,42 +631,50 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
rinfo.OpType = dobj.OpType
rinfo.endpoint = tgt.EndpointURL().Host
rinfo.secure = tgt.EndpointURL().Scheme == "https"
// A purge leaves ReplicationStatus empty: the metadata writer interprets
// that as preserving the entire creation-status block, including targets
// outside this fan-out. Only VersionPurgeStatus records the purge outcome.
defer func() {
if rinfo.ReplicationStatus == replication.Completed && tgt.ResetID != "" && dobj.OpType == replication.ExistingObjectReplicationType {
completed := rinfo.ReplicationStatus == replication.Completed
if isPurge {
completed = rinfo.VersionPurgeStatus == replication.VersionPurgeComplete
}
if completed && tgt.ResetID != "" && dobj.OpType == replication.ExistingObjectReplicationType {
rinfo.ResyncTimestamp = fmt.Sprintf("%s;%s", UTCNow().Format(http.TimeFormat), tgt.ResetID)
}
}()
if dobj.VersionID == "" && rinfo.PrevReplicationStatus == replication.Completed && dobj.OpType != replication.ExistingObjectReplicationType {
if !isPurge && rinfo.PrevReplicationStatus == replication.Completed && dobj.OpType != replication.ExistingObjectReplicationType {
rinfo.ReplicationStatus = rinfo.PrevReplicationStatus
return rinfo
}
if dobj.VersionID != "" && rinfo.VersionPurgeStatus == replication.VersionPurgeComplete {
if isPurge && rinfo.VersionPurgeStatus == replication.VersionPurgeComplete {
return rinfo
}
if globalBucketTargetSys.isOffline(tgt.EndpointURL()) {
replLogOnceIf(ctx, fmt.Errorf("remote target is offline for bucket:%s arn:%s", dobj.Bucket, tgt.ARN), "replication-target-offline-delete-"+tgt.ARN)
rinfo.Err = fmt.Errorf("remote target is offline for bucket:%s arn:%s", dobj.Bucket, tgt.ARN)
replLogOnceIf(ctx, rinfo.Err, "replication-target-offline-delete-"+tgt.ARN)
sendEvent(eventArgs{
BucketName: dobj.Bucket,
Object: ObjectInfo{
Bucket: dobj.Bucket,
Name: dobj.ObjectName,
VersionID: dobj.VersionID,
VersionID: versionID,
DeleteMarker: dobj.DeleteMarker,
},
UserAgent: "Internal: [Replication]",
Host: globalLocalNodeName,
EventName: event.ObjectReplicationNotTracked,
})
if dobj.VersionID == "" {
rinfo.ReplicationStatus = replication.Failed
} else {
if isPurge {
rinfo.VersionPurgeStatus = replication.VersionPurgeFailed
} else {
rinfo.ReplicationStatus = replication.Failed
}
return rinfo
}
// early return if already replicated delete marker for existing object replication/ healing delete markers
if dobj.DeleteMarkerVersionID != "" {
if !isPurge && dobj.DeleteMarkerVersionID != "" {
toi, err := tgt.StatObject(ctx, tgt.Bucket, dobj.ObjectName, minio.StatObjectOptions{
VersionID: versionID,
Internal: minio.AdvancedGetOptions{
@@ -662,16 +686,10 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
switch {
case isErrMethodNotAllowed(serr):
// delete marker already replicated
if dobj.VersionID == "" && rinfo.VersionPurgeStatus.Empty() {
rinfo.ReplicationStatus = replication.Completed
return rinfo
}
rinfo.ReplicationStatus = replication.Completed
return rinfo
case isErrObjectNotFound(serr), isErrVersionNotFound(serr):
// version being purged is already not found on target.
if !rinfo.VersionPurgeStatus.Empty() {
rinfo.VersionPurgeStatus = replication.VersionPurgeComplete
return rinfo
}
// The marker still needs to be created on the target.
case isErrReadQuorum(serr), isErrWriteQuorum(serr):
// destination has some quorum issues, perform removeObject() anyways
// to complete the operation.
@@ -691,7 +709,7 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
rmErr := tgt.RemoveObject(ctx, tgt.Bucket, dobj.ObjectName, minio.RemoveObjectOptions{
VersionID: versionID,
Internal: minio.AdvancedRemoveOptions{
ReplicationDeleteMarker: dobj.DeleteMarkerVersionID != "",
ReplicationDeleteMarker: !isPurge && dobj.DeleteMarkerVersionID != "",
ReplicationMTime: dobj.DeleteMarkerMTime.Time,
ReplicationStatus: minio.ReplicationStatusReplica,
ReplicationRequest: true, // always set this to distinguish between `mc mirror` replication and serverside
@@ -699,20 +717,20 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
})
if rmErr != nil {
rinfo.Err = rmErr
if dobj.VersionID == "" {
rinfo.ReplicationStatus = replication.Failed
} else {
if isPurge {
rinfo.VersionPurgeStatus = replication.VersionPurgeFailed
} else {
rinfo.ReplicationStatus = replication.Failed
}
replLogIf(ctx, fmt.Errorf("unable to replicate delete marker to %s: %s/%s(%s): %w", tgt.EndpointURL(), tgt.Bucket, dobj.ObjectName, versionID, rmErr))
if rmErr != nil && minio.IsNetworkOrHostDown(rmErr, true) && !globalBucketTargetSys.isOffline(tgt.EndpointURL()) {
globalBucketTargetSys.markOffline(tgt.EndpointURL())
}
} else {
if dobj.VersionID == "" {
rinfo.ReplicationStatus = replication.Completed
} else {
if isPurge {
rinfo.VersionPurgeStatus = replication.VersionPurgeComplete
} else {
rinfo.ReplicationStatus = replication.Completed
}
}
return rinfo
@@ -781,6 +799,18 @@ func (m caseInsensitiveMap) Lookup(key string) (string, bool) {
return "", false
}
// replicationTaggingTimestamp carries a recorded removal even when tags are
// empty. Only legacy nonempty tags use ModTime; absence is not a tombstone.
func replicationTaggingTimestamp(objInfo ObjectInfo) (time.Time, error) {
if stamp, ok := caseInsensitiveMap(objInfo.UserDefined).Lookup(ReservedMetadataPrefixLower + TaggingTimestamp); ok {
return time.Parse(time.RFC3339Nano, stamp)
}
if objInfo.UserTags != "" {
return objInfo.ModTime, nil
}
return time.Time{}, nil
}
func putReplicationOpts(ctx context.Context, sc string, objInfo ObjectInfo) (putOpts minio.PutObjectOptions, isMP bool, err error) {
meta := make(map[string]string)
isSSEC := crypto.SSEC.IsEncrypted(objInfo.UserDefined)
@@ -850,17 +880,12 @@ func putReplicationOpts(ctx context.Context, sc string, objInfo ObjectInfo) (put
tag, _ := tags.ParseObjectTags(objInfo.UserTags)
if tag != nil {
putOpts.UserTags = tag.ToMap()
// set tag timestamp in opts
tagTimestamp := objInfo.ModTime
if tagTmstampStr, ok := objInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]; ok {
tagTimestamp, err = time.Parse(time.RFC3339Nano, tagTmstampStr)
if err != nil {
return putOpts, false, err
}
}
putOpts.Internal.TaggingTimestamp = tagTimestamp
}
}
putOpts.Internal.TaggingTimestamp, err = replicationTaggingTimestamp(objInfo)
if err != nil {
return putOpts, false, err
}
lkMap := caseInsensitiveMap(objInfo.UserDefined)
if lang, ok := lkMap.Lookup(xhttp.ContentLanguage); ok {
@@ -1003,6 +1028,12 @@ func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replicati
if (oi2.UserTagCount > 0 && !reflect.DeepEqual(oi2Map, t.ToMap())) || (oi2.UserTagCount != len(t.ToMap())) {
return replicateMetadata
}
// HEAD does not report the tag revision. Equal values can hide a newer
// deletion or re-addition, so scheduled metadata/heal work must deliver it.
// Completed objects are still excluded by the existing scanner gates.
if _, ok := caseInsensitiveMap(oi1.UserDefined).Lookup(ReservedMetadataPrefixLower + TaggingTimestamp); ok {
return replicateMetadata
}
// Compare only necessary headers
compareKeys := []string{
@@ -1269,9 +1300,6 @@ func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI Obje
oi.UserDefined[targetResetHeader(rinfo.Arn)] = rinfo.ResyncTimestamp
}
}
if ri.UserTags != "" {
oi.UserDefined[xhttp.AmzObjectTagging] = ri.UserTags
}
return dsc, nil
},
}
@@ -1689,14 +1717,11 @@ applyAction:
if _, ok := lkMap.Lookup(xhttp.AmzObjectLockRetainUntilDate); ok {
dstOpts.Internal.RetentionTimestamp = objInfo.ModTime
}
if objInfo.UserTags != "" {
dstOpts.Internal.TaggingTimestamp = objInfo.ModTime
}
if tagTmStr, ok := lkMap.Lookup(ReservedMetadataPrefixLower + TaggingTimestamp); ok {
ondiskTimestamp, err := time.Parse(time.RFC3339, tagTmStr)
if err == nil {
dstOpts.Internal.TaggingTimestamp = ondiskTimestamp
}
dstOpts.Internal.TaggingTimestamp, rinfo.Err = replicationTaggingTimestamp(objInfo)
if rinfo.Err != nil {
rinfo.ReplicationStatus = replication.Failed
replLogIf(ctx, fmt.Errorf("invalid tagging timestamp for object %s/%s(%s): %w", bucket, object, objInfo.VersionID, rinfo.Err))
return rinfo
}
if retTmStr, ok := lkMap.Lookup(ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp); ok {
ondiskTimestamp, err := time.Parse(time.RFC3339, retTmStr)
@@ -1916,11 +1941,26 @@ func filterReplicationStatusMetadata(metadata map[string]string) map[string]stri
// DeletedObjectReplicationInfo has info on deleted object
type DeletedObjectReplicationInfo struct {
DeletedObject
Bucket string
EventType string
OpType replication.Type
ResetID string
TargetArn string
Bucket string
EventType string
OpType replication.Type
ResetID string
TargetArn string
RetryCount int
}
// isVersionPurge also recognizes the old marker-shaped purge task. Use the
// operation's state, rather than one target's possibly missing purge entry.
func (di DeletedObjectReplicationInfo) isVersionPurge() bool {
return di.VersionID != "" || di.DeleteMarkerVersionID != "" && !di.VersionPurgeStatus().Empty()
}
// Purge metadata uses COMPLETE; operation statistics and audit use COMPLETED.
func purgeReplicationStatus(status VersionPurgeStatusType) replication.StatusType {
if replication.StatusType(status) == replication.CompletedLegacy {
return replication.Completed
}
return replication.StatusType(status)
}
// ToMRFEntry returns the relevant info needed by MRF
@@ -1930,9 +1970,10 @@ func (di DeletedObjectReplicationInfo) ToMRFEntry() MRFReplicateEntry {
versionID = di.VersionID
}
return MRFReplicateEntry{
Bucket: di.Bucket,
Object: di.ObjectName,
versionID: versionID,
Bucket: di.Bucket,
Object: di.ObjectName,
versionID: versionID,
RetryCount: di.RetryCount,
}
}
@@ -2424,6 +2465,7 @@ func (p *ReplicationPool) queueReplicaDeleteTask(doi DeletedObjectReplicationInf
case <-p.ctx.Done():
case ch <- doi:
default:
doi.RetryCount++
p.queueMRFSave(doi.ToMRFEntry())
p.mu.RLock()
prio := p.priority
@@ -3787,9 +3829,10 @@ func queueReplicationHeal(ctx context.Context, bucket string, oi ObjectInfo, rcf
DeleteMarkerMTime: DeleteMarkerMTime{roi.ModTime},
DeleteMarker: roi.DeleteMarker,
},
Bucket: roi.Bucket,
OpType: replication.HealReplicationType,
EventType: ReplicateHealDelete,
Bucket: roi.Bucket,
OpType: replication.HealReplicationType,
EventType: ReplicateHealDelete,
RetryCount: retryCount,
}
// heal delete marker replication failure or versioned delete replication failure
if roi.ReplicationStatus == replication.Pending ||
@@ -4072,7 +4115,12 @@ func (p *ReplicationPool) queueMRFHeal() error {
VersionID: vID,
})
cancel()
if err != nil {
// A versioned marker lookup returns its metadata with a 405. Only
// accept that error with a real, matching marker identity.
validMarker := isErrMethodNotAllowed(err) && oi.DeleteMarker &&
vID != "" && oi.VersionID == vID && !oi.ModTime.IsZero() &&
oi.Bucket == e.Bucket && oi.Name != "" && oi.Name == decodeDirObject(e.Object)
if err != nil && !validMarker || oi.Name == "" {
continue
}
+1
View File
@@ -445,6 +445,7 @@ func buildServerCtxt(ctx *cli.Context, ctxt *serverCtxt) (err error) {
ctxt.SendBufSize = ctx.Int("send-buf-size")
ctxt.RecvBufSize = ctx.Int("recv-buf-size")
ctxt.IdleTimeout = ctx.Duration("idle-timeout")
ctxt.ReadHeaderTimeout = ctx.Duration("read-header-timeout")
ctxt.UserTimeout = ctx.Duration("conn-user-timeout")
if conf := ctx.String("config"); len(conf) > 0 {
+1 -1
View File
@@ -1179,7 +1179,7 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
switch {
case gerr == nil:
reconcileStoredObjectLock(fi.Metadata, storedObjectLockState(curr.UserDefined))
reconcileStoredObjectTags(fi.Metadata, curr.UserDefined)
reconcileStoredObjectTags(fi.Metadata, curr.UserTags, curr.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
case isErrVersionNotFound(gerr) || isErrObjectNotFound(gerr):
// No existing version to order against: keep the upload's own accepted
// lock, including a pre-upgrade upload that persisted values without
+6 -2
View File
@@ -135,7 +135,7 @@ func (er erasureObjects) CopyObject(ctx context.Context, srcBucket, srcObject, d
if dstOpts.ReplicaLockReconcile {
reconcileStoredObjectLock(srcInfo.UserDefined, storedObjectLockState(fi.Metadata))
reconcileStoredObjectTags(srcInfo.UserDefined, fi.Metadata)
reconcileStoredObjectTags(srcInfo.UserDefined, fi.Metadata[xhttp.AmzObjectTagging], fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp])
}
filterOnlineDisksInplace(fi, metaArr, onlineDisks)
@@ -1311,7 +1311,7 @@ func (er erasureObjects) putObject(ctx context.Context, bucket string, object st
// existing version contributes independently ordered lock and tags.
if opts.ReplicaLockReconcile && err == nil {
reconcileStoredObjectLock(opts.UserDefined, storedObjectLockState(obj.UserDefined))
reconcileStoredObjectTags(opts.UserDefined, obj.UserDefined)
reconcileStoredObjectTags(opts.UserDefined, obj.UserTags, obj.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
}
}
@@ -2327,7 +2327,11 @@ func (er erasureObjects) PutObjectTags(ctx context.Context, bucket, object strin
fi.Metadata[xhttp.AmzObjectTagging] = tags
fi.ReplicationState = opts.PutReplicationState()
stamp := monotonicTaggingTimestamp(opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp])
maps.Copy(fi.Metadata, opts.UserDefined)
if stamp != "" {
fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = stamp
}
if err = er.updateObjectMeta(ctx, bucket, object, fi, onlineDisks); err != nil {
return ObjectInfo{}, toObjectErr(err, bucket, object)
+35 -11
View File
@@ -126,10 +126,9 @@ func mergedPoolObjectInfo(copies []PoolObjInfo) ObjectInfo {
stamp, _ := time.Parse(time.RFC3339Nano, ts)
if olderThan(oi.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], stamp) {
oi.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = ts
oi.UserDefined[xhttp.AmzObjectTagging] = copy.ObjInfo.UserDefined[xhttp.AmzObjectTagging]
oi.UserTags = copy.ObjInfo.UserTags
}
}
oi.UserTags = oi.UserDefined[xhttp.AmzObjectTagging]
return oi
}
@@ -190,6 +189,12 @@ func (z *erasureServerPools) updatePoolMetadata(ctx context.Context, bucket, obj
changes[key] = ""
}
}
// Metadata callbacks may explicitly replace tags in the raw write map.
// Otherwise retain the merged value from the ObjectInfo read model.
tags, ok := updated.UserDefined[xhttp.AmzObjectTagging]
if !ok {
tags = updated.UserTags
}
state := storedObjectLockState(updated.UserDefined)
opts.VersionID = updated.VersionID
if opts.VersionID == "" {
@@ -199,11 +204,13 @@ func (z *erasureServerPools) updatePoolMetadata(ctx context.Context, bucket, obj
opts.EvalMetadataFn = func(oi *ObjectInfo, _ error) (ReplicateDecision, error) {
maps.Copy(oi.UserDefined, changes)
replaceObjectLockMetadata(oi.UserDefined, state)
for _, key := range []string{xhttp.AmzObjectTagging, ReservedMetadataPrefixLower + TaggingTimestamp} {
value, exists := updated.UserDefined[key]
if exists || oi.UserDefined[key] != "" {
oi.UserDefined[key] = value
}
// Reassemble the tag value and its ordering timestamp for storage.
if tags != "" || oi.UserTags != "" {
oi.UserDefined[xhttp.AmzObjectTagging] = tags
}
key := ReservedMetadataPrefixLower + TaggingTimestamp
if value, exists := updated.UserDefined[key]; exists || oi.UserDefined[key] != "" {
oi.UserDefined[key] = value
}
return ReplicateDecision{}, nil
}
@@ -220,19 +227,36 @@ func (z *erasureServerPools) updatePoolMetadata(ctx context.Context, bucket, obj
return primary, nil
}
func reconcileStoredObjectTags(metadata, stored map[string]string) {
// Pass the stored tag value explicitly: ObjectInfo.UserDefined excludes it,
// whereas FileInfo.Metadata retains the raw storage key.
func reconcileStoredObjectTags(metadata map[string]string, storedTags, storedTimestamp string) {
key := ReservedMetadataPrefixLower + TaggingTimestamp
stamp, err := time.Parse(time.RFC3339Nano, stored[key])
stamp, err := time.Parse(time.RFC3339Nano, storedTimestamp)
if err != nil {
return
}
incoming, err := time.Parse(time.RFC3339Nano, metadata[key])
if err != nil || !stamp.Before(incoming) {
metadata[key] = stored[key]
metadata[xhttp.AmzObjectTagging] = stored[xhttp.AmzObjectTagging]
metadata[key] = storedTimestamp
metadata[xhttp.AmzObjectTagging] = storedTags
}
}
// Local tagging mutations must advance the revision they overwrite, even when
// a request's clock or lock acquisition order is behind the stored revision.
// Replica writes use reconcileStoredObjectTags instead of minting a revision.
func monotonicTaggingTimestamp(incoming, stored string) string {
requested, err := time.Parse(time.RFC3339Nano, incoming)
if err != nil {
return incoming
}
current, err := time.Parse(time.RFC3339Nano, stored)
if err != nil || requested.After(current) {
return incoming
}
return current.Add(time.Nanosecond).UTC().Format(time.RFC3339Nano)
}
// A restored version still owns its tier reference even while IsRemote is
// false. Only the last copy of a reference may schedule its contents for GC.
func sharesTierObject(oi ObjectInfo, copies []PoolObjInfo) bool {
@@ -0,0 +1,172 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"testing"
"time"
)
// A multipart completion's If-Match must be evaluated against the logical
// latest object across pools, not against the copy local to a pool.
//
// Multi-pool write placement is not sticky (getPoolIdx picks by available
// space even for existing objects), so an upload and a newer overwrite of
// the same name routinely end up in different pools. Uploads are pinned to
// their pools directly: routing through z.NewMultipartUpload would make the
// placement depend on the space-weighted random choice.
func TestPoolsMultipartConditionalUsesLogicalLatest(t *testing.T) {
z, bucket := consistencyPools(t)
ctx := t.Context()
ifMatch := func(etag string) (opts ObjectOptions) {
return ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return oi.ETag != etag
},
}
}
uploadPart := func(t *testing.T, bucket, object, uploadID string) []CompletePart {
t.Helper()
pi, err := z.PutObjectPart(ctx, bucket, object, uploadID, 1,
mustGetPutObjReader(t, bytes.NewBufferString("part"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
return []CompletePart{{PartNumber: 1, ETag: pi.ETag}}
}
// Scenario A: the uploaded If-Match carries the stale ETag of the pool-0
// copy while the logical latest object lives in pool 1. The completion
// must fail with 412 instead of shadowing the newer logical state.
objectA := "cond-mp-stale-etag"
base := time.Now()
oldA := putConsistencyObject(t, z, bucket, objectA, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpA, err := z.serverPools[0].NewMultipartUpload(ctx, bucket, objectA, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerA := putConsistencyObject(t, z, bucket, objectA, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
latest, _, err := z.getLatestObjectInfoWithIdx(ctx, bucket, objectA, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if latest.ETag != newerA.ETag {
t.Fatalf("logical latest should be the pool-1 copy: got %s want %s", latest.ETag, newerA.ETag)
}
if _, err = z.CompleteMultipartUpload(ctx, bucket, objectA, mpA.UploadID,
uploadPart(t, bucket, objectA, mpA.UploadID), ifMatch(oldA.ETag)); err == nil {
t.Fatal("If-Match with the stale pool-0 ETag must not complete over the newer pool-1 object")
} else if _, ok := err.(PreConditionFailed); !ok {
t.Fatalf("expected PreconditionFailed, got %v", err)
}
if latest, _, err = z.getLatestObjectInfoWithIdx(ctx, bucket, objectA, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if latest.ETag != newerA.ETag {
t.Fatalf("the newer pool-1 object must remain the logical latest, got %s", latest.ETag)
}
// Scenario B: the uploaded If-Match carries the logical latest ETag (the
// pool-1 copy) while the upload sits next to the stale pool-0 copy. The
// precondition is satisfied, so completion must succeed and its result
// must become the logical latest.
objectB := "cond-mp-latest-etag"
putConsistencyObject(t, z, bucket, objectB, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpB, err := z.serverPools[0].NewMultipartUpload(ctx, bucket, objectB, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerB := putConsistencyObject(t, z, bucket, objectB, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
oiB, err := z.CompleteMultipartUpload(ctx, bucket, objectB, mpB.UploadID,
uploadPart(t, bucket, objectB, mpB.UploadID), ifMatch(newerB.ETag))
if err != nil {
t.Fatalf("If-Match with the logical latest ETag must complete, got %v", err)
}
if oiB.ETag == "" {
t.Fatal("completion returned an empty ETag")
}
if latest, _, err = z.getLatestObjectInfoWithIdx(ctx, bucket, objectB, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if latest.ETag != oiB.ETag {
t.Fatalf("the completed object must be the logical latest: got %s want %s", latest.ETag, oiB.ETag)
}
// Scenario C: the upload lives in pool 1 with the newer copy while pool 0
// holds the stale one. A set-local evaluation order would let pool 0's
// stale copy fail the request before pool 1 is reached; the logical
// latest ETag must complete.
objectC := "cond-mp-upload-other-pool"
putConsistencyObject(t, z, bucket, objectC, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpC, err := z.serverPools[1].NewMultipartUpload(ctx, bucket, objectC, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerC := putConsistencyObject(t, z, bucket, objectC, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
if _, err = z.CompleteMultipartUpload(ctx, bucket, objectC, mpC.UploadID,
uploadPart(t, bucket, objectC, mpC.UploadID), ifMatch(newerC.ETag)); err != nil {
t.Fatalf("If-Match with the logical latest ETag must complete regardless of upload pool, got %v", err)
}
// Scenario D: pool 1 holds the newer copy but cannot be read. An
// unreadable pool may contain the newest state, so the unverifiable
// condition must fail the request rather than pass it against pool 0's
// stale ETag.
objectD := "cond-mp-unreadable-pool"
oldD := putConsistencyObject(t, z, bucket, objectD, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpD, err := z.serverPools[0].NewMultipartUpload(ctx, bucket, objectD, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerD := putConsistencyObject(t, z, bucket, objectD, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
if latest, _, err = z.getLatestObjectInfoWithIdx(ctx, bucket, objectD, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if latest.ETag != newerD.ETag {
t.Fatalf("logical latest before faulting pool 1 should be its copy: got %s want %s", latest.ETag, newerD.ETag)
}
set := z.serverPools[1].getHashedSet(objectD)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = consistencyReadFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: objectD}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
_, err = z.CompleteMultipartUpload(ctx, bucket, objectD, mpD.UploadID,
uploadPart(t, bucket, objectD, mpD.UploadID), ifMatch(oldD.ETag))
if !isErrReadQuorum(err) {
t.Fatalf("expected an insufficient read quorum error, got %v", err)
}
}
@@ -0,0 +1,341 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"encoding/xml"
"errors"
"fmt"
"io"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
)
func multipartConditionRequest(t *testing.T, router http.Handler, method, target, body string, headers map[string]string) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, target, int64(len(body)), strings.NewReader(body), globalActiveCred.AccessKey, globalActiveCred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
return rec
}
// The server writes the wire spelling ETag directly into Header; unlike a
// network response, httptest's map has not canonicalized it to Etag.
func multipartConditionResponseETag(rec *httptest.ResponseRecorder) string {
for key, values := range rec.Header() {
if strings.EqualFold(key, xhttp.ETag) && len(values) > 0 {
return strings.Trim(values[0], "\"")
}
}
return ""
}
func multipartConditionUpload(t *testing.T, z *erasureServerPools, bucket, object string, owner int, opts ObjectOptions) (string, []CompletePart) {
t.Helper()
mp, err := z.serverPools[owner].NewMultipartUpload(t.Context(), bucket, object, opts)
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[owner].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1, mustGetPutObjReader(t, bytes.NewBufferString("replacement"), 11, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
return mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}
}
func multipartConditionCompleteBody(parts []CompletePart) string {
data, err := xml.Marshal(CompleteMultipartUpload{Parts: parts})
if err != nil {
panic(err)
}
return string(data)
}
func multipartConditionError(t *testing.T, rec *httptest.ResponseRecorder, code string) {
t.Helper()
decoder := xml.NewDecoder(strings.NewReader(rec.Body.String()))
var response APIErrorResponse
if err := decoder.Decode(&response); err != nil || response.Code != code {
t.Fatalf("expected %s error, got %q: %v", code, rec.Body.String(), err)
}
if err := decoder.Decode(&response); err != io.EOF {
t.Fatalf("expected exactly one error response, got %q: %v", rec.Body.String(), err)
}
}
// State is deliberately placed per pool; the final request uses signed HTTP
// and the real handler, precondition callback, erasure metadata and rename.
func TestPoolsMultipartConditionalHTTPMatrix(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
for owner := range 2 {
for _, localOld := range []bool{false, true} {
for _, condition := range []string{"match-old", "match-current", "none-match"} {
t.Run(fmt.Sprintf("owner=%d/old=%t/%s", owner, localOld, condition), func(t *testing.T) {
object := fmt.Sprintf("http-%d-%t-%s", owner, localOld, condition)
oldETag := "old-does-not-exist"
if localOld {
oldETag = putConsistencyObject(t, z, bucket, object, owner, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)}).ETag
}
id, parts := multipartConditionUpload(t, z, bucket, object, owner, ObjectOptions{})
current := putConsistencyObject(t, z, bucket, object, 1-owner, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
head := multipartConditionRequest(t, router, http.MethodHead, getGetObjectURL("", bucket, object), "", nil)
if head.Code != 200 || multipartConditionResponseETag(head) != current.ETag {
t.Fatalf("bad HEAD: %d %v", head.Code, head.Header())
}
h := map[string]string{xhttp.IfMatch: "\"" + oldETag + "\""}
want := 412
if condition == "match-current" {
h[xhttp.IfMatch] = "\"" + current.ETag + "\""
want = 200
}
if condition == "none-match" {
h = map[string]string{xhttp.IfNoneMatch: "*"}
}
rec := multipartConditionRequest(t, router, http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, id), multipartConditionCompleteBody(parts), h)
t.Logf("HEAD current=%s, requested=%v, complete HTTP=%d", current.ETag, h, rec.Code)
if rec.Code != want {
t.Errorf("want HTTP %d, got %d: %s", want, rec.Code, rec.Body.String())
}
if want == 412 {
multipartConditionError(t, rec, "PreconditionFailed")
got := multipartConditionRequest(t, router, http.MethodGet, getGetObjectURL("", bucket, object), "", nil)
if got.Code != 200 || got.Body.String() != "current" {
t.Errorf("rejected request must preserve current data: %d %q", got.Code, got.Body.String())
}
if _, err := z.serverPools[owner].ListObjectParts(t.Context(), bucket, object, id, 0, 10, ObjectOptions{}); err != nil {
t.Errorf("rejected request consumed upload: %v", err)
}
}
})
}
}
}
}
func TestPoolsMultipartConditionalHTTPAbsentObject(t *testing.T) {
for _, deleted := range []bool{false, true} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("delete-marker=%t/if-match=%t", deleted, match), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{VersioningEnabled: deleted})
if err != nil {
t.Fatal(err)
}
object := "http-absent"
if deleted {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
_, err = z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)})
if err != nil {
t.Fatal(err)
}
}
id, parts := multipartConditionUpload(t, z, bucket, object, 0, ObjectOptions{Versioned: deleted})
headers := map[string]string{xhttp.IfNoneMatch: "*"}
want := http.StatusOK
if match {
headers = map[string]string{xhttp.IfMatch: "\"missing\""}
want = http.StatusNotFound
}
rec := multipartConditionRequest(t, router, http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, id), multipartConditionCompleteBody(parts), headers)
if rec.Code != want {
t.Fatalf("expected HTTP %d, got %d: %s", want, rec.Code, rec.Body.String())
}
if match {
multipartConditionError(t, rec, "NoSuchKey")
if _, err := z.serverPools[0].ListObjectParts(t.Context(), bucket, object, id, 0, 10, ObjectOptions{}); err != nil {
t.Errorf("rejected request consumed upload: %v", err)
}
}
})
}
}
}
// All writes below use ordinary signed S3 requests. The result must be correct
// for every placement; the matrix above deterministically covers split pools.
func TestPoolsMultipartConditionalHTTPNormalRouting(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
object := "normal-routing"
url := getPutObjectURL("", bucket, object)
old := multipartConditionRequest(t, router, http.MethodPut, url, "old-data", nil)
if old.Code != http.StatusOK {
t.Fatalf("initial PUT %d: %s", old.Code, old.Body.String())
}
init := multipartConditionRequest(t, router, http.MethodPost, url+"?uploads", "", nil)
if init.Code != http.StatusOK {
t.Fatalf("init %d: %s", init.Code, init.Body.String())
}
var mp InitiateMultipartUploadResponse
if err := xml.Unmarshal(init.Body.Bytes(), &mp); err != nil {
t.Fatal(err)
}
part := multipartConditionRequest(t, router, http.MethodPut, getPutObjectPartURL("", bucket, object, mp.UploadID, "1"), "replacement", nil)
if part.Code != http.StatusOK {
t.Fatalf("part %d: %s", part.Code, part.Body.String())
}
parts := []CompletePart{{PartNumber: 1, ETag: multipartConditionResponseETag(part)}}
newer := multipartConditionRequest(t, router, http.MethodPut, url, "newer-data", nil)
if newer.Code != http.StatusOK {
t.Fatalf("new PUT %d: %s", newer.Code, newer.Body.String())
}
head := multipartConditionRequest(t, router, http.MethodHead, url, "", nil)
if head.Code != http.StatusOK || multipartConditionResponseETag(head) != multipartConditionResponseETag(newer) {
t.Fatalf("HEAD did not pick new object: %d %v", head.Code, head.Header())
}
rec := multipartConditionRequest(t, router, http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, mp.UploadID), multipartConditionCompleteBody(parts), map[string]string{xhttp.IfMatch: "\"" + multipartConditionResponseETag(old) + "\""})
if rec.Code != http.StatusPreconditionFailed {
t.Errorf("stale If-Match should be HTTP 412, got %d: %s", rec.Code, rec.Body.String())
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "newer-data" {
t.Errorf("conditional completion changed newer data: %d %q", get.Code, get.Body.String())
}
}
func TestPoolsMultipartConditionalUnreadablePool(t *testing.T) {
z, bucket := consistencyPools(t)
object := "quorum-with-readable-copy"
old := putConsistencyObject(t, z, bucket, object, 0, "readable", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "hidden-newer", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
id, parts := multipartConditionUpload(t, z, bucket, object, 0, ObjectOptions{})
set := z.serverPools[1].getHashedSet(object)
original := set.getDisks
disks := append([]StorageAPI(nil), original()...)
for i := range disks {
disks[i] = consistencyReadFaultDisk{StorageAPI: disks[i], bucket: bucket, object: object}
}
set.getDisks = func() []StorageAPI { return disks }
defer func() { set.getDisks = original }()
read, readErr := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
t.Logf("ordinary GET lookup: etag=%s err=%v", read.ETag, readErr)
called := 0
_, err := z.CompleteMultipartUpload(t.Context(), bucket, object, id, parts, ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { called++; return oi.ETag != old.ETag }})
if !isErrReadQuorum(err) {
t.Errorf("conditional write must fail on unreadable pool, got %v", err)
}
if called != 0 {
t.Errorf("callback evaluated without complete state: %d calls", called)
}
}
func TestPoolsMultipartConditionalLatestVersionAndCallbackOnce(t *testing.T) {
for _, explicit := range []bool{false, true} {
t.Run(fmt.Sprintf("explicit=%t", explicit), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "tie-version"
opts := ObjectOptions{MTime: UTCNow().Add(-time.Hour), Versioned: explicit}
if explicit {
opts.VersionID = mustGetUUID()
}
current := putConsistencyObject(t, z, bucket, object, 0, "first", opts)
putConsistencyObject(t, z, bucket, object, 1, "second", opts)
if explicit {
current = putConsistencyObject(t, z, bucket, object, 1, "latest-other-version", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
}
id, parts := multipartConditionUpload(t, z, bucket, object, 1, opts)
called := 0
opts.CheckPrecondFn = func(oi ObjectInfo) bool { called++; return oi.ETag != current.ETag }
opts.MTime = time.Time{}
opts.HasIfMatch = true
_, err := z.CompleteMultipartUpload(t.Context(), bucket, object, id, parts, opts)
if err != nil {
t.Errorf("logical current object should match: %v", err)
}
if called != 1 {
t.Errorf("condition evaluated %d times; want exactly once", called)
}
})
}
}
func TestPoolsMultipartConditionalConcurrentCompletes(t *testing.T) {
z, bucket := consistencyPools(t)
object := "concurrent-completes"
old := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "old", ObjectOptions{MTime: old.ModTime})
ids := make([]string, 2)
parts := make([][]CompletePart, 2)
for i := range 2 {
ids[i], parts[i] = multipartConditionUpload(t, z, bucket, object, i, ObjectOptions{})
}
entered := make(chan struct{})
release := make(chan struct{})
var gate, releaseOnce sync.Once
defer releaseOnce.Do(func() { close(release) })
errs := make([]error, 2)
var wg sync.WaitGroup
// Let pool 1's completion hold the object lock before pool 0's starts.
// Once pool 1 commits, a set-local read in pool 0 would still see the
// old ETag and incorrectly accept the second completion.
wg.Add(1)
go func() {
defer wg.Done()
_, errs[1] = z.CompleteMultipartUpload(t.Context(), bucket, object, ids[1], parts[1], ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool {
gate.Do(func() { close(entered); <-release })
return oi.ETag != old.ETag
}})
}()
select {
case <-entered:
case <-time.After(10 * time.Second):
t.Fatal("first completion did not enter its condition callback")
}
started := make(chan struct{})
wg.Add(1)
go func() {
defer wg.Done()
close(started)
_, errs[0] = z.CompleteMultipartUpload(t.Context(), bucket, object, ids[0], parts[0], ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != old.ETag }})
}()
<-started
releaseOnce.Do(func() { close(release) })
wg.Wait()
success, failed := 0, 0
for _, err := range errs {
var p PreConditionFailed
switch {
case err == nil:
success++
case errors.As(err, &p):
failed++
default:
t.Errorf("unexpected completion error %v", err)
}
}
if success != 1 || failed != 1 {
t.Errorf("CAS writers: success=%d conditional failures=%d errors=%v", success, failed, errs)
}
}
@@ -0,0 +1,135 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"errors"
"fmt"
"testing"
"time"
)
func TestPoolsMultipartConditionMatrix(t *testing.T) {
for owner := range 2 {
for _, withOld := range []bool{false, true} {
for _, condition := range []string{"match-current", "match-old", "none-match-any"} {
t.Run(fmt.Sprintf("upload-pool=%d/old-copy=%t/%s", owner, withOld, condition), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-multipart"
oldETag := "arbitrary-old"
if withOld {
old := putConsistencyObject(t, z, bucket, object, owner, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
oldETag = old.ETag
}
mp, err := z.serverPools[owner].NewMultipartUpload(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[owner].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1, mustGetPutObjReader(t, bytes.NewBufferString("replacement"), 11, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
current := putConsistencyObject(t, z, bucket, object, 1-owner, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
visible, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil || visible.ETag != current.ETag {
t.Fatalf("invalid current state: %v", err)
}
opts := ObjectOptions{HasIfMatch: condition != "none-match-any", CheckPrecondFn: func(oi ObjectInfo) bool {
switch condition {
case "match-current":
return oi.ETag != current.ETag
case "match-old":
return oi.ETag != oldETag
default:
return true
}
}}
_, err = z.CompleteMultipartUpload(t.Context(), bucket, object, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, opts)
if condition == "match-current" {
if err != nil {
t.Errorf("correct logical If-Match rejected: %v", err)
}
} else {
var expected PreConditionFailed
if !errors.As(err, &expected) {
t.Errorf("logical condition must fail, got %v", err)
}
}
})
}
}
}
}
func TestPoolsMultipartConditionBoundaries(t *testing.T) {
for _, state := range []string{"missing", "latest-delete-marker", "unreadable-other-pool"} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("%s/if-match=%t", state, match), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-boundary"
versioned := state == "latest-delete-marker"
if versioned {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
if _, err := z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)}); err != nil {
t.Fatal(err)
}
}
mp, err := z.serverPools[0].NewMultipartUpload(t.Context(), bucket, object, ObjectOptions{Versioned: versioned})
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[0].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1, mustGetPutObjReader(t, bytes.NewBufferString("new"), 3, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if state == "unreadable-other-pool" {
set := z.serverPools[1].getHashedSet(object)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = consistencyReadFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: object}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
}
opts := ObjectOptions{Versioned: versioned, HasIfMatch: match, CheckPrecondFn: func(ObjectInfo) bool { return true }}
_, err = z.CompleteMultipartUpload(t.Context(), bucket, object, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, opts)
switch {
case state == "unreadable-other-pool":
if !isErrReadQuorum(err) {
t.Errorf("unreadable pool must not mean absence: %v", err)
}
case match:
if !isErrObjectNotFound(err) {
t.Errorf("If-Match against logical absence should report absence: %v", err)
}
default:
if err != nil {
t.Errorf("If-None-Match against logical absence must succeed: %v", err)
}
}
if err != nil {
if _, lerr := z.serverPools[0].ListObjectParts(t.Context(), bucket, object, mp.UploadID, 0, 10, ObjectOptions{}); lerr != nil {
t.Errorf("failed condition consumed upload: %v", lerr)
}
}
})
}
}
}
@@ -0,0 +1,721 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"crypto/md5"
"encoding/base64"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// Change allocation capacity only; metadata and data use real fixture disks.
// Restoring the adapters also lets encrypted fixtures move the write target
// between requests without changing the production allocation policy.
type conditionalPutCapacityDisk struct {
StorageAPI
full bool
}
func (d conditionalPutCapacityDisk) DiskInfo(ctx context.Context, opts DiskInfoOptions) (DiskInfo, error) {
info, err := d.StorageAPI.DiskInfo(ctx, opts)
info.Total, info.Used = 1<<40, 0
if d.full {
info.Used = info.Total - (1 << 20)
}
info.Free = info.Total - info.Used
return info, err
}
// Keep getDisks immutable while background IAM and storage readers use it.
// GetDisks takes this same mutex when it copies the backing disk list.
func conditionalPutSwapDisks(pool *erasureSets, object string, wrap func(StorageAPI) StorageAPI) func() {
setIndex := pool.getHashedSet(object).setIndex
pool.erasureDisksMu.Lock()
previous := pool.erasureDisks[setIndex]
disks := append([]StorageAPI(nil), previous...)
for i, disk := range disks {
disks[i] = wrap(disk)
}
pool.erasureDisks[setIndex] = disks
pool.erasureDisksMu.Unlock()
return func() {
pool.erasureDisksMu.Lock()
pool.erasureDisks[setIndex] = previous
pool.erasureDisksMu.Unlock()
}
}
func conditionalPutPool(t *testing.T, z *erasureServerPools, object string, target int) func() {
t.Helper()
var restore []func()
for i, pool := range z.serverPools {
restore = append(restore, conditionalPutSwapDisks(pool, object, func(disk StorageAPI) StorageAPI {
return conditionalPutCapacityDisk{StorageAPI: disk, full: i != target}
}))
}
return func() {
for _, fn := range restore {
fn()
}
}
}
func conditionalPutBucket(t *testing.T, z *erasureServerPools, mode string) (string, http.Handler) {
t.Helper()
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{VersioningEnabled: mode != "unversioned"})
if err != nil {
t.Fatal(err)
}
if mode == "suspended" {
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig,
[]byte(`<VersioningConfiguration><Status>Suspended</Status></VersioningConfiguration>`)); err != nil {
t.Fatal(err)
}
}
return bucket, router
}
func TestPoolsConditionalPutHTTP(t *testing.T) {
for _, mode := range []string{"unversioned", "versioned", "suspended"} {
t.Run(mode, func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, mode)
for target := range 2 {
for _, tc := range []struct {
name string
oldCopy, localCurrent bool
condition string
status int
}{
{"stale-etag-accepted", true, false, "old", http.StatusPreconditionFailed},
{"current-etag-rejected", true, false, "current", http.StatusOK},
{"create-only-overwrites-other-pool", false, false, "none", http.StatusPreconditionFailed},
{"current-etag-missing-in-write-pool", false, false, "current", http.StatusOK},
{"control-current-in-write-pool", true, true, "current", http.StatusOK},
{"control-stale-etag-rejected", true, true, "old", http.StatusPreconditionFailed},
{"none-match-current-etag", true, false, "none-current", http.StatusPreconditionFailed},
{"none-match-old-etag", true, false, "none-old", http.StatusOK},
} {
t.Run(fmt.Sprintf("target=%d/%s", target, tc.name), func(t *testing.T) {
object := fmt.Sprintf("%d-%s", target, tc.name)
currentPool := 1 - target
if tc.localCurrent {
currentPool = target
}
opts := ObjectOptions{Versioned: mode == "versioned", VersionSuspended: mode == "suspended", MTime: UTCNow().Add(-time.Hour)}
oldETag := "absent-old"
if tc.oldCopy {
oldETag = putConsistencyObject(t, z, bucket, object, 1-currentPool, "old", opts).ETag
}
opts.MTime = UTCNow().Add(-time.Minute)
current := putConsistencyObject(t, z, bucket, object, currentPool, "current", opts)
defer conditionalPutPool(t, z, object, target)()
idx, err := z.getWritePoolIdx(t.Context(), bucket, object, 11, false)
if err != nil || idx != target {
t.Fatalf("allocation target=%d: idx=%d err=%v", target, idx, err)
}
headers := map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", current.ETag)}
if tc.condition == "old" {
headers[xhttp.IfMatch] = fmt.Sprintf("%q", oldETag)
}
if tc.condition == "none" {
headers = map[string]string{xhttp.IfNoneMatch: "*"}
}
if tc.condition == "none-current" {
headers = map[string]string{xhttp.IfNoneMatch: fmt.Sprintf("%q", current.ETag)}
}
if tc.condition == "none-old" {
headers = map[string]string{xhttp.IfNoneMatch: fmt.Sprintf("%q", oldETag)}
}
url := getPutObjectURL("", bucket, object)
before := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if before.Code != http.StatusOK || before.Body.String() != "current" || multipartConditionResponseETag(before) != current.ETag {
t.Fatalf("invalid current object: %d %q %v", before.Code, before.Body.String(), before.Header())
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != tc.status {
t.Errorf("PUT status=%d want=%d: %s", put.Code, tc.status, put.Body.String())
}
if put.Code == http.StatusPreconditionFailed {
multipartConditionError(t, put, "PreconditionFailed")
if multipartConditionResponseETag(put) != current.ETag || put.Header().Get(xhttp.LastModified) != before.Header().Get(xhttp.LastModified) {
t.Errorf("412 headers do not describe current object: %v", put.Header())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
wantBody, wantETag := "current", current.ETag
if tc.status == http.StatusOK {
wantBody, wantETag = "replacement", fmt.Sprintf("%x", md5.Sum([]byte("replacement")))
}
t.Logf("PUT %d; GET %d bytes=%q ETag=%s", put.Code, get.Code, get.Body.String(), multipartConditionResponseETag(get))
if get.Code != http.StatusOK || get.Body.String() != wantBody || multipartConditionResponseETag(get) != wantETag {
t.Errorf("GET=%d bytes=%q ETag=%s; want %q %s", get.Code, get.Body.String(), multipartConditionResponseETag(get), wantBody, wantETag)
}
})
}
}
})
}
}
func TestPoolsConditionalPutHTTPAbsence(t *testing.T) {
for _, state := range []string{"missing", "uuid-marker", "null-marker"} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("%s/match=%t", state, match), func(t *testing.T) {
z, _ := consistencyPools(t)
mode := "versioned"
if state == "null-marker" {
mode = "suspended"
}
bucket, router := conditionalPutBucket(t, z, mode)
object := "absent-key"
if state != "missing" {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
vid := mustGetUUID()
if state == "null-marker" {
vid = nullVersionID
}
if _, err := z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: vid, DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)}); err != nil {
t.Fatal(err)
}
}
defer conditionalPutPool(t, z, object, 0)()
headers := map[string]string{xhttp.IfNoneMatch: "*"}
want := http.StatusOK
if match {
headers = map[string]string{xhttp.IfMatch: "*"}
want = http.StatusNotFound
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url, "new", headers)
if put.Code != want {
t.Fatalf("PUT %d want %d: %s", put.Code, want, put.Body.String())
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if match {
multipartConditionError(t, put, "NoSuchKey")
if get.Code != http.StatusNotFound {
t.Fatalf("failed PUT exposed data: %d %s", get.Code, get.Body.String())
}
} else if get.Code != http.StatusOK || get.Body.String() != "new" || multipartConditionResponseETag(get) != multipartConditionResponseETag(put) {
t.Fatalf("successful create GET: %d %q %v", get.Code, get.Body.String(), get.Header())
}
})
}
}
}
func TestPoolsConditionalPutUnreadable(t *testing.T) {
for faultPool := range 2 {
for _, present := range []bool{false, true} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("fault-pool=%d/present=%t/match=%t", faultPool, present, match), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
object := "unreadable-key"
if present {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
}
defer conditionalPutPool(t, z, object, 0)()
restoreFault := conditionalPutSwapDisks(z.serverPools[faultPool], object, func(disk StorageAPI) StorageAPI {
return consistencyReadFaultDisk{StorageAPI: disk, bucket: bucket, object: object}
})
defer restoreFault()
headers := map[string]string{xhttp.IfNoneMatch: "*"}
if match {
headers = map[string]string{xhttp.IfMatch: "*"}
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != http.StatusServiceUnavailable {
t.Errorf("unverified PUT must fail: %d %s", put.Code, put.Body.String())
}
called := 0
_, err := z.PutObject(t.Context(), bucket, object, mustGetPutObjReader(t, strings.NewReader("replacement"), 11, "", ""), ObjectOptions{HasIfMatch: match, CheckPrecondFn: func(ObjectInfo) bool { called++; return false }})
if !isErrReadQuorum(err) || called != 0 {
t.Errorf("lookup error=%v callback calls=%d", err, called)
}
restoreFault()
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if present {
if get.Code != http.StatusOK || get.Body.String() != "current" || multipartConditionResponseETag(get) != fmt.Sprintf("%x", md5.Sum([]byte("current"))) {
t.Fatalf("failed PUT changed object: %d %q", get.Code, get.Body.String())
}
} else if get.Code != http.StatusNotFound {
t.Fatalf("failed PUT created object: %d %q", get.Code, get.Body.String())
}
})
}
}
}
}
func TestPoolsConditionalPutEncryptedETag(t *testing.T) {
for _, kind := range []string{"SSE-C", "SSE-S3", "SSE-KMS"} {
t.Run(kind, func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
oldKMS, oldTLS := GlobalKMS, globalIsTLS
GlobalKMS, globalIsTLS = kms.NewStub("conditional-put-key"), true
defer func() { GlobalKMS, globalIsTLS = oldKMS, oldTLS }()
headers := map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}
readHeaders := map[string]string{}
if kind == "SSE-C" {
key := bytes.Repeat([]byte{0x42}, 32)
digest := md5.Sum(key)
headers = map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(digest[:]),
}
readHeaders = maps.Clone(headers)
}
if kind == "SSE-KMS" {
headers[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionKMS
headers[xhttp.AmzServerSideEncryptionKmsID] = "conditional-put-key"
}
object := "encrypted-key"
url := getPutObjectURL("", bucket, object)
restore := conditionalPutPool(t, z, object, 0)
old := multipartConditionRequest(t, router, http.MethodPut, url, "old", headers)
restore()
restore = conditionalPutPool(t, z, object, 1)
current := multipartConditionRequest(t, router, http.MethodPut, url, "current", headers)
restore()
if old.Code != http.StatusOK || current.Code != http.StatusOK {
t.Fatalf("encrypted setup: %d %s / %d %s", old.Code, old.Body.String(), current.Code, current.Body.String())
}
defer conditionalPutPool(t, z, object, 0)()
for _, stale := range []bool{true, false} {
h := maps.Clone(headers)
h[xhttp.IfMatch] = fmt.Sprintf("%q", multipartConditionResponseETag(current))
wantStatus, wantBody, wantETag := http.StatusOK, "replacement", ""
if stale {
h[xhttp.IfMatch] = fmt.Sprintf("%q", multipartConditionResponseETag(old))
wantStatus, wantBody, wantETag = http.StatusPreconditionFailed, "current", multipartConditionResponseETag(current)
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", h)
if put.Code != wantStatus {
t.Fatalf("encrypted condition: %d want %d: %s", put.Code, wantStatus, put.Body.String())
}
if !stale {
wantETag = multipartConditionResponseETag(put)
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", readHeaders)
if get.Code != http.StatusOK || get.Body.String() != wantBody || multipartConditionResponseETag(get) != wantETag {
t.Fatalf("encrypted GET: %d %q %v", get.Code, get.Body.String(), get.Header())
}
}
})
}
}
func TestPoolsConditionalPutVersionSelection(t *testing.T) {
for _, kind := range []string{"public-version", "replica", "replica-preserve-etag", "movement", "no-lock", "tie", "draining"} {
t.Run(kind, func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "version-selection"
addressed := putConsistencyObject(t, z, bucket, object, 1, "addressed", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
opts := ObjectOptions{Versioned: true, VersionID: addressed.VersionID, HasIfMatch: true}
want := current
switch kind {
case "replica":
opts.ReplicaLockReconcile, opts.ReplicationRequest = true, true
want = addressed
case "replica-preserve-etag":
opts.PreserveETag = addressed.ETag
opts.ReplicaLockReconcile, opts.ReplicationRequest = true, true
want = addressed
case "movement":
opts.DataMovement, opts.SrcPoolIdx = true, 1
want = addressed
case "tie":
want = putConsistencyObject(t, z, bucket, object, 0, "tie-winner", ObjectOptions{Versioned: true, MTime: current.ModTime})
case "draining":
z.poolMetaMutex.Lock()
z.poolMeta.Pools[1].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
}
defer conditionalPutPool(t, z, object, 0)()
ctx := t.Context()
if kind == "no-lock" {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
t.Fatal(err)
}
defer lk.Unlock(lkctx)
ctx, opts.NoLock = lkctx.Context(), true
}
called := 0
opts.UserDefined = make(map[string]string)
opts.CheckPrecondFn = func(oi ObjectInfo) bool {
called++
if oi.ETag != want.ETag || oi.VersionID != want.VersionID {
t.Errorf("comparison ETag/version=%s/%s want %s/%s", oi.ETag, oi.VersionID, want.ETag, want.VersionID)
}
return oi.ETag != want.ETag
}
oi, err := z.PutObject(ctx, bucket, object, mustGetPutObjReader(t, strings.NewReader("replacement"), 11, "", ""), opts)
if err != nil || called != 1 {
t.Fatalf("PUT err=%v callback calls=%d", err, called)
}
if oi.VersionID != addressed.VersionID {
t.Fatalf("destination version changed: %s", oi.VersionID)
}
if opts.PreserveETag != "" && oi.ETag != opts.PreserveETag {
t.Fatalf("PreserveETag changed: %s", oi.ETag)
}
})
}
}
func TestPoolsConditionalPutReplicaDuplicateHTTP(t *testing.T) {
for _, null := range []bool{false, true} {
t.Run(fmt.Sprintf("null=%t", null), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "replica-duplicate"
vid := mustGetUUID()
if null {
vid = nullVersionID
}
addressed := putConsistencyObject(t, z, bucket, object, 1, "addressed", ObjectOptions{Versioned: true, VersionID: vid, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
xhttp.MinIOSourceETag: addressed.ETag,
xhttp.MinIOSourceMTime: addressed.ModTime.Format(time.RFC3339Nano),
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url+"?versionId="+vid, "addressed", headers)
if put.Code != http.StatusPreconditionFailed {
t.Fatalf("replica duplicate: %d %s", put.Code, put.Body.String())
}
multipartConditionError(t, put, "PreconditionFailed")
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "current" || multipartConditionResponseETag(get) != current.ETag {
t.Fatalf("duplicate changed current object: %d %q", get.Code, get.Body.String())
}
get = multipartConditionRequest(t, router, http.MethodGet, url+"?versionId="+vid, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "addressed" || multipartConditionResponseETag(get) != addressed.ETag {
t.Fatalf("duplicate changed addressed version: %d %q", get.Code, get.Body.String())
}
})
}
}
func TestPoolsConditionalPutConcurrentHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
for _, match := range []bool{false, true} {
for iteration := range 5 {
t.Run(fmt.Sprintf("match=%t/iteration=%d", match, iteration), func(t *testing.T) {
object := fmt.Sprintf("concurrent-%t-%d", match, iteration)
headers := map[string]string{xhttp.IfNoneMatch: "*"}
if match {
oi := putConsistencyObject(t, z, bucket, object, 1, "old", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
headers = map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", oi.ETag)}
}
defer conditionalPutPool(t, z, object, 0)()
start := make(chan struct{})
results := make(chan *httptest.ResponseRecorder, 2)
url := getPutObjectURL("", bucket, object)
for i := range 2 {
body := fmt.Sprintf("writer-%d", i)
req, err := newTestSignedRequestV4(http.MethodPut, url, int64(len(body)), strings.NewReader(body), globalActiveCred.AccessKey, globalActiveCred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
go func() {
<-start
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
results <- rec
}()
}
close(start)
success, failed, winnerETag := 0, 0, ""
for range 2 {
result := <-results
switch result.Code {
case http.StatusOK:
success++
winnerETag = multipartConditionResponseETag(result)
case http.StatusPreconditionFailed:
failed++
multipartConditionError(t, result, "PreconditionFailed")
default:
t.Errorf("unexpected PUT %d: %s", result.Code, result.Body.String())
}
}
if success != 1 || failed != 1 {
t.Fatalf("success=%d precondition failures=%d", success, failed)
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || multipartConditionResponseETag(get) != winnerETag || fmt.Sprintf("%x", md5.Sum(get.Body.Bytes())) != winnerETag {
t.Fatalf("winner lost: %d %q %v", get.Code, get.Body.String(), get.Header())
}
})
}
}
}
func TestPoolsConditionalPutSerializesMutation(t *testing.T) {
for _, deletion := range []bool{false, true} {
t.Run(fmt.Sprintf("delete=%t", deletion), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-mutation"
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
gate := &consistencyGateReader{Reader: strings.NewReader("replacement"), entered: make(chan struct{}), resume: make(chan struct{})}
release := func() { gate.release.Do(func() { close(gate.resume) }) }
defer release()
reader := mustGetPutObjReader(t, gate, 11, "", "")
written := make(chan error, 1)
go func() {
_, err := z.PutObject(ctx, bucket, object, reader, ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != current.ETag }})
written <- err
}()
select {
case <-gate.entered:
case err := <-written:
t.Fatalf("PUT failed before body read: %v", err)
case <-ctx.Done():
t.Fatal(ctx.Err())
}
mutated := make(chan error, 1)
go func() {
var err error
if deletion {
_, err = z.DeleteObject(ctx, bucket, object, ObjectOptions{})
} else {
_, err = z.PutObjectMetadata(ctx, bucket, object, ObjectOptions{EvalMetadataFn: func(oi *ObjectInfo, _ error) (ReplicateDecision, error) {
oi.UserDefined["x-amz-meta-after-put"] = "present"
return ReplicateDecision{}, nil
}})
}
mutated <- err
}()
select {
case err := <-mutated:
t.Fatalf("mutation escaped PUT lock: %v", err)
case <-time.After(100 * time.Millisecond):
}
release()
if err := <-written; err != nil {
t.Fatal(err)
}
if err := <-mutated; err != nil {
t.Fatal(err)
}
oi, err := z.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if deletion {
if !isErrObjectNotFound(err) {
t.Fatalf("delete lost: %v", err)
}
} else if err != nil || oi.UserDefined["x-amz-meta-after-put"] != "present" || oi.ETag != fmt.Sprintf("%x", md5.Sum([]byte("replacement"))) {
t.Fatalf("metadata/PUT lost: %+v %v", oi, err)
}
})
}
}
func TestSinglePoolConditionalPutHTTP(t *testing.T) {
ctx, cancel := context.WithCancel(t.Context())
obj, dirs, err := prepareErasure16(ctx)
if err != nil {
cancel()
t.Fatal(err)
}
z := obj.(*erasureServerPools)
t.Cleanup(func() { cancel(); z.Shutdown(context.Background()); removeRoots(dirs) })
if !z.SinglePool() {
t.Fatal("fixture is not a single pool")
}
bucket, router := conditionalPutBucket(t, z, "unversioned")
object := "single-pool-condition"
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object)
for _, tc := range []struct {
body, match, none string
status int
}{
{"missing", "*", "", http.StatusNotFound},
{"first", "", "*", http.StatusOK},
{"blocked", "", "*", http.StatusPreconditionFailed},
{"blocked", "stale", "", http.StatusPreconditionFailed},
{"second", fmt.Sprintf("%x", md5.Sum([]byte("first"))), "", http.StatusOK},
} {
rec := multipartConditionRequest(t, router, http.MethodPut, url, tc.body, map[string]string{xhttp.IfMatch: tc.match, xhttp.IfNoneMatch: tc.none})
if rec.Code != tc.status {
t.Fatalf("PUT %d want %d: %s", rec.Code, tc.status, rec.Body.String())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "second" {
t.Fatalf("single-pool GET: %d %q", get.Code, get.Body.String())
}
}
// An internal replica callback without an addressed version is not a public
// condition. Preserve its availability when another pool is unreadable. An
// addressed replica already requires all pools for lock/tag reconciliation.
func TestPoolsConditionalPutReplicaAvailability(t *testing.T) {
for _, addressed := range []bool{false, true} {
t.Run(fmt.Sprintf("addressed=%t", addressed), func(t *testing.T) {
z, _ := consistencyPools(t)
mode := "unversioned"
if addressed {
mode = "versioned"
}
bucket, router := conditionalPutBucket(t, z, mode)
object := "replica-availability"
oi := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: addressed, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
restoreFault := conditionalPutSwapDisks(z.serverPools[1], object, func(disk StorageAPI) StorageAPI {
return consistencyReadFaultDisk{StorageAPI: disk, bucket: bucket, object: object}
})
defer restoreFault()
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
xhttp.MinIOSourceETag: fmt.Sprintf("%x", md5.Sum([]byte("replacement"))),
}
url := getPutObjectURL("", bucket, object)
wantStatus, wantBody := http.StatusOK, "replacement"
if addressed {
url += "?versionId=" + oi.VersionID
wantStatus, wantBody = http.StatusServiceUnavailable, "old"
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != wantStatus {
t.Fatalf("replica PUT: %d want %d: %s", put.Code, wantStatus, put.Body.String())
}
restoreFault()
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != wantBody {
t.Fatalf("replica GET: %d %q", get.Code, get.Body.String())
}
})
}
}
func TestPoolsConditionalPutDestinationVersionHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "client-version"
old := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object) + "?versionId=" + old.VersionID
for _, stale := range []bool{true, false} {
tag, status := current.ETag, http.StatusOK
if stale {
tag, status = old.ETag, http.StatusPreconditionFailed
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", tag)})
if put.Code != status {
t.Fatalf("version-addressed public PUT: %d want %d: %s", put.Code, status, put.Body.String())
}
if got := put.Header()[xhttp.AmzVersionID]; !stale && (len(got) != 1 || got[0] != old.VersionID) {
t.Fatalf("write version changed: %v", put.Header())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "replacement" {
t.Fatalf("addressed GET: %d %q", get.Code, get.Body.String())
}
}
func TestPoolsConditionalPutDeleteMarkerTie(t *testing.T) {
for markerPool := range 2 {
t.Run(fmt.Sprintf("marker-pool=%d", markerPool), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "marker-tie"
mtime := UTCNow().Add(-time.Minute)
putConsistencyObject(t, z, bucket, object, 1-markerPool, "live", ObjectOptions{Versioned: true, MTime: mtime})
if _, err := z.serverPools[markerPool].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: mtime}); err != nil {
t.Fatal(err)
}
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object)
before := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
want := http.StatusPreconditionFailed
if markerPool == 0 {
want = http.StatusOK
if before.Code != http.StatusNotFound {
t.Fatalf("GET tie: %d", before.Code)
}
} else if before.Code != http.StatusOK {
t.Fatalf("GET tie: %d", before.Code)
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", map[string]string{xhttp.IfNoneMatch: "*"})
if put.Code != want {
t.Fatalf("PUT tie: %d want %d: %s", put.Code, want, put.Body.String())
}
})
}
}
// Overlap fixture changes with the real IAM Walk reader instead of relying on
// its periodic refresh timer to expose an unsynchronized disk-adapter swap.
func TestPoolsConditionalPutFixtureConcurrentIAM(t *testing.T) {
z, _ := consistencyPools(t)
conditionalPutBucket(t, z, "unversioned")
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
started, finished := make(chan struct{}), make(chan error, 1)
iam := globalIAMSys
go func() {
close(started)
for range 20 {
if err := iam.Load(ctx, false); err != nil {
finished <- err
return
}
}
finished <- nil
}()
<-started
for i := range 5000 {
restore := conditionalPutPool(t, z, "fixture-concurrent-iam", i%2)
restore()
}
if err := <-finished; err != nil {
t.Fatal(err)
}
}
+459
View File
@@ -0,0 +1,459 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"context"
"fmt"
"maps"
"strings"
"testing"
xhttp "github.com/minio/minio/internal/http"
)
// The fixture places copies directly in real erasure pools. It models a
// duplicated version; it does not claim to exercise a rebalance workflow.
func TestPoolsMetadataUpdatePreservesTags(t *testing.T) {
z, bucket := consistencyPools(t)
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, test := range []struct {
name string
tags, stamps [2]string
winner int
single bool
}{
{name: "newer-secondary", tags: [2]string{"key=old", "key=new"}, stamps: [2]string{old, recent}, winner: 1},
{name: "newer-primary", tags: [2]string{"key=new", "key=old"}, stamps: [2]string{recent, old}},
{name: "empty-secondary", tags: [2]string{"key=old", ""}, stamps: [2]string{old, recent}, winner: 1},
{name: "empty-primary", tags: [2]string{"", "key=old"}, stamps: [2]string{recent, old}},
{name: "single-copy-primary", tags: [2]string{"key=only", ""}, stamps: [2]string{recent, ""}, single: true},
{name: "single-copy-secondary", tags: [2]string{"", "key=only"}, stamps: [2]string{"", recent}, winner: 1, single: true},
{name: "legacy-single-copy", tags: [2]string{"key=legacy", ""}, single: true},
{name: "legacy-duplicates", tags: [2]string{"key=legacy", "key=legacy"}},
} {
t.Run(test.name, func(t *testing.T) {
object := test.name
var original ObjectInfo
for pool := range 2 {
if test.single && pool != test.winner {
continue
}
metadata := map[string]string{
xhttp.AmzObjectTagging: test.tags[pool],
"copy-local": fmt.Sprint(pool),
}
if test.stamps[pool] != "" {
metadata[timestamp] = test.stamps[pool]
}
original = putConsistencyObject(t, z, bucket, object, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, UserDefined: metadata,
})
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: original.VersionID})
if err != nil || got.UserTags != test.tags[pool] || got.UserDefined[timestamp] != test.stamps[pool] {
t.Fatalf("pool %d fixture: tags=%q timestamp=%q err=%v", pool, got.UserTags, got.UserDefined[timestamp], err)
}
if _, exists := got.UserDefined[xhttp.AmzObjectTagging]; exists {
t.Fatal("fixture must use the cleaned ObjectInfo representation")
}
}
wantTags, wantStamp := test.tags[test.winner], test.stamps[test.winner]
called := 0
got, err := z.PutObjectMetadata(t.Context(), bucket, object, ObjectOptions{
VersionID: original.VersionID, MTime: original.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
called++
if current.UserTags != wantTags || current.UserDefined[timestamp] != wantStamp {
t.Errorf("callback tags=%q timestamp=%q; want %q %q", current.UserTags, current.UserDefined[timestamp], wantTags, wantStamp)
}
current.UserDefined["unrelated-update"] = "preserved"
return ReplicateDecision{}, nil
},
})
if err != nil || called != 1 {
t.Fatalf("metadata update: %v, callbacks=%d", err, called)
}
if got.UserTags != wantTags || got.UserDefined[timestamp] != wantStamp {
t.Errorf("response tags=%q timestamp=%q; want %q %q", got.UserTags, got.UserDefined[timestamp], wantTags, wantStamp)
}
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: original.VersionID})
if test.single && pool != test.winner {
if !isErrVersionNotFound(err) {
t.Errorf("metadata update created another copy: %v", err)
}
continue
}
if err != nil {
t.Fatal(err)
}
t.Logf("pool %d persisted tags=%q timestamp=%q", pool, got.UserTags, got.UserDefined[timestamp])
if got.UserTags != wantTags || got.UserDefined[timestamp] != wantStamp {
t.Errorf("pool %d persisted tags=%q timestamp=%q; want %q %q", pool, got.UserTags, got.UserDefined[timestamp], wantTags, wantStamp)
}
if got.UserDefined["unrelated-update"] != "preserved" || got.UserDefined["copy-local"] != fmt.Sprint(pool) {
t.Errorf("pool %d lost unrelated metadata: %v", pool, got.UserDefined)
}
}
})
}
}
func TestReplicaWritesPreserveTagOrdering(t *testing.T) {
z, bucket := consistencyPools(t)
// Pool allocation checks the host's used-space percentage. Present only
// its free space as fixture capacity; all reads and writes still use the
// real disks. This keeps unrelated host disk usage out of the tag test.
for _, pool := range z.serverPools {
for _, set := range pool.sets {
getDisks := set.getDisks
disks := append([]StorageAPI(nil), getDisks()...)
for i := range disks {
disks[i] = tagTestCapacityDisk{StorageAPI: disks[i]}
}
set.getDisks = func() []StorageAPI { return disks }
t.Cleanup(func() { set.getDisks = getDisks })
}
}
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, path := range []struct {
name, operation string
direct bool
winner int
}{
{name: "set-put", operation: "put", direct: true},
{name: "set-multipart", operation: "multipart", direct: true},
{name: "set-copy", operation: "copy", direct: true},
{name: "pools-put", operation: "put", winner: 1},
{name: "pools-multipart", operation: "multipart", winner: 1},
{name: "pools-copy-primary", operation: "copy"},
{name: "pools-copy-secondary", operation: "copy", winner: 1},
} {
for _, test := range []struct {
name, storedTags, incomingTags, storedStamp, incomingStamp, wantTags string
}{
{"stored-newer", "key=stored", "key=incoming", recent, old, "key=stored"},
{"stored-deleted", "", "key=incoming", recent, old, ""},
{"incoming-newer", "key=stored", "key=incoming", old, recent, "key=incoming"},
{"incoming-deleted", "key=stored", "", old, recent, ""},
} {
t.Run(path.name+"/"+test.name, func(t *testing.T) {
object := path.name + "-" + test.name
var original ObjectInfo
for pool := range 2 {
if path.direct && pool != 0 {
continue
}
metadata := map[string]string{
xhttp.AmzObjectTagging: "key=older-copy",
timestamp: "2026-09-09T08:00:00Z",
}
if pool == path.winner {
metadata[xhttp.AmzObjectTagging] = test.storedTags
metadata[timestamp] = test.storedStamp
}
original = putConsistencyObject(t, z, bucket, object, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, UserDefined: metadata,
})
}
opts := ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, ReplicaLockReconcile: true,
UserDefined: map[string]string{
xhttp.AmzObjectTagging: test.incomingTags,
timestamp: test.incomingStamp,
},
}
var got ObjectInfo
var err error
switch path.operation {
case "put":
put := z.PutObject
if path.direct {
put = z.serverPools[0].PutObject
}
got, err = put(t.Context(), bucket, object, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), opts)
case "multipart":
// Persist the incoming tags with the upload, before completion
// reconciles the destination version through its real resolver.
mp, err := z.serverPools[0].NewMultipartUpload(t.Context(), bucket, object, opts)
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[0].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1,
mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
complete := z.CompleteMultipartUpload
if path.direct {
complete = z.serverPools[0].CompleteMultipartUpload
}
got, err = complete(t.Context(), bucket, object, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, ObjectOptions{
Versioned: true, MTime: original.ModTime, ReplicaLockReconcile: true,
})
if err != nil {
t.Fatal(err)
}
case "copy":
src := original
src.metadataOnly = true
src.UserDefined = maps.Clone(opts.UserDefined)
copyObject := z.CopyObject
if path.direct {
copyObject = z.serverPools[0].CopyObject
}
got, err = copyObject(t.Context(), bucket, object, bucket, object, src, ObjectOptions{VersionID: original.VersionID}, opts)
}
if err != nil {
t.Fatal(err)
}
if got.UserTags != test.wantTags || got.UserDefined[timestamp] != recent {
t.Errorf("response tags=%q timestamp=%q; want %q %q", got.UserTags, got.UserDefined[timestamp], test.wantTags, recent)
}
copies := 0
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: original.VersionID})
if isErrVersionNotFound(err) {
continue
}
if err != nil {
t.Fatal(err)
}
copies++
if got.UserTags != test.wantTags || got.UserDefined[timestamp] != recent {
t.Errorf("pool %d persisted tags=%q timestamp=%q; want %q %q", pool, got.UserTags, got.UserDefined[timestamp], test.wantTags, recent)
}
}
if copies != 1 {
t.Errorf("replacement left %d copies; want 1", copies)
}
})
}
}
}
type tagTestCapacityDisk struct{ StorageAPI }
func (d tagTestCapacityDisk) DiskInfo(ctx context.Context, opts DiskInfoOptions) (DiskInfo, error) {
info, err := d.StorageAPI.DiskInfo(ctx, opts)
info.Total, info.Used = info.Free, 0
return info, err
}
func TestMergedPoolObjectInfoTagOrdering(t *testing.T) {
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, test := range []struct {
name, firstStamp, secondStamp, secondTags string
winner int
}{
{"newer", old, recent, "key=second", 1},
{"newer-removal", old, recent, "", 1},
{"equal", recent, recent, "key=second", 0},
{"unordered", "", "", "key=second", 0},
{"missing-first", "", recent, "key=second", 1},
{"missing-second", recent, "", "key=second", 0},
{"invalid-first", "invalid", recent, "key=second", 1},
{"invalid-second", recent, "invalid", "key=second", 0},
} {
t.Run(test.name, func(t *testing.T) {
tagValues := []string{"key=first", test.secondTags}
stamps := []string{test.firstStamp, test.secondStamp}
copies := make([]PoolObjInfo, 2)
before := make([]map[string]string, 2)
for i := range copies {
fi := FileInfo{Metadata: map[string]string{xhttp.AmzObjectTagging: tagValues[i]}}
if stamps[i] != "" {
fi.Metadata[timestamp] = stamps[i]
}
copies[i] = PoolObjInfo{Index: i, ObjInfo: fi.ToObjectInfo("bucket", "object", true)}
before[i] = maps.Clone(copies[i].ObjInfo.UserDefined)
}
got := mergedPoolObjectInfo(copies)
if got.UserTags != tagValues[test.winner] || got.UserDefined[timestamp] != stamps[test.winner] {
t.Errorf("merged tags=%q timestamp=%q; want %q %q", got.UserTags, got.UserDefined[timestamp], tagValues[test.winner], stamps[test.winner])
}
if _, exists := got.UserDefined[xhttp.AmzObjectTagging]; exists {
t.Error("merged ObjectInfo leaked the raw tagging key into UserDefined")
}
for i := range copies {
if !maps.Equal(copies[i].ObjInfo.UserDefined, before[i]) || copies[i].ObjInfo.UserTags != tagValues[i] {
t.Errorf("merge mutated input copy %d", i)
}
}
})
}
}
func TestPoolsMetadataCallbackReplacesTags(t *testing.T) {
z, bucket := consistencyPools(t)
const timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
for _, test := range []struct{ name, tags string }{
{"replace", "key=callback"},
{"remove", ""},
} {
t.Run(test.name, func(t *testing.T) {
var original ObjectInfo
for pool := range 2 {
original = putConsistencyObject(t, z, bucket, test.name, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime,
UserDefined: map[string]string{
xhttp.AmzObjectTagging: []string{"key=old", "key=new"}[pool],
timestamp: []string{"2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"}[pool],
},
})
}
const updatedStamp = "2026-09-09T11:00:00Z"
got, err := z.PutObjectMetadata(t.Context(), bucket, test.name, ObjectOptions{
VersionID: original.VersionID, MTime: original.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
if current.UserTags != "key=new" {
t.Errorf("callback read tags=%q; want key=new", current.UserTags)
}
if _, exists := current.UserDefined[xhttp.AmzObjectTagging]; exists {
t.Error("callback received the raw tagging key")
}
current.UserDefined[xhttp.AmzObjectTagging] = test.tags
current.UserDefined[timestamp] = updatedStamp
return ReplicateDecision{}, nil
},
})
if err != nil {
t.Fatal(err)
}
if got.UserTags != test.tags || got.UserDefined[timestamp] != updatedStamp {
t.Errorf("callback update response tags=%q timestamp=%q", got.UserTags, got.UserDefined[timestamp])
}
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, test.name, ObjectOptions{VersionID: original.VersionID})
if err != nil || got.UserTags != test.tags || got.UserDefined[timestamp] != updatedStamp {
t.Errorf("pool %d did not persist callback tags: tags=%q timestamp=%q err=%v", pool, got.UserTags, got.UserDefined[timestamp], err)
}
}
})
}
}
func TestReconcileStoredObjectTagOrdering(t *testing.T) {
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, test := range []struct {
name, storedStamp, incomingStamp, storedTags string
wantStored bool
}{
{"stored-newer", recent, old, "key=stored", true},
{"incoming-newer", old, recent, "key=stored", false},
{"equal", recent, recent, "key=stored", true},
{"equal-removal", recent, recent, "", true},
{"missing-stored", "", recent, "key=stored", false},
{"missing-incoming", recent, "", "key=stored", true},
{"invalid-stored", "invalid", recent, "key=stored", false},
{"invalid-incoming", recent, "invalid", "key=stored", true},
} {
t.Run(test.name, func(t *testing.T) {
metadata := map[string]string{
xhttp.AmzObjectTagging: "key=incoming",
timestamp: test.incomingStamp,
"unrelated": "preserved",
}
reconcileStoredObjectTags(metadata, test.storedTags, test.storedStamp)
wantTags, wantStamp := "key=incoming", test.incomingStamp
if test.wantStored {
wantTags, wantStamp = test.storedTags, test.storedStamp
}
if metadata[xhttp.AmzObjectTagging] != wantTags || metadata[timestamp] != wantStamp || metadata["unrelated"] != "preserved" {
t.Errorf("reconciled metadata=%v; want tags=%q timestamp=%q and unrelated field preserved", metadata, wantTags, wantStamp)
}
})
}
}
func TestPoolsMetadataUpdatePreservesAbsentTags(t *testing.T) {
z, bucket := consistencyPools(t)
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, removal := range []bool{false, true} {
t.Run(fmt.Sprintf("timestamp-only-removal=%t", removal), func(t *testing.T) {
object := fmt.Sprintf("absent-tags-%t", removal)
var original ObjectInfo
checkStored := func(pool int, wantKey bool, wantTags, wantStamp string) {
t.Helper()
infos, errs := readAllFileInfo(t.Context(), z.serverPools[pool].getHashedSet(object).getDisks(), "", bucket, object, original.VersionID, false, false)
for disk, info := range infos {
if errs[disk] != nil {
t.Fatal(errs[disk])
}
tags, exists := info.Metadata[xhttp.AmzObjectTagging]
if exists != wantKey || tags != wantTags || info.Metadata[timestamp] != wantStamp {
t.Errorf("pool %d disk %d raw tagging key=%t value=%q stamp=%q; want %t %q %q", pool, disk, exists, tags, info.Metadata[timestamp], wantKey, wantTags, wantStamp)
}
}
}
for pool := range 2 {
metadata := map[string]string{}
if removal {
metadata[timestamp] = recent
if pool == 1 {
metadata[xhttp.AmzObjectTagging] = "key=old"
metadata[timestamp] = old
}
}
original = putConsistencyObject(t, z, bucket, object, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, UserDefined: metadata,
})
checkStored(pool, removal && pool == 1, metadata[xhttp.AmzObjectTagging], metadata[timestamp])
}
_, err := z.PutObjectMetadata(t.Context(), bucket, object, ObjectOptions{
VersionID: original.VersionID, MTime: original.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
current.UserDefined["unrelated-update"] = "preserved"
return ReplicateDecision{}, nil
},
})
if err != nil {
t.Fatal(err)
}
wantStamp := ""
if removal {
wantStamp = recent
}
for pool := range 2 {
// A previously non-empty key needs an explicit empty value to
// propagate deletion; an absent key should remain absent.
checkStored(pool, removal && pool == 1, "", wantStamp)
}
})
}
}
+85 -1
View File
@@ -1149,6 +1149,40 @@ func (z *erasureServerPools) PutObject(ctx context.Context, bucket string, objec
}
opts.NoLock = true
// Public write conditions compare the logical current object while the
// pools-layer write lock is held. The destination selected by capacity may
// be empty or stale, and draining pools can still hold the current object.
// Replica callbacks retain their existing addressed-version semantics and
// metadata reconciliation at the set layer.
if opts.CheckPrecondFn != nil && !opts.ReplicationRequest &&
!opts.ReplicaLockReconcile && !opts.DataMovement {
copies, lerr := z.objectPoolInfos(ctx, bucket, object, ObjectOptions{
VersionID: "", // Compare the current object, not the write's version.
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoAuditLog: true,
})
var latest ObjectInfo
if lerr == nil {
latest = copies[0].ObjInfo
if latest.DeleteMarker {
lerr = toObjectErr(errFileNotFound, bucket, object)
}
}
// An unreadable pool may hold the newest object; it is not absence.
if lerr != nil && !isErrObjectNotFound(lerr) && !isErrVersionNotFound(lerr) {
return ObjectInfo{}, lerr
}
if lerr == nil && opts.CheckPrecondFn(latest) {
return ObjectInfo{}, PreConditionFailed{}
}
if lerr != nil && opts.HasIfMatch {
return ObjectInfo{}, lerr
}
// Do not repeat an accepted condition against the destination's copy.
opts.CheckPrecondFn = nil
}
idx, err := z.getWritePoolIdx(ctx, bucket, object, data.Size(), true)
if err != nil {
return ObjectInfo{}, err
@@ -1447,7 +1481,7 @@ func (z *erasureServerPools) CopyObject(ctx context.Context, srcBucket, srcObjec
}
stored := mergedPoolObjectInfo(copies)
reconcileStoredObjectLock(srcInfo.UserDefined, storedObjectLockState(stored.UserDefined))
reconcileStoredObjectTags(srcInfo.UserDefined, stored.UserDefined)
reconcileStoredObjectTags(srcInfo.UserDefined, stored.UserTags, stored.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
idx := copies[0].Index
oi, err := z.serverPools[idx].CopyObject(ctx, srcBucket, srcObject, dstBucket, dstObject, srcInfo, srcOpts, dstOpts)
if err == nil {
@@ -2154,6 +2188,47 @@ func (z *erasureServerPools) CompleteMultipartUpload(ctx context.Context, bucket
defer lk.Unlock(lkctx)
}
opts.NoLock = true
// A conditional completion must be evaluated against the logical
// latest object across pools, under the object write lock held for
// this operation. The pool hosting the upload may only hold a stale
// duplicate, so its set-local check would both accept an outdated
// ETag and reject the current one. An unreadable pool is not
// absence: it may hold the newest copy, so a read that cannot be
// verified fails the request instead of passing the condition.
// Once satisfied, the callback is cleared so the set layer does not
// re-evaluate it against its local copy.
if opts.CheckPrecondFn != nil {
copies, lerr := z.objectPoolInfos(ctx, bucket, encodeDirObject(object), ObjectOptions{
// Conditions always compare the logical current object,
// independently of the completion's destination version.
VersionID: "",
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoAuditLog: true,
})
var latest ObjectInfo
if lerr == nil {
latest = copies[0].ObjInfo
if latest.DeleteMarker {
// A delete-marker latest reads as an absent key, matching
// the set layer's getObjectInfo.
lerr = toObjectErr(errFileNotFound, bucket, object)
}
}
if lerr == nil && opts.CheckPrecondFn(latest) {
return ObjectInfo{}, PreConditionFailed{}
}
if lerr != nil && !isErrVersionNotFound(lerr) && !isErrObjectNotFound(lerr) {
return ObjectInfo{}, lerr
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if lerr != nil && opts.HasIfMatch && (isErrObjectNotFound(lerr) || isErrVersionNotFound(lerr)) {
return ObjectInfo{}, lerr
}
opts.CheckPrecondFn = nil
}
}
// Hold write locks to verify uploaded parts, also disallows any
@@ -3010,6 +3085,15 @@ func (z *erasureServerPools) PutObjectTags(ctx context.Context, bucket, object s
if err != nil {
return ObjectInfo{}, err
}
// Ordinary reads and replication can return any owning pool. Persist one
// revision beyond all copies, so the returned value and every pool agree.
if stamp := opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]; stamp != "" {
for _, copy := range copies {
stamp = monotonicTaggingTimestamp(stamp, copy.ObjInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
}
opts.UserDefined = cloneMSS(opts.UserDefined)
opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = stamp
}
opts.NoLock = true
opts.VersionID = copies[0].ObjInfo.VersionID
if opts.VersionID == "" {
+33 -23
View File
@@ -246,16 +246,6 @@ func extractMetadata(ctx context.Context, mimesHeader ...textproto.MIMEHeader) (
// extractMetadata extracts metadata from map values.
func extractMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
return extractMetadataFromMimeWithReplication(ctx, v, m, false)
}
// extractReplicationMetadataFromMime restores replication-only metadata after the
// caller has validated that the request is a trusted replication write.
func extractReplicationMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
return extractMetadataFromMimeWithReplication(ctx, v, m, true)
}
func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIMEHeader, m map[string]string, allowReplication bool) error {
if v == nil {
bugLogIf(ctx, errInvalidArgument)
return errInvalidArgument
@@ -267,18 +257,14 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
nv[http.CanonicalHeaderKey(k)] = kv
}
// Save all supported headers.
// Save ordinary object metadata. Replication-only headers are restored only
// after the request has been validated as a trusted replication write.
for _, supportedHeader := range supportedHeaders {
value, ok := nv[http.CanonicalHeaderKey(supportedHeader)]
if ok {
if v, ok := replicationToInternalHeaders[supportedHeader]; ok {
if !allowReplication {
continue
}
m[v] = strings.Join(value, ",")
} else {
m[supportedHeader] = strings.Join(value, ",")
}
if _, ok := replicationToInternalHeaders[supportedHeader]; ok {
continue
}
if value, ok := nv[http.CanonicalHeaderKey(supportedHeader)]; ok {
m[supportedHeader] = strings.Join(value, ",")
}
}
@@ -287,8 +273,7 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
if !stringsHasPrefixFold(key, prefix) {
continue
}
value, ok := nv[http.CanonicalHeaderKey(key)]
if ok {
if value, ok := nv[http.CanonicalHeaderKey(key)]; ok {
m[key] = strings.Join(value, ",")
break
}
@@ -297,6 +282,31 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
return nil
}
// extractReplicationMetadataFromMime restores replication-only metadata after the
// caller has validated that the request is a trusted replication write.
func extractReplicationMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
if v == nil {
bugLogIf(ctx, errInvalidArgument)
return errInvalidArgument
}
nv := make(textproto.MIMEHeader, len(v))
for k, kv := range v {
// Canonicalize all headers, to remove any duplicates.
nv[http.CanonicalHeaderKey(k)] = kv
}
// Ordinary object metadata belongs to the caller. Re-extracting it would
// undo normalization (such as removing aws-chunked) or copy an outer
// Snowball archive's metadata onto its individual entries.
for header, internalHeader := range replicationToInternalHeaders {
if value, ok := nv[http.CanonicalHeaderKey(header)]; ok {
m[internalHeader] = strings.Join(value, ",")
}
}
return nil
}
// Returns access credentials in the request Authorization header.
func getReqAccessCred(r *http.Request, region string) (cred auth.Credentials) {
cred, _, _ = getReqAccessKeyV4(r, region, serviceS3)
+12 -1
View File
@@ -254,6 +254,9 @@ func TestExtractMetadataFromRequestKeepsQueryCompatibility(t *testing.T) {
func TestExtractReplicationMetadataHeaders(t *testing.T) {
header := http.Header{
"Content-Type": []string{"application/wasm"},
"Content-Encoding": []string{"aws-chunked"},
"X-Amz-Meta-Source": []string{"client"},
"X-Minio-Replication-Server-Side-Encryption-Sealed-Key": []string{"sealed-key"},
"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm": []string{"DAREv2-HMAC-SHA256"},
"X-Minio-Replication-Server-Side-Encryption-Iv": []string{"iv"},
@@ -262,12 +265,17 @@ func TestExtractReplicationMetadataHeaders(t *testing.T) {
ReplicationSsecChecksumHeader: []string{"checksum"},
}
metadata := make(map[string]string)
metadata := map[string]string{
"content-type": "application/wasm",
"x-amz-meta-source": "client",
}
if err := extractReplicationMetadataFromMime(t.Context(), textproto.MIMEHeader(header), metadata); err != nil {
t.Fatalf("failed to extract replication metadata: %v", err)
}
expected := map[string]string{
"content-type": "application/wasm",
"x-amz-meta-source": "client",
"X-Minio-Internal-Server-Side-Encryption-Sealed-Key": "sealed-key",
"X-Minio-Internal-Server-Side-Encryption-Seal-Algorithm": "DAREv2-HMAC-SHA256",
"X-Minio-Internal-Server-Side-Encryption-Iv": "iv",
@@ -279,6 +287,9 @@ func TestExtractReplicationMetadataHeaders(t *testing.T) {
if !reflect.DeepEqual(metadata, expected) {
t.Fatalf("unexpected replication metadata: expected %#v, got %#v", expected, metadata)
}
if _, ok := metadata["content-encoding"]; ok {
t.Fatalf("replication metadata restored transport content-encoding: %#v", metadata)
}
}
func TestGetCopyObjectMetadataFromHeaderReplication(t *testing.T) {
+111
View File
@@ -0,0 +1,111 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"errors"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
func TestIAMCredentialRetention(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
secret, err := getTokenSigningKey()
mustIAM(t, err)
parent := "external-idp-parent"
credential := func(exp time.Time) auth.Credentials {
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": exp.Unix(), parentClaim: parent}, secret)
mustIAM(t, err)
cred.ParentUser = parent
return cred
}
// Disablement of an external identity must include cached STS,
// which are kept separately from regular and service accounts.
cred := credential(UTCNow().Add(time.Hour))
_, err = sys.SetTempUser(ctx, cred.AccessKey, cred, "")
mustIAM(t, err)
mustIAM(t, sys.store.DeleteUsers(ctx, []string{parent}))
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(cred.AccessKey, stsUser))
mustIAM(t, err)
if !r.Deleted || !r.ExpiresAt.Equal(cred.Expiration.Add(globalMaxSkewTime)) || r.Credentials.SessionToken != "" || r.Credentials.SecretKey != "" {
t.Fatal("early STS revocation lost its retention boundary or retained a secret")
}
if _, ok := sys.store.GetUser(cred.AccessKey); ok {
t.Fatal("external disablement left the STS cache live")
}
_, err = sys.SetTempUser(withIAMReplicationTime(ctx, UTCNow().Add(time.Minute)), cred.AccessKey, cred, "")
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("same revoked token was reissued by replay: %v", err)
}
var mp MappedPolicy
err = sys.store.loadIAMConfig(ctx, &mp, getMappedPolicyPath(cred.AccessKey, stsUser, false))
if !errors.Is(err, errConfigNotFound) {
t.Fatalf("random STS key produced a permanent mapping: %v", err)
}
// Seed genuinely expired immutable tokens, as an ordinary startup
// loader sees them. Natural expiry leaves no permanent tombstone.
expired := credential(UTCNow().Add(-time.Hour))
path := getUserIdentityPath(expired.AccessKey, stsUser)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: expired, UpdatedAt: UTCNow().Add(-2 * time.Hour)}, path))
_ = sys.store.loadUser(ctx, expired.AccessKey, stsUser, make(map[string]UserIdentity))
var u UserIdentity
if err := sys.store.loadIAMConfig(ctx, &u, path); !errors.Is(err, errConfigNotFound) {
t.Fatalf("natural expiration retained a random key: %v", err)
}
// A retained early-revocation record is collectable only after the
// immutable token's expiration plus the skew allowance.
tomb := UserIdentity{Version: 1, Deleted: true, UpdatedAt: UTCNow().Add(-2 * time.Hour), ExpiresAt: expired.Expiration.Add(globalMaxSkewTime)}
mustIAM(t, sys.store.saveIAMConfig(ctx, &tomb, path))
_ = sys.store.loadUser(ctx, expired.AccessKey, stsUser, make(map[string]UserIdentity))
if err := sys.store.loadIAMConfig(ctx, &u, path); !errors.Is(err, errConfigNotFound) {
t.Fatalf("expired STS revocation not collected: %v", err)
}
if _, ok := sys.store.revisionIndex().snapshot()[path]; ok {
t.Fatal("expired STS retained an index entry")
}
})
}
}
func TestIAMPolicyDeletionRemainsExplicit(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
mustIAM(t, sys.DeletePolicy(ctx, "misspelled-policy", true))
r, err := loadIAMRevision(ctx, sys.store, getPolicyDocPath("misspelled-policy"))
mustIAM(t, err)
if r.Deleted {
t.Fatal("local nonexistent policy created a tombstone")
}
p, err := sys.store.GetPolicy("readwrite")
mustIAM(t, err)
if err := sys.DeletePolicy(ctx, "readwrite", true); err == nil {
t.Fatal("local pristine builtin policy became deletable")
}
_, err = sys.SetPolicy(ctx, "readwrite", p)
mustIAM(t, err)
mustIAM(t, sys.DeletePolicy(ctx, "readwrite", true))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
if _, err := sys.store.GetPolicy("readwrite"); !errors.Is(err, errNoSuchPolicy) {
t.Fatalf("reload restored an explicitly deleted override: %v", err)
}
_, err = sys.SetPolicy(ctx, "readwrite", p)
mustIAM(t, err)
if _, err := sys.store.GetPolicy("readwrite"); err != nil {
t.Fatal("explicit policy recreation failed", err)
}
mustIAM(t, globalSiteReplicationSys.PeerAddPolicyHandler(ctx, "remote-unknown-policy", nil, UTCNow()))
r, err = loadIAMRevision(ctx, sys.store, getPolicyDocPath("remote-unknown-policy"))
mustIAM(t, err)
if !r.Deleted {
t.Fatal("replicated unknown deletion lost its version")
}
})
}
}
+65 -71
View File
@@ -26,7 +26,6 @@ import (
"sync"
"time"
jsoniter "github.com/json-iterator/go"
"github.com/minio/minio-go/v7/pkg/set"
"github.com/minio/minio/internal/config"
"github.com/minio/minio/internal/kms"
@@ -62,6 +61,7 @@ type IAMEtcdStore struct {
sync.RWMutex
*iamCache
index iamRevisionIndex
usersSysType UsersSysType
@@ -69,13 +69,17 @@ type IAMEtcdStore struct {
}
func newIAMEtcdStore(client *etcd.Client, usersSysType UsersSysType) *IAMEtcdStore {
return &IAMEtcdStore{
store := &IAMEtcdStore{
iamCache: newIamCache(),
client: client,
usersSysType: usersSysType,
}
store.revisions = &store.index
return store
}
func (ies *IAMEtcdStore) revisionIndex() *iamRevisionIndex { return &ies.index }
func (ies *IAMEtcdStore) rlock() *iamCache {
ies.RLock()
return ies.iamCache
@@ -103,6 +107,7 @@ func (ies *IAMEtcdStore) saveIAMConfig(ctx context.Context, item any, itemPath s
if err != nil {
return err
}
plain := data
if GlobalKMS != nil {
data, err = config.EncryptBytes(GlobalKMS, data, kms.Context{
minioMetaBucket: path.Join(minioMetaBucket, itemPath),
@@ -111,24 +116,28 @@ func (ies *IAMEtcdStore) saveIAMConfig(ctx context.Context, item any, itemPath s
return err
}
}
return saveKeyEtcd(ctx, ies.client, itemPath, data, opts...)
if err := saveKeyEtcd(ctx, ies.client, itemPath, data, opts...); err != nil {
return err
}
ies.index.observe(itemPath, plain)
return nil
}
func getIAMConfig(item any, data []byte, itemPath string) error {
data, err := decryptData(data, itemPath)
func (ies *IAMEtcdStore) decodeIAMConfig(item any, data []byte, path string) error {
data, err := decryptData(data, path)
if err != nil {
return err
}
json := jsoniter.ConfigCompatibleWithStandardLibrary
ies.index.observe(path, data)
return json.Unmarshal(data, item)
}
func (ies *IAMEtcdStore) loadIAMConfig(ctx context.Context, item any, path string) error {
data, err := readKeyEtcd(ctx, ies.client, path)
data, err := ies.loadIAMConfigBytes(ctx, path)
if err != nil {
return err
}
return getIAMConfig(item, data, path)
return json.Unmarshal(data, item)
}
func (ies *IAMEtcdStore) loadIAMConfigBytes(ctx context.Context, path string) ([]byte, error) {
@@ -136,11 +145,19 @@ func (ies *IAMEtcdStore) loadIAMConfigBytes(ctx context.Context, path string) ([
if err != nil {
return nil, err
}
return decryptData(data, path)
data, err = decryptData(data, path)
if err == nil {
ies.index.observe(path, data)
}
return data, err
}
func (ies *IAMEtcdStore) deleteIAMConfig(ctx context.Context, path string) error {
return deleteKeyEtcd(ctx, ies.client, path)
if err := deleteKeyEtcd(ctx, ies.client, path); err != nil {
return err
}
ies.index.forget(path)
return nil
}
func (ies *IAMEtcdStore) loadPolicyDocWithRetry(ctx context.Context, policy string, m map[string]PolicyDoc, _ int) error {
@@ -162,6 +179,9 @@ func (ies *IAMEtcdStore) loadPolicyDoc(ctx context.Context, policy string, m map
return err
}
if p.Deleted {
return errNoSuchPolicy
}
m[policy] = p
return nil
}
@@ -181,7 +201,11 @@ func (ies *IAMEtcdStore) getPolicyDocKV(ctx context.Context, kvs *mvccpb.KeyValu
return err
}
ies.index.observe(string(kvs.Key), data)
policy := extractPathPrefixAndSuffix(string(kvs.Key), iamConfigPoliciesPrefix, path.Base(string(kvs.Key)))
if p.Deleted {
return errNoSuchPolicy
}
m[policy] = p
return nil
}
@@ -207,7 +231,7 @@ func (ies *IAMEtcdStore) loadPolicyDocs(ctx context.Context, m map[string]Policy
func (ies *IAMEtcdStore) getUserKV(ctx context.Context, userkv *mvccpb.KeyValue, userType IAMUserType, m map[string]UserIdentity, basePrefix string) error {
var u UserIdentity
err := getIAMConfig(&u, userkv.Value, string(userkv.Key))
err := ies.decodeIAMConfig(&u, userkv.Value, string(userkv.Key))
if err != nil {
if err == errConfigNotFound {
return errNoSuchUser
@@ -219,10 +243,14 @@ func (ies *IAMEtcdStore) getUserKV(ctx context.Context, userkv *mvccpb.KeyValue,
}
func (ies *IAMEtcdStore) addUser(ctx context.Context, user string, userType IAMUserType, u UserIdentity, m map[string]UserIdentity) error {
if u.Deleted {
if userType == stsUser && !u.ExpiresAt.IsZero() && UTCNow().After(u.ExpiresAt) {
bestEffortIAMExpiration(ctx, ies, getUserIdentityPath(user, userType))
}
return errNoSuchUser
}
if u.Credentials.IsExpired() {
// Delete expired identity.
deleteKeyEtcd(ctx, ies.client, getUserIdentityPath(user, userType))
deleteKeyEtcd(ctx, ies.client, getMappedPolicyPath(user, userType, false))
bestEffortIAMExpiration(ctx, ies, getUserIdentityPath(user, userType))
return nil
}
if u.Credentials.AccessKey == "" {
@@ -231,16 +259,17 @@ func (ies *IAMEtcdStore) addUser(ctx context.Context, user string, userType IAMU
if u.Credentials.SessionToken != "" {
jwtClaims, err := extractJWTClaims(u)
if err != nil {
if u.Credentials.IsTemp() {
// We should delete such that the client can re-request
// for the expiring credentials.
deleteKeyEtcd(ctx, ies.client, getUserIdentityPath(user, userType))
deleteKeyEtcd(ctx, ies.client, getMappedPolicyPath(user, userType, false))
}
// A temporarily unavailable signing key is not proof of expiration.
return nil
}
u.Credentials.Claims = jwtClaims.Map()
}
if err := checkIAMParentRevision(ctx, ies, u.Credentials); err != nil {
if errors.Is(err, errIAMStaleUpdate) {
return errNoSuchUser
}
return err
}
if u.Credentials.Description == "" {
u.Credentials.Description = u.Credentials.Comment
}
@@ -258,6 +287,9 @@ func (ies *IAMEtcdStore) loadSecretKey(ctx context.Context, user string, userTyp
}
return "", err
}
if u.Deleted {
return "", errNoSuchUser
}
return u.Credentials.SecretKey, nil
}
@@ -274,6 +306,7 @@ func (ies *IAMEtcdStore) loadUser(ctx context.Context, user string, userType IAM
}
func (ies *IAMEtcdStore) loadUsers(ctx context.Context, userType IAMUserType, m map[string]UserIdentity) error {
ctx = withIAMExpirationCleanup(ctx)
var basePrefix string
switch userType {
case svcUser:
@@ -312,6 +345,9 @@ func (ies *IAMEtcdStore) loadGroup(ctx context.Context, group string, m map[stri
}
return err
}
if gi.Deleted {
return errNoSuchGroup
}
m[group] = gi
return nil
}
@@ -349,13 +385,16 @@ func (ies *IAMEtcdStore) loadMappedPolicy(ctx context.Context, name string, user
}
return err
}
if !ies.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
func getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy], basePrefix string) error {
func (ies *IAMEtcdStore) getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy], basePrefix string) error {
var p MappedPolicy
err := getIAMConfig(&p, kv.Value, string(kv.Key))
err := ies.decodeIAMConfig(&p, kv.Value, string(kv.Key))
if err != nil {
if err == errConfigNotFound {
return errNoSuchPolicy
@@ -363,6 +402,9 @@ func getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy],
return err
}
name := extractPathPrefixAndSuffix(string(kv.Key), basePrefix, ".json")
if !ies.index.mappingAllowed(string(kv.Key), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
@@ -392,61 +434,13 @@ func (ies *IAMEtcdStore) loadMappedPolicies(ctx context.Context, userType IAMUse
// Parse all policies mapping to create the proper data model
for _, kv := range r.Kvs {
if err = getMappedPolicy(kv, m, basePrefix); err != nil && !errors.Is(err, errNoSuchPolicy) {
if err = ies.getMappedPolicy(kv, m, basePrefix); err != nil && !errors.Is(err, errNoSuchPolicy) {
return err
}
}
return nil
}
func (ies *IAMEtcdStore) savePolicyDoc(ctx context.Context, policyName string, p PolicyDoc) error {
return ies.saveIAMConfig(ctx, &p, getPolicyDocPath(policyName))
}
func (ies *IAMEtcdStore) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp MappedPolicy, opts ...options) error {
return ies.saveIAMConfig(ctx, mp, getMappedPolicyPath(name, userType, isGroup), opts...)
}
func (ies *IAMEtcdStore) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u UserIdentity, opts ...options) error {
return ies.saveIAMConfig(ctx, u, getUserIdentityPath(name, userType), opts...)
}
func (ies *IAMEtcdStore) saveGroupInfo(ctx context.Context, name string, gi GroupInfo) error {
return ies.saveIAMConfig(ctx, gi, getGroupInfoPath(name))
}
func (ies *IAMEtcdStore) deletePolicyDoc(ctx context.Context, name string) error {
err := ies.deleteIAMConfig(ctx, getPolicyDocPath(name))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (ies *IAMEtcdStore) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
err := ies.deleteIAMConfig(ctx, getMappedPolicyPath(name, userType, isGroup))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (ies *IAMEtcdStore) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
err := ies.deleteIAMConfig(ctx, getUserIdentityPath(name, userType))
if err == errConfigNotFound {
err = errNoSuchUser
}
return err
}
func (ies *IAMEtcdStore) deleteGroupInfo(ctx context.Context, name string) error {
err := ies.deleteIAMConfig(ctx, getGroupInfoPath(name))
if err == errConfigNotFound {
err = errNoSuchGroup
}
return err
}
func (ies *IAMEtcdStore) watch(ctx context.Context, keyPath string) <-chan iamWatchEvent {
ch := make(chan iamWatchEvent)
+164
View File
@@ -0,0 +1,164 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"maps"
"slices"
"time"
"github.com/minio/minio-go/v7/pkg/set"
)
type (
iamGroupGrantsKey struct{}
iamGroupMutationKey struct{}
iamGroupMutation struct {
Members []string
Remove bool
StatusOnly bool
}
)
// Merge the intended mutation with the record read under the distributed
// revision lock, not the older cache used to prepare the request.
func mergeIAMGroupMutation(ctx context.Context, previous GroupInfo, next *GroupInfo) {
op, ok := ctx.Value(iamGroupMutationKey{}).(iamGroupMutation)
if !ok || previous.Deleted || previous.Version == 0 {
return
}
members := set.CreateStringSet(previous.Members...)
grants := maps.Clone(previous.MemberGrants)
if grants == nil {
grants = make(map[string]time.Time)
}
switch {
case op.StatusOnly:
// Only the status changes.
case op.Remove:
for _, member := range op.Members {
members.Remove(member)
delete(grants, member)
}
next.Status = previous.Status
default:
requested := set.CreateStringSet(next.Members...)
for _, member := range op.Members {
if !requested.Contains(member) {
continue
}
at := next.MemberGrants[member]
if at.Before(grants[member]) {
continue
}
members.Add(member)
grants[member] = at
}
next.Status = previous.Status
}
next.Members, next.MemberGrants = members.ToSlice(), grants
slices.Sort(next.Members)
}
// A non-nil map is supplied by the versioned peer envelope, including for
// snapshots. Missing times are unknown, never the snapshot's newer timestamp.
func withIAMGroupGrants(ctx context.Context, grants map[string]time.Time) context.Context {
return context.WithValue(ctx, iamGroupGrantsKey{}, grants)
}
func (c *iamCache) effectiveGroupMembers(gi GroupInfo) []string {
var members []string
for _, member := range gi.Members {
if c.groupMemberAllowed(member, gi.MemberGrants[member], gi.RevokedBefore) {
members = append(members, member)
}
}
return members
}
func (c *iamCache) effectiveUserGroups(user string) []string {
var groups []string
for group := range c.iamUserGroupMemberships[user] {
gi, ok := c.iamGroupsMap[group]
r := c.revisions.get(getGroupInfoPath(group))
if r.RevokedBefore.After(gi.RevokedBefore) {
gi.RevokedBefore = r.RevokedBefore
}
if ok && !r.Deleted && c.groupMemberAllowed(user, gi.MemberGrants[user], gi.RevokedBefore) {
groups = append(groups, group)
}
}
return groups
}
func (c *iamCache) addGroupMembers(ctx context.Context, gi GroupInfo, members []string) (GroupInfo, error) {
grants, versioned := ctx.Value(iamGroupGrantsKey{}).(map[string]time.Time)
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(gi.RevokedBefore) {
gi.RevokedBefore = boundary
}
origin, replicated := iamReplicationTime(ctx)
gi.Members = slices.Clone(gi.Members)
gi.MemberGrants = maps.Clone(gi.MemberGrants)
if gi.MemberGrants == nil {
gi.MemberGrants = make(map[string]time.Time)
}
current := set.CreateStringSet(gi.Members...)
gi.UpdatedAt = UTCNow()
if replicated {
gi.UpdatedAt = origin
}
for _, member := range members {
at := gi.UpdatedAt
r := c.userRevocation(member)
if replicated {
switch {
case versioned:
at = grants[member]
if at.After(origin) {
return gi, errInvalidArgument
}
case !r.RevokedBefore.IsZero() || r.Deleted || !gi.RevokedBefore.IsZero():
// Legacy snapshots cannot prove a post-revocation grant.
continue
case current.Contains(member):
continue
}
if !c.groupMemberAllowed(member, at, gi.RevokedBefore) {
continue
}
} else {
if current.Contains(member) && c.groupMemberAllowed(member, gi.MemberGrants[member], gi.RevokedBefore) {
continue // Editing the group is not reissuing every grant.
}
if !at.After(gi.RevokedBefore) {
at = gi.RevokedBefore.Add(time.Nanosecond)
}
if !at.After(r.RevokedBefore) {
at = r.RevokedBefore.Add(time.Nanosecond)
}
if !at.After(gi.MemberGrants[member]) {
at = gi.MemberGrants[member].Add(time.Nanosecond)
}
if at.After(gi.UpdatedAt) {
gi.UpdatedAt = at
}
}
u, ok := c.iamUsersMap[member]
if !ok {
return gi, errNoSuchUser
}
if u.Credentials.IsTemp() || u.Credentials.IsServiceAccount() {
return gi, errIAMActionNotAllowed
}
if previous := gi.MemberGrants[member]; previous.After(at) {
continue
}
current.Add(member)
gi.MemberGrants[member] = at
}
gi.Members = current.ToSlice()
slices.Sort(gi.Members)
return gi, nil
}
+75
View File
@@ -0,0 +1,75 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"testing"
"time"
)
func TestIAMHealingResumesAfterLeadershipLoss(t *testing.T) {
previous := globalLeaderLock
locks := make(chan LockContext)
globalLeaderLock = &sharedLock{lockContext: locks}
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
c := &SiteReplicationSys{}
go func() { c.startHealRoutine(ctx, nil); close(done) }()
t.Cleanup(func() {
cancel()
select {
case <-done:
case locks <- LockContext{ctx: ctx}:
}
<-done
globalLeaderLock = previous
})
first, loseFirst := context.WithCancel(ctx)
defer loseFirst()
select {
case locks <- LockContext{ctx: first}:
case <-time.After(time.Second):
t.Fatal("healer did not acquire its first leader context")
}
loseFirst() // A transient quorum loss cancels the distributed lease.
select {
case <-done:
t.Fatal("healer permanently exited after temporary leadership loss")
case locks <- LockContext{ctx: ctx}:
case <-time.After(time.Second):
t.Fatal("healer did not wait for reacquired leadership")
}
// Shutdown must also interrupt the wait for leadership after lease loss.
cancel()
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("healer did not stop with its owning context")
}
}
func TestIAMHealingLeadershipWaitCancels(t *testing.T) {
previous := globalLeaderLock
locks := make(chan LockContext)
globalLeaderLock = &sharedLock{lockContext: locks}
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
c := &SiteReplicationSys{}
go func() { c.startHealRoutine(ctx, nil); close(done) }()
cancel()
t.Cleanup(func() {
select {
case <-done:
case locks <- LockContext{ctx: ctx}:
}
<-done
globalLeaderLock = previous
})
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("healer ignored shutdown while waiting for leadership")
}
}
+60
View File
@@ -0,0 +1,60 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
)
// Measures steady-state index traversal, sorting and the capability request.
// The network peer acknowledges real batches but performs no disk I/O; this
// benchmark deliberately does not claim durable catch-up throughput.
func BenchmarkIAMRevisionConvergedHealing(b *testing.B) {
for _, n := range []int{1000, 10000} {
b.Run(fmt.Sprint(n), func(b *testing.B) {
ctx, sys, _ := prepareIAMRevisionFixture(b)
_, err := sys.CreateUser(ctx, "benchmark-sync", madmin.AddOrUpdateUserReq{SecretKey: "valid-sync-password", Status: madmin.AccountEnabled})
mustIAM(b, err)
for i := range n {
at := UTCNow().Add(time.Duration(i) * time.Nanosecond)
data, err := json.Marshal(iamRevision{Deleted: true, UpdatedAt: at, RevokedBefore: at})
mustIAM(b, err)
sys.store.revisionIndex().observe(getUserIdentityPath(fmt.Sprintf("deleted-%06d", i), regUser), data)
}
var puts atomic.Int64
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
w.WriteHeader(http.StatusOK)
return
}
if r.Method == http.MethodPut {
puts.Add(1)
}
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: "benchmark-peer", Digest: "constant"}})
}))
defer server.Close()
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "benchmark-sync", Peers: map[string]madmin.PeerInfo{globalDeploymentID(): {DeploymentID: globalDeploymentID(), Name: "local"}, "remote": {DeploymentID: "remote", Name: "remote", Endpoint: server.URL}}}}
mustIAM(b, c.healIAMDeletions(ctx))
before := puts.Load()
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
mustIAM(b, c.healIAMDeletions(ctx))
}
b.StopTimer()
b.ReportMetric(float64(puts.Load()-before)/float64(b.N), "PUT/op")
if puts.Load() != before {
b.Fatal("steady-state healing replayed acknowledged records")
}
})
}
}
+56 -61
View File
@@ -45,6 +45,7 @@ type IAMObjectStore struct {
sync.RWMutex
*iamCache
index iamRevisionIndex
usersSysType UsersSysType
@@ -52,13 +53,17 @@ type IAMObjectStore struct {
}
func newIAMObjectStore(objAPI ObjectLayer, usersSysType UsersSysType) *IAMObjectStore {
return &IAMObjectStore{
store := &IAMObjectStore{
iamCache: newIamCache(),
objAPI: objAPI,
usersSysType: usersSysType,
}
store.revisions = &store.index
return store
}
func (iamOS *IAMObjectStore) revisionIndex() *iamRevisionIndex { return &iamOS.index }
func (iamOS *IAMObjectStore) rlock() *iamCache {
iamOS.RLock()
return iamOS.iamCache
@@ -87,6 +92,7 @@ func (iamOS *IAMObjectStore) saveIAMConfig(ctx context.Context, item any, objPat
if err != nil {
return err
}
plain := data
if GlobalKMS != nil {
data, err = config.EncryptBytes(GlobalKMS, data, kms.Context{
minioMetaBucket: path.Join(minioMetaBucket, objPath),
@@ -95,7 +101,11 @@ func (iamOS *IAMObjectStore) saveIAMConfig(ctx context.Context, item any, objPat
return err
}
}
return saveConfig(ctx, iamOS.objAPI, objPath, data)
if err := saveConfig(ctx, iamOS.objAPI, objPath, data); err != nil {
return err
}
iamOS.index.observe(objPath, plain)
return nil
}
func decryptData(data []byte, objPath string) ([]byte, error) {
@@ -133,6 +143,7 @@ func (iamOS *IAMObjectStore) loadIAMConfigBytesWithMetadata(ctx context.Context,
if err != nil {
return nil, meta, err
}
iamOS.index.observe(objPath, data)
return data, meta, nil
}
@@ -146,7 +157,11 @@ func (iamOS *IAMObjectStore) loadIAMConfig(ctx context.Context, item any, objPat
}
func (iamOS *IAMObjectStore) deleteIAMConfig(ctx context.Context, path string) error {
return deleteConfig(ctx, iamOS.objAPI, path)
if err := deleteConfig(ctx, iamOS.objAPI, path); err != nil {
return err
}
iamOS.index.forget(path)
return nil
}
func (iamOS *IAMObjectStore) loadPolicyDocWithRetry(ctx context.Context, policy string, m map[string]PolicyDoc, retries int) error {
@@ -171,6 +186,10 @@ func (iamOS *IAMObjectStore) loadPolicyDocWithRetry(ctx context.Context, policy
return err
}
if p.Deleted {
return errNoSuchPolicy
}
if p.Version == 0 {
// This means that policy was in the old version (without any
// timestamp info). We fetch the mod time of the file and save
@@ -200,6 +219,10 @@ func (iamOS *IAMObjectStore) loadPolicy(ctx context.Context, policy string) (Pol
return p, err
}
if p.Deleted {
return PolicyDoc{}, errNoSuchPolicy
}
if p.Version == 0 {
// This means that policy was in the old version (without any
// timestamp info). We fetch the mod time of the file and save
@@ -245,6 +268,9 @@ func (iamOS *IAMObjectStore) loadSecretKey(ctx context.Context, user string, use
}
return "", err
}
if u.Deleted {
return "", errNoSuchUser
}
return u.Credentials.SecretKey, nil
}
@@ -258,10 +284,15 @@ func (iamOS *IAMObjectStore) loadUserIdentity(ctx context.Context, user string,
return u, err
}
if u.Deleted {
if userType == stsUser && !u.ExpiresAt.IsZero() && UTCNow().After(u.ExpiresAt) {
bestEffortIAMExpiration(ctx, iamOS, getUserIdentityPath(user, userType))
}
return UserIdentity{}, errNoSuchUser
}
if u.Credentials.IsExpired() {
// Delete expired identity - ignoring errors here.
iamOS.deleteIAMConfig(ctx, getUserIdentityPath(user, userType))
iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(user, userType, false))
bestEffortIAMExpiration(ctx, iamOS, getUserIdentityPath(user, userType))
return u, errNoSuchUser
}
@@ -272,16 +303,18 @@ func (iamOS *IAMObjectStore) loadUserIdentity(ctx context.Context, user string,
if u.Credentials.SessionToken != "" {
jwtClaims, err := extractJWTClaims(u)
if err != nil {
if u.Credentials.IsTemp() {
// We should delete such that the client can re-request
// for the expiring credentials.
iamOS.deleteIAMConfig(ctx, getUserIdentityPath(user, userType))
iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(user, userType, false))
}
return u, errNoSuchUser
// During startup the site signing key may not be available yet.
// Reject this load without deleting a credential that has not expired.
return UserIdentity{}, errNoSuchUser
}
u.Credentials.Claims = jwtClaims.Map()
}
if err := checkIAMParentRevision(ctx, iamOS, u.Credentials); err != nil {
if errors.Is(err, errIAMStaleUpdate) {
return UserIdentity{}, errNoSuchUser
}
return UserIdentity{}, err
}
if u.Credentials.Description == "" {
u.Credentials.Description = u.Credentials.Comment
@@ -320,6 +353,7 @@ func (iamOS *IAMObjectStore) loadUser(ctx context.Context, user string, userType
}
func (iamOS *IAMObjectStore) loadUsers(ctx context.Context, userType IAMUserType, m map[string]UserIdentity) error {
ctx = withIAMExpirationCleanup(ctx)
var basePrefix string
switch userType {
case svcUser:
@@ -354,6 +388,9 @@ func (iamOS *IAMObjectStore) loadGroup(ctx context.Context, group string, m map[
}
return err
}
if g.Deleted {
return errNoSuchGroup
}
m[group] = g
return nil
}
@@ -391,6 +428,9 @@ func (iamOS *IAMObjectStore) loadMappedPolicyWithRetry(ctx context.Context, name
goto retry
}
if !iamOS.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
@@ -405,6 +445,9 @@ func (iamOS *IAMObjectStore) loadMappedPolicyInternal(ctx context.Context, name
}
return p, err
}
if !iamOS.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return MappedPolicy{}, errNoSuchPolicy
}
return p, nil
}
@@ -824,54 +867,6 @@ func (iamOS *IAMObjectStore) loadAllFromObjStore(ctx context.Context, cache *iam
return nil
}
func (iamOS *IAMObjectStore) savePolicyDoc(ctx context.Context, policyName string, p PolicyDoc) error {
return iamOS.saveIAMConfig(ctx, &p, getPolicyDocPath(policyName))
}
func (iamOS *IAMObjectStore) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp MappedPolicy, opts ...options) error {
return iamOS.saveIAMConfig(ctx, mp, getMappedPolicyPath(name, userType, isGroup), opts...)
}
func (iamOS *IAMObjectStore) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u UserIdentity, opts ...options) error {
return iamOS.saveIAMConfig(ctx, u, getUserIdentityPath(name, userType), opts...)
}
func (iamOS *IAMObjectStore) saveGroupInfo(ctx context.Context, name string, gi GroupInfo) error {
return iamOS.saveIAMConfig(ctx, gi, getGroupInfoPath(name))
}
func (iamOS *IAMObjectStore) deletePolicyDoc(ctx context.Context, name string) error {
err := iamOS.deleteIAMConfig(ctx, getPolicyDocPath(name))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (iamOS *IAMObjectStore) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
err := iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(name, userType, isGroup))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (iamOS *IAMObjectStore) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
err := iamOS.deleteIAMConfig(ctx, getUserIdentityPath(name, userType))
if err == errConfigNotFound {
err = errNoSuchUser
}
return err
}
func (iamOS *IAMObjectStore) deleteGroupInfo(ctx context.Context, name string) error {
err := iamOS.deleteIAMConfig(ctx, getGroupInfoPath(name))
if err == errConfigNotFound {
err = errNoSuchGroup
}
return err
}
// Lists objects in the minioMetaBucket at the given path prefix. All returned
// items have the pathPrefix removed from their names.
func listIAMConfigItems(ctx context.Context, objAPI ObjectLayer, pathPrefix string) <-chan itemOrErr[string] {
+133
View File
@@ -0,0 +1,133 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/grid"
"github.com/pgsty/silo-pkg/v3/policy"
)
// Two independent IAM caches share the same real object backend, as sibling
// nodes do. Deliver the actual peer handler only after the source committed.
func TestIAMPeerDeleteNotificationReloadsCommittedState(t *testing.T) {
for _, name := range []string{"deleted", "recreated", "recreated_without_grant"} {
recreate := name != "deleted"
t.Run(name, func(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
globalObjLayerMutex.Lock()
globalObjectAPI = obj
globalObjLayerMutex.Unlock()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
source := globalIAMSys
const user = "peer-reload-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "original-test-password", Status: madmin.AccountEnabled}
_, err = source.CreateUser(ctx, user, req)
must(err)
_, err = source.PolicyDBSet(ctx, user, "readwrite", regUser, false)
must(err)
_, err = source.AddUsersToGroup(ctx, "peer-reload-group", []string{user})
must(err)
_, err = source.PolicyDBSet(ctx, "peer-reload-group", "readwrite", regUser, true)
must(err)
svc, _, err := source.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{accessKey: "peer-reload-service", secretKey: "service-test-password"})
must(err)
signingKey, err := getTokenSigningKey()
must(err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: user}, signingKey)
must(err)
sts.ParentUser = user
_, err = source.SetTempUser(ctx, sts.AccessKey, sts, "")
must(err)
siblingStore := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
must(siblingStore.LoadIAMCache(ctx, true))
must(siblingStore.UserNotificationHandler(ctx, sts.AccessKey, stsUser))
for _, key := range []string{user, svc.AccessKey, sts.AccessKey} {
if _, ok := siblingStore.GetUser(key); !ok {
t.Fatalf("fixture did not load %s", key)
}
}
must(source.DeleteUser(ctx, user, false))
if recreate {
req.SecretKey = "recreated-test-password"
_, err = source.CreateUser(ctx, user, req)
must(err)
if name == "recreated" {
_, err = source.PolicyDBSet(ctx, user, "readonly", regUser, false)
must(err)
}
}
sibling := &IAMSys{store: siblingStore, usersSysType: MinIOUsersSysType}
globalIAMSys = sibling
defer func() { globalIAMSys = source }()
server := &peerRESTServer{}
for range 2 {
_, remoteErr := server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
}
if recreate {
u, ok := siblingStore.GetUser(user)
if !ok || u.Credentials.SecretKey != req.SecretKey {
t.Fatal("delayed deletion notification removed the recreated user")
}
loaded := make(map[string]UserIdentity)
must(source.store.loadUser(ctx, user, regUser, loaded))
if loaded[user].Credentials.SecretKey != req.SecretKey {
t.Fatal("notification changed the persisted recreated identity")
}
if allowed := sibling.IsAllowed(policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}); allowed != (name == "recreated") {
t.Fatal("notification did not load the recreated user's current grant")
}
}
if sibling.IsAllowed(policy.Args{AccountName: user, Action: policy.PutObjectAction, BucketName: "bucket", ObjectName: "object"}) {
t.Fatal("notification retained an old direct or group grant")
}
for _, key := range []string{svc.AccessKey, sts.AccessKey} {
if _, ok := siblingStore.GetUser(key); ok {
t.Fatalf("notification retained a revoked child: %s", key)
}
}
if !recreate {
for _, key := range []string{user} {
if _, ok := siblingStore.GetUser(key); ok {
t.Fatalf("notification retained a revoked cached identity: %s", key)
}
}
if sibling.IsAllowed(policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}) {
t.Fatal("notification retained the user's old grant")
}
cache := siblingStore.rlock()
member := cache.iamUserGroupMemberships[user].Contains("peer-reload-group")
siblingStore.runlock()
if member {
t.Fatal("notification retained the deleted user's group membership")
}
}
})
}
}
+131
View File
@@ -0,0 +1,131 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"fmt"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
)
// Uses APIs shared with the pre-revision tree so the same benchmark can be
// overlaid on that tree for a comparable local baseline.
func prepareIAMPerformanceFixture(b *testing.B) (context.Context, *IAMSys) {
b.Helper()
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
disks, err := getRandomDisks(1)
if err != nil {
b.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
b.Fatal(err)
}
initAllSubsystems(ctx)
globalIAMSys.initStore(obj, nil)
if err := globalIAMSys.Load(ctx, true); err != nil {
b.Fatal(err)
}
b.Cleanup(func() { cancel(); obj.Shutdown(context.Background()); os.RemoveAll(disks[0]); resetTestGlobals() })
return ctx, globalIAMSys
}
func BenchmarkIAMCachedCredential(b *testing.B) {
for _, kind := range []string{"user", "service", "sts"} {
b.Run(kind, func(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
const parent = "benchmark-parent"
_, err := sys.CreateUser(ctx, parent, madmin.AddOrUpdateUserReq{SecretKey: "benchmark-user-password", Status: madmin.AccountEnabled})
if err != nil {
b.Fatal(err)
}
key := parent
if kind == "service" {
c, _, err := sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "benchmark-service", secretKey: "benchmark-service-password"})
if err != nil {
b.Fatal(err)
}
key = c.AccessKey
}
if kind == "sts" {
secret, err := getTokenSigningKey()
if err != nil {
b.Fatal(err)
}
c, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
if err != nil {
b.Fatal(err)
}
c.ParentUser = parent
if _, err := sys.SetTempUser(ctx, c.AccessKey, c, ""); err != nil {
b.Fatal(err)
}
key = c.AccessKey
}
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
if _, ok := sys.store.GetUser(key); !ok {
b.Fatal("credential missing")
}
}
})
}
}
func BenchmarkIAMSetTempUser(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
const parent = "benchmark-sts-parent"
_, err := sys.CreateUser(ctx, parent, madmin.AddOrUpdateUserReq{SecretKey: "benchmark-user-password", Status: madmin.AccountEnabled})
if err != nil {
b.Fatal(err)
}
secret, err := getTokenSigningKey()
if err != nil {
b.Fatal(err)
}
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
if err != nil {
b.Fatal(err)
}
cred.ParentUser = parent
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
if _, err := sys.SetTempUser(ctx, cred.AccessKey, cred, "readwrite"); err != nil {
b.Fatal(err)
}
}
}
// Run with -benchtime=1x. Preparation is outside the timer; each measured load
// sees a fresh set of expired reusable service-account records.
func BenchmarkIAMColdLoadExpiredServices(b *testing.B) {
for _, count := range []int{100, 1000} {
b.Run(fmt.Sprint(count), func(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
b.ReportAllocs()
for i := 0; i < b.N; i++ {
b.StopTimer()
for j := 0; j < count; j++ {
key := fmt.Sprintf("expired-benchmark-%d-%d", i, j)
u := UserIdentity{Version: 1, UpdatedAt: UTCNow().Add(-2 * time.Hour), Credentials: auth.Credentials{AccessKey: key, SecretKey: "expired-benchmark-password", ParentUser: "absent-idp-parent", Expiration: UTCNow().Add(-time.Hour), Status: auth.AccountOn}}
if err := sys.store.saveIAMConfig(ctx, &u, getUserIdentityPath(key, svcUser)); err != nil {
b.Fatal(err)
}
}
b.StartTimer()
if err := sys.store.LoadIAMCache(ctx, true); err != nil {
b.Fatal(err)
}
}
})
}
}
+168
View File
@@ -0,0 +1,168 @@
package cmd
import (
"context"
"errors"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/pgsty/silo-pkg/v3/policy"
)
func TestReviewIAMRevokedUserReplay(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "review-revoked-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "review-valid-password", Status: madmin.AccountEnabled}
created, err := globalIAMSys.CreateUser(ctx, user, req)
if err != nil {
t.Fatal(err)
}
policyAt, err := globalIAMSys.PolicyDBSet(ctx, user, "readwrite", regUser, false)
if err != nil {
t.Fatal(err)
}
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "review-bucket", ObjectName: "review-object"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("seed must allow object read")
}
if err := globalIAMSys.DeleteUser(ctx, user, false); err != nil {
t.Fatal(err)
}
if err := globalIAMSys.store.LoadIAMCache(ctx, false); err != nil {
t.Fatal(err)
}
if globalIAMSys.IsAllowed(args) {
t.Fatal("deletion did not remove initial permission")
}
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, created); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Errorf("revoked user restored by an older replicated create, GetUserInfo error = %v", err)
}
if err := globalSiteReplicationSys.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: user, UserType: int(regUser), Policy: "readwrite"}, policyAt); err != nil {
t.Fatal(err)
}
if globalIAMSys.IsAllowed(args) {
t.Error("older replicated identity and policy events restored revoked S3 read permission")
}
}
func TestReviewIAMSourceTimestampOrder(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "review-ordered-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "review-valid-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour)
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin); err != nil {
t.Fatal(err)
}
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(time.Minute)); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Fatalf("newer source deletion skipped after delayed creation, GetUserInfo error = %v", err)
}
}
// A user's old group grant must not return after deletion and deliberate recreation.
func TestR3CandidateOldGroupReplayAfterRecreation(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user, group := "r3-group-member", "r3-granting-group"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-r3-user-password", Status: madmin.AccountEnabled}
peer := &globalSiteReplicationSys
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin))
add := &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}
must(peer.PeerGroupInfoChangeHandler(ctx, add, origin.Add(time.Minute)))
must(peer.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: group, IsGroup: true, UserType: int(regUser), Policy: "readwrite"}, origin.Add(time.Minute)))
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "r3-bucket", ObjectName: "probe"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("fixture must grant through group")
}
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin.Add(3*time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if globalIAMSys.IsAllowed(args) {
t.Fatal("recreation must start without deleted group membership")
}
must(peer.PeerGroupInfoChangeHandler(ctx, add, origin.Add(time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if globalIAMSys.IsAllowed(args) {
t.Fatal("old group event restored the deleted user's read grant after recreation and durable reload")
}
}
// The user delete is also a revocation of its earlier group memberships.
func TestR3CandidateLateDeleteRetainsOldGroupGrant(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user, group := "r3-late-member", "r3-late-group"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-r3-user-password", Status: madmin.AccountEnabled}
peer := &globalSiteReplicationSys
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin))
must(peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}, origin.Add(time.Minute)))
must(peer.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: group, IsGroup: true, UserType: int(regUser), Policy: "readwrite"}, origin.Add(time.Minute)))
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "r3-bucket", ObjectName: "probe"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("fixture must grant through group")
}
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin.Add(3*time.Minute)))
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if _, ok := globalIAMSys.GetUser(ctx, user); !ok {
t.Fatal("newer identity must survive")
}
if globalIAMSys.IsAllowed(args) {
t.Fatal("late user deletion retained the older group grant on the recreated identity")
}
}
+321
View File
@@ -0,0 +1,321 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"sort"
"strings"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
xhttp "github.com/minio/minio/internal/http"
"github.com/pgsty/silo-pkg/v3/policy"
)
const (
iamRevisionProtocol = 1
iamRevisionPeerPath = "/v3/site-replication/peer/iam-revisions"
iamUserBoundaryType = "silo-user-revocation"
iamGroupBoundaryType = "silo-group-revocation"
maxIAMRevisionBatch = 128
)
var iamRevisionInstance = mustGetUUID()
type iamUserBoundary struct {
User string `json:"user"`
Before time.Time `json:"before"`
}
type iamGroupBoundary struct {
Group string `json:"group"`
Before time.Time `json:"before"`
}
// The server owns this additive protocol, without changing the client SDK or
// overloading a policy/document field. Old servers reject the dedicated route
// before applying any change that would lose revocation or member metadata.
type iamReplicationItem struct {
madmin.SRIAMItem
GroupGrants map[string]time.Time `json:"groupGrants,omitempty"`
GroupSnapshot bool `json:"groupSnapshot,omitempty"`
UserRevocation *iamUserBoundary `json:"userRevocation,omitempty"`
GroupRevocation *iamGroupBoundary `json:"groupRevocation,omitempty"`
RevokedBefore time.Time `json:"revokedBefore,omitempty"`
}
type iamRevisionBatch struct {
Version int `json:"version"`
Items []iamReplicationItem `json:"items"`
}
type iamRevisionStatus struct {
Version int `json:"version"`
Node string `json:"node"`
Instance string `json:"instance"`
Digest string `json:"digest"`
}
type iamRevisionResponse struct {
iamRevisionStatus
Errors []string `json:"errors,omitempty"`
}
type iamRevisionBatchError struct{ failures []string }
func (e *iamRevisionBatchError) Error() string {
return "IAM revision batch: " + strings.Join(e.failures, "; ")
}
type iamRevisionProgress struct {
Instances map[string]string
Acknowledged map[string]string
}
type iamRevisionMetrics struct {
healFailures atomic.Uint64
healLastSuccess atomic.Int64
healDurationMillis atomic.Int64
}
func iamRevisionDigest(items map[string]iamRevision) string {
paths := make([]string, 0, len(items))
for path := range items {
paths = append(paths, path)
}
sort.Strings(paths)
h := sha256.New()
for _, path := range paths {
r := items[path]
fmt.Fprintf(h, "%q %s %t %s\n", path, r.timestamp().UTC().Format(time.RFC3339Nano), r.Deleted, r.RevokedBefore.UTC().Format(time.RFC3339Nano))
}
return hex.EncodeToString(h.Sum(nil))
}
func (store *IAMStoreSys) iamRevisionStatus() iamRevisionStatus {
node := globalLocalNodeName
if node == "" {
node = "local"
}
return iamRevisionStatus{Version: iamRevisionProtocol, Node: node, Instance: iamRevisionInstance, Digest: store.revisionIndex().digest()}
}
func executeIAMRevisionRequest(ctx context.Context, client *madmin.AdminClient, method string, batch *iamRevisionBatch) (status iamRevisionStatus, err error) {
var content []byte
if batch != nil {
content, err = json.Marshal(batch)
if err != nil {
return status, err
}
}
resp, err := client.ExecuteMethod(ctx, method, madmin.RequestData{RelPath: iamRevisionPeerPath, QueryValues: url.Values{"api-version": {madmin.SiteReplAPIVersion}}, Content: content})
if resp != nil {
defer xhttp.DrainBody(resp.Body)
}
if err != nil {
return status, err
}
if resp.StatusCode != http.StatusOK {
var remote madmin.ErrorResponse
if json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&remote) == nil && remote.Code != "" {
return status, remote
}
return status, fmt.Errorf("IAM revision protocol requires upgraded peers: %s", resp.Status)
}
var response iamRevisionResponse
if err = json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&response); err != nil {
return status, err
}
status = response.iamRevisionStatus
if status.Version != iamRevisionProtocol || status.Node == "" || status.Instance == "" || status.Digest == "" {
return status, errors.New("peer did not acknowledge the IAM revision protocol")
}
if len(response.Errors) != 0 {
return status, &iamRevisionBatchError{failures: response.Errors}
}
return status, nil
}
type (
iamRecordBoundaryKey struct{}
iamGroupSnapshotKey struct{}
)
func (c *SiteReplicationSys) replicationItem(ctx context.Context, item madmin.SRIAMItem) (iamReplicationItem, error) {
out := iamReplicationItem{SRIAMItem: item}
if item.Type == madmin.SRIAMItemSvcAcc && item.SvcAccChange != nil {
var key string
if item.SvcAccChange.Create != nil {
key = item.SvcAccChange.Create.AccessKey
} else if item.SvcAccChange.Update != nil {
key = item.SvcAccChange.Update.AccessKey
}
if key != "" {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getUserIdentityPath(key, svcUser))
if err != nil {
return out, err
}
if r.Deleted || r.timestamp().After(item.UpdatedAt) {
return out, errIAMStaleUpdate
}
out.RevokedBefore = r.RevokedBefore
}
}
if item.Type == madmin.SRIAMItemGroupInfo && item.GroupInfo != nil && !item.GroupInfo.UpdateReq.IsRemove {
out.GroupSnapshot = true
var gi GroupInfo
if err := globalIAMSys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath(item.GroupInfo.UpdateReq.Group)); err != nil {
return out, err
}
// The matching persisted snapshot carries member grant times. If a
// later write won before sending, propagate that whole newer state.
if gi.Deleted {
return out, errIAMStaleUpdate
}
out.UpdatedAt = gi.UpdatedAt
out.RevokedBefore = gi.RevokedBefore
out.GroupInfo = &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: item.GroupInfo.UpdateReq.Group, Status: madmin.GroupStatus(gi.Status)}}
cache := globalIAMSys.store.rlock()
out.GroupInfo.UpdateReq.Members = cache.effectiveGroupMembers(gi)
globalIAMSys.store.runlock()
out.GroupGrants = make(map[string]time.Time, len(out.GroupInfo.UpdateReq.Members))
for _, member := range out.GroupInfo.UpdateReq.Members {
out.GroupGrants[member] = gi.MemberGrants[member]
}
}
if item.Type == madmin.SRIAMItemGroupInfo && item.GroupInfo != nil && item.GroupInfo.UpdateReq.IsRemove && len(item.GroupInfo.UpdateReq.Members) == 0 {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getGroupInfoPath(item.GroupInfo.UpdateReq.Group))
if err != nil {
return out, err
}
if !r.Deleted && !r.RevokedBefore.IsZero() {
out.Type, out.GroupInfo = iamGroupBoundaryType, nil
out.GroupRevocation = &iamGroupBoundary{Group: item.GroupInfo.UpdateReq.Group, Before: r.RevokedBefore}
out.UpdatedAt = r.RevokedBefore
}
}
if item.Type == madmin.SRIAMItemIAMUser && item.IAMUser != nil {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getUserIdentityPath(item.IAMUser.AccessKey, regUser))
if err != nil {
return out, err
}
if item.IAMUser.IsDeleteReq && !r.Deleted && !r.RevokedBefore.IsZero() {
out.Type = iamUserBoundaryType
out.IAMUser = nil
out.UserRevocation = &iamUserBoundary{User: item.IAMUser.AccessKey, Before: r.RevokedBefore}
out.UpdatedAt = r.RevokedBefore
} else if !item.IAMUser.IsDeleteReq {
if r.Deleted || r.timestamp().After(item.UpdatedAt) {
return out, errIAMStaleUpdate
}
out.RevokedBefore = r.RevokedBefore
}
}
return out, nil
}
func applyIAMReplicationItem(ctx context.Context, item iamReplicationItem) error {
if item.GroupInfo != nil {
if item.GroupSnapshot {
ctx = context.WithValue(ctx, iamGroupSnapshotKey{}, true)
}
// A nil map also explicitly denotes unknown legacy grants. Do not
// turn an unrelated group edit into a new grant after a revocation.
ctx = withIAMGroupGrants(ctx, item.GroupGrants)
}
if !item.RevokedBefore.IsZero() {
if item.RevokedBefore.After(item.UpdatedAt) {
return errSRInvalidRequest(errInvalidArgument)
}
ctx = context.WithValue(ctx, iamRecordBoundaryKey{}, item.RevokedBefore)
}
switch item.Type {
case iamUserBoundaryType:
if item.UserRevocation == nil || item.UserRevocation.User == "" || item.UserRevocation.Before.IsZero() {
return errSRInvalidRequest(errInvalidArgument)
}
return iamReplicationError(globalIAMSys.DeleteUser(withIAMReplicationTime(ctx, item.UserRevocation.Before), item.UserRevocation.User, true))
case iamGroupBoundaryType:
if item.GroupRevocation == nil || item.GroupRevocation.Group == "" || item.GroupRevocation.Before.IsZero() {
return errSRInvalidRequest(errInvalidArgument)
}
_, err := globalIAMSys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, item.GroupRevocation.Before), item.GroupRevocation.Group, nil)
return iamReplicationError(err)
case madmin.SRIAMItemPolicy:
if len(item.Policy) == 0 {
return globalSiteReplicationSys.PeerAddPolicyHandler(ctx, item.Name, nil, item.UpdatedAt)
}
p, err := policy.ParseConfig(bytes.NewReader(item.Policy))
if err != nil {
return err
}
if p.IsEmpty() {
p = nil
}
return globalSiteReplicationSys.PeerAddPolicyHandler(ctx, item.Name, p, item.UpdatedAt)
case madmin.SRIAMItemSvcAcc:
return globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, item.SvcAccChange, item.UpdatedAt)
case madmin.SRIAMItemPolicyMapping:
return globalSiteReplicationSys.PeerPolicyMappingHandler(ctx, item.PolicyMapping, item.UpdatedAt)
case madmin.SRIAMItemSTSAcc:
return globalSiteReplicationSys.PeerSTSAccHandler(ctx, item.STSCredential, item.UpdatedAt)
case madmin.SRIAMItemIAMUser:
return globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, item.IAMUser, item.UpdatedAt)
case madmin.SRIAMItemGroupInfo:
return globalSiteReplicationSys.PeerGroupInfoChangeHandler(ctx, item.GroupInfo, item.UpdatedAt)
default:
return errSRInvalidRequest(errInvalidArgument)
}
}
func (a adminAPIHandlers) SRPeerIAMRevisions(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
if obj, _ := validateAdminReq(ctx, w, r, policy.SiteReplicationOperationAction); obj == nil {
return
}
var failures []string
if r.Method == http.MethodPut {
var batch iamRevisionBatch
if err := parseJSONBody(ctx, r.Body, &batch, ""); err != nil {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, err), r.URL)
return
}
if batch.Version != iamRevisionProtocol || len(batch.Items) == 0 || len(batch.Items) > maxIAMRevisionBatch {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, errSRInvalidRequest(errInvalidArgument)), r.URL)
return
}
for i, item := range batch.Items {
if err := applyIAMReplicationItem(ctx, item); err != nil {
failures = append(failures, fmt.Sprintf("item %d (%s): %v", i, item.Type, err))
}
}
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: globalIAMSys.store.iamRevisionStatus(), Errors: failures})
}
// A site endpoint can balance requests across nodes sharing durable IAM state.
// Switching between known node incarnations preserves ACKs; a new incarnation
// conservatively invalidates them so restoring an old backend cannot inherit
// acknowledgements from before the restore.
func (p *iamRevisionProgress) observePeer(status iamRevisionStatus) {
if p.Instances == nil {
p.Instances = make(map[string]string)
}
if p.Instances[status.Node] != status.Instance || p.Acknowledged == nil {
p.Instances[status.Node] = status.Instance
p.Acknowledged = make(map[string]string)
}
}
+161
View File
@@ -0,0 +1,161 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
)
func TestIAMRevisionProtocolDoesNotFallBackToLegacy(t *testing.T) {
var requests atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requests.Add(1)
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" {
t.Errorf("unsafe fallback path: %s", r.URL.Path)
}
w.WriteHeader(http.StatusNotFound)
_, _ = w.Write([]byte(`{"Code":"NotImplemented","Message":"old server"}`))
}))
defer server.Close()
client, err := madmin.New(strings.TrimPrefix(server.URL, "http://"), "test-access", "valid-test-secret", false)
mustIAM(t, err)
_, err = executeIAMRevisionRequest(context.Background(), client, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: []iamReplicationItem{{SRIAMItem: madmin.SRIAMItem{Type: iamUserBoundaryType}, UserRevocation: &iamUserBoundary{User: "recreated", Before: UTCNow()}}}})
if err == nil || requests.Load() != 1 {
t.Fatalf("old peer must reject without fallback, err=%v requests=%d", err, requests.Load())
}
}
type iamNoHealingScanStore struct{ IAMStorageAPI }
func (s *iamNoHealingScanStore) listIAMConfigPaths(context.Context) ([]string, error) {
panic("healing must use the loaded revision index")
}
func TestIAMRevisionHealingAcknowledgements(t *testing.T) {
for _, balanced := range []bool{false, true} {
t.Run(fmt.Sprintf("load_balanced_%t", balanced), func(t *testing.T) { testIAMRevisionHealingAcknowledgements(t, balanced) })
}
}
func testIAMRevisionHealingAcknowledgements(t *testing.T, balanced bool) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "ack-sync", madmin.AddOrUpdateUserReq{SecretKey: "valid-sync-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
for i := range maxIAMRevisionBatch*2 + 1 {
at := UTCNow().Add(time.Duration(i) * time.Nanosecond)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Deleted: true, UpdatedAt: at, RevokedBefore: at}, getUserIdentityPath(fmt.Sprintf("ack-%04d", i), regUser)))
}
sys.store.IAMStorageAPI = &iamNoHealingScanStore{IAMStorageAPI: sys.store.IAMStorageAPI}
var mu sync.Mutex
var puts, gets int
var applied int
instance := "boot-1"
failSecondBatch := true
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
w.WriteHeader(http.StatusOK)
return
}
mu.Lock()
defer mu.Unlock()
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" {
t.Errorf("unexpected request: %s", r.URL.Path)
w.WriteHeader(404)
return
}
var failures []string
if r.Method == http.MethodGet {
gets++
} else {
puts++
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(400)
return
}
if len(batch.Items) > maxIAMRevisionBatch {
t.Error("batch exceeds limit")
}
if failSecondBatch && puts == 2 {
failures = []string{"injected item error"}
} else {
applied += len(batch.Items)
}
}
node := "node-1"
if balanced {
node = fmt.Sprintf("node-%d", (gets+puts)%2+1)
}
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: node, Instance: node + instance, Digest: fmt.Sprintf("%d", applied)}, Errors: failures})
}))
defer server.Close()
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "ack-sync", Peers: map[string]madmin.PeerInfo{globalDeploymentID(): {DeploymentID: globalDeploymentID(), Name: "local"}, "remote": {DeploymentID: "remote", Name: "remote", Endpoint: server.URL}}}}
if err := c.healIAMDeletions(ctx); err == nil {
t.Fatal("item failure was hidden")
}
mu.Lock()
if puts != 3 || applied != maxIAMRevisionBatch+1 {
t.Errorf("failed middle batch blocked later revocations: puts=%d applied=%d", puts, applied)
}
failSecondBatch = false
applied++ // Unrelated remote mutation changes its digest.
mu.Unlock()
at := UTCNow()
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Deleted: true, UpdatedAt: at, RevokedBefore: at}, getUserIdentityPath("ack-new-local", regUser)))
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
if puts != 5 {
t.Errorf("did not resume at unacknowledged batch: puts=%d", puts)
}
mu.Unlock()
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
if puts != 5 || gets != 3 {
t.Errorf("converged records were replayed: puts=%d gets=%d", puts, gets)
}
instance = "boot-2"
mu.Unlock()
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
defer mu.Unlock()
if puts != 8 {
t.Fatalf("peer restart reused an old acknowledgement: puts=%d", puts)
}
}
func TestIAMRevisionIndexRebuildsFromStorage(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
const user = "index-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-parent-password", Status: madmin.AccountEnabled}
_, err := sys.CreateUser(ctx, user, req)
mustIAM(t, err)
mustIAM(t, sys.DeleteUser(ctx, user, false))
before := sys.store.revisionIndex().snapshot()
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, store.LoadIAMCache(ctx, true))
if iamRevisionDigest(before) != iamRevisionDigest(store.revisionIndex().snapshot()) {
t.Fatal("ordinary IAM loading did not restore the deletion index")
}
_, err = store.AddUser(ctx, user, req)
mustIAM(t, err)
r := store.revisionIndex().get(getUserIdentityPath(user, regUser))
if r.Deleted || r.RevokedBefore.IsZero() {
t.Fatal("recreation discarded the retained boundary")
}
if r.Credentials.SecretKey != "" || r.Credentials.SessionToken != "" {
t.Fatal("index retained credentials")
}
}
+393
View File
@@ -0,0 +1,393 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"errors"
"fmt"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
"github.com/pgsty/silo-pkg/v3/policy"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/namespace"
)
func prepareIAMRevisionFixture(t testing.TB, backend ...string) (context.Context, *IAMSys, ObjectLayer) {
t.Helper()
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
disks, err := getRandomDisks(1)
mustIAM(t, err)
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
mustIAM(t, err)
initAllSubsystems(ctx)
// Deliberately omit the periodic refresh goroutine. Fault injection can
// replace this fixture's storage interface without racing initialization.
var client *etcd.Client
if len(backend) != 0 && backend[0] == "etcd" {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if endpoint == "" {
cancel()
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
client, err = etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: 5 * time.Second})
mustIAM(t, err)
prefix := fmt.Sprintf("/silo-boundary-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(client.KV, prefix)
client.Watcher = namespace.NewWatcher(client.Watcher, prefix)
t.Cleanup(func() { client.Delete(context.Background(), "", etcd.WithPrefix()); client.Close() })
}
globalIAMSys.initStore(obj, client)
mustIAM(t, globalIAMSys.Load(ctx, true))
t.Cleanup(func() { cancel(); obj.Shutdown(context.Background()); os.RemoveAll(disks[0]); resetTestGlobals() })
return ctx, globalIAMSys, obj
}
func mustIAM(t testing.TB, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
var errIAMInjectedWrite = errors.New("injected IAM persistence failure")
type iamFailingCleanupStore struct {
IAMStorageAPI
parentPath string
beforeCommit bool
}
func (s *iamFailingCleanupStore) saveIAMConfig(ctx context.Context, item any, path string, opts ...options) error {
if s.beforeCommit || path != s.parentPath {
return errIAMInjectedWrite
}
return s.IAMStorageAPI.saveIAMConfig(ctx, item, path, opts...)
}
func TestIAMRevocationCommitBoundary(t *testing.T) {
for _, before := range []bool{true, false} {
name := "after_identity_commit"
if before {
name = "before_identity_commit"
}
t.Run(name, func(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
const user = "commit-boundary-user"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, req)
mustIAM(t, err)
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin.Add(time.Minute)), user, "readwrite", regUser, false)
mustIAM(t, err)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(time.Minute)), "commit-group", []string{user})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "commit-group", "readwrite", regUser, true)
mustIAM(t, err)
child, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), user, nil, newServiceAccountOpts{accessKey: "commit-child", secretKey: "valid-child-password"})
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("fixture has no grant")
}
siblingStore := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, siblingStore.LoadIAMCache(ctx, true))
sibling := &IAMSys{store: siblingStore, usersSysType: MinIOUsersSysType}
tg, err := grid.SetupTestGrid(2)
mustIAM(t, err)
defer tg.Cleanup()
var notifications atomic.Int32
mustIAM(t, deleteUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
notifications.Add(1)
if err := sibling.LoadUserAfterDelete(ctx, r.Get(peerRESTUser)); err != nil {
return grid.NoPayload{}, grid.NewRemoteErr(err)
}
return grid.NoPayload{}, nil
}))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
mustIAM(t, err)
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{host: host, gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) }}}}
original := sys.store.IAMStorageAPI
sys.store.IAMStorageAPI = &iamFailingCleanupStore{IAMStorageAPI: original, parentPath: getUserIdentityPath(user, regUser), beforeCommit: before}
boundary := origin.Add(2 * time.Minute)
err = sys.DeleteUser(withIAMReplicationTime(ctx, boundary), user, true)
if !errors.Is(err, errIAMInjectedWrite) {
t.Fatalf("expected write failure, got %v", err)
}
sys.store.IAMStorageAPI = original
r, err := loadIAMRevision(ctx, original, getUserIdentityPath(user, regUser))
mustIAM(t, err)
if before {
if r.Deleted || !sys.IsAllowed(args) || !sibling.IsAllowed(args) || notifications.Load() != 0 {
t.Fatal("failure before commit changed the identity or grant")
}
return
}
if !r.Deleted || !r.RevokedBefore.Equal(boundary) {
t.Fatal("cleanup failure lost durable revocation")
}
if sys.IsAllowed(args) || sibling.IsAllowed(args) || notifications.Load() != 1 {
t.Fatal("cleanup failure retained old permission")
}
// Subsequent fixture writes need no additional RPC handlers.
globalNotificationSys = &NotificationSys{}
// Recreate after the partial cleanup. The old mapping, group member
// and child still exist in storage; none may authorize this identity.
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), user, req)
mustIAM(t, err)
reloaded := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, reloaded.LoadIAMCache(ctx, true))
fresh := &IAMSys{store: reloaded, usersSysType: MinIOUsersSysType}
if fresh.IsAllowed(args) {
t.Fatal("cold reload restored partially cleaned-up grants")
}
if _, ok := reloaded.GetUser(child.AccessKey); ok {
t.Fatal("cold reload restored the old child")
}
gd, err := reloaded.GetGroupDescription("commit-group")
mustIAM(t, err)
if len(gd.Members) != 0 {
t.Fatalf("listing exposed a revoked group relation: %v", gd.Members)
}
_, err = sys.AddUsersToGroup(ctx, "commit-group", []string{user})
mustIAM(t, err)
if !sys.IsAllowed(args) {
t.Fatal("explicit new group grant was not accepted")
}
})
}
}
func TestIAMGroupGrantVersionsSurviveSnapshotsAndRecreation(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
for _, user := range []string{"grant-alice", "grant-bob"} {
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, req)
mustIAM(t, err)
}
grant := origin.Add(time.Minute)
_, err := sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), "grant-group", []string{"grant-alice"})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "grant-group", "readwrite", regUser, true)
mustIAM(t, err)
boundary := origin.Add(2 * time.Minute)
mustIAM(t, sys.DeleteUser(withIAMReplicationTime(ctx, boundary), "grant-alice", false))
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), "grant-alice", req)
mustIAM(t, err)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(4*time.Minute)), "grant-group", []string{"grant-bob"})
mustIAM(t, err)
_, err = sys.SetGroupStatus(withIAMReplicationTime(ctx, origin.Add(5*time.Minute)), "grant-group", true)
mustIAM(t, err)
var gi GroupInfo
mustIAM(t, sys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath("grant-group")))
if !gi.MemberGrants["grant-alice"].Equal(grant) {
t.Fatal("unrelated group edits refreshed an old grant")
}
args := policy.Args{AccountName: "grant-alice", Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
for _, stale := range []time.Time{grant, boundary, {}} {
item := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, UpdatedAt: origin.Add(6 * time.Minute), GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: "grant-group", Members: []string{"grant-alice", "grant-bob"}}}}, GroupSnapshot: true, GroupGrants: map[string]time.Time{"grant-alice": stale, "grant-bob": origin.Add(4 * time.Minute)}}
mustIAM(t, applyIAMReplicationItem(ctx, item))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
if sys.IsAllowed(args) {
t.Fatalf("snapshot restored revoked grant %s", stale)
}
gd, err := sys.GetGroupDescription("grant-group")
mustIAM(t, err)
if len(gd.Members) != 1 || gd.Members[0] != "grant-bob" {
t.Fatalf("inconsistent effective members: %v", gd.Members)
}
}
// Only an explicit post-revocation grant restores access.
freshAt, err := sys.AddUsersToGroup(ctx, "grant-group", []string{"grant-alice"})
mustIAM(t, err)
if !sys.IsAllowed(args) {
t.Fatal("explicit regrant rejected")
}
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
mustIAM(t, sys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath("grant-group")))
if !gi.MemberGrants["grant-alice"].Equal(freshAt) {
t.Fatal("new grant version was not persisted")
}
if !gi.MemberGrants["grant-bob"].Equal(origin.Add(4 * time.Minute)) {
t.Fatal("regranting Alice changed Bob's grant")
}
}
func TestIAMGroupRevocationCommitAndRecreation(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) { testIAMGroupRevocationCommitAndRecreation(t, backend) })
}
}
func testIAMGroupRevocationCommitAndRecreation(t *testing.T, backend string) {
ctx, sys, obj := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
user, group := "group-boundary-user", "group-boundary"
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
grant, boundary := origin.Add(time.Minute), origin.Add(2*time.Minute)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), group, []string{user})
mustIAM(t, err)
// A newer mapping must not veto the authoritative group deletion.
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), group, "readwrite", regUser, true)
mustIAM(t, err)
_, err = sys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, boundary), group, nil)
mustIAM(t, err)
r, err := loadIAMRevision(ctx, sys.store, getGroupInfoPath(group))
mustIAM(t, err)
if !r.Deleted || !r.RevokedBefore.Equal(boundary) {
t.Fatal("newer mapping swallowed group deletion")
}
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(4*time.Minute)), group, nil)
mustIAM(t, err)
for _, at := range []time.Time{grant, boundary, {}} {
item := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, UpdatedAt: origin.Add(5 * time.Minute), GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}}, GroupSnapshot: true, GroupGrants: map[string]time.Time{user: at}}
mustIAM(t, applyIAMReplicationItem(ctx, item))
gd, err := sys.GetGroupDescription(group)
mustIAM(t, err)
if len(gd.Members) != 0 {
t.Fatalf("group recreation restored grant %s", at)
}
}
_, err = sys.AddUsersToGroup(ctx, group, []string{user})
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("explicit group regrant was rejected")
}
// The newer live snapshot may arrive before an older group deletion.
lateBoundary := origin.Add(6 * time.Minute)
_, err = sys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, lateBoundary), group, nil)
mustIAM(t, err)
r, err = loadIAMRevision(ctx, sys.store, getGroupInfoPath(group))
mustIAM(t, err)
if r.Deleted || !r.RevokedBefore.Equal(lateBoundary) {
t.Fatal("late deletion lost the live group's revocation boundary")
}
// The old mapping is now revoked; a new explicit mapping restores access.
if sys.IsAllowed(args) {
t.Fatal("late group boundary retained an old mapping")
}
_, err = sys.PolicyDBSet(ctx, group, "readwrite", regUser, true)
mustIAM(t, err)
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
if es, ok := sys.store.IAMStorageAPI.(*IAMEtcdStore); ok {
store.IAMStorageAPI = newIAMEtcdStore(es.client, MinIOUsersSysType)
}
mustIAM(t, store.LoadIAMCache(ctx, true))
fresh := &IAMSys{store: store, usersSysType: MinIOUsersSysType}
if !fresh.IsAllowed(args) {
t.Fatal("reload lost explicit grants after a retained group boundary")
}
item, err := globalSiteReplicationSys.replicationItem(ctx, madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group}}, UpdatedAt: r.timestamp()})
mustIAM(t, err)
if !item.RevokedBefore.Equal(lateBoundary) || !item.GroupGrants[user].After(lateBoundary) {
t.Fatal("group snapshot lost revision metadata")
}
}
// A committed revision is observable before all cached dependents have been
// cleaned up. Every authorization read must apply that boundary in this window.
func TestIAMCachedMappingHonorsCommittedRevision(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
parent := "cached-external-parent"
_, err := sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), parent, "readwrite", stsUser, false)
mustIAM(t, err)
policies, err := sys.PolicyDBGet(parent)
mustIAM(t, err)
if len(policies) == 0 {
t.Fatal("fixture has no STS-parent mapping")
}
mustIAM(t, sys.store.saveIAMConfig(ctx, &MappedPolicy{Version: 1, Deleted: true, UpdatedAt: origin.Add(time.Minute)}, getMappedPolicyPath(parent, stsUser, false)))
policies, err = sys.PolicyDBGet(parent)
mustIAM(t, err)
if len(policies) != 0 {
t.Fatal("cached STS mapping ignored its own namespace tombstone")
}
user, group := "cached-group-user", "cached-group"
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin), user, madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
grant := origin.Add(5 * time.Minute)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), group, []string{user})
mustIAM(t, err)
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), group, "readwrite", regUser, true)
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("fixture has no group grant")
}
// A late deletion preserves the newer member grant but revokes the older
// policy mapping. Simulate the interval before mapping cleanup completes.
gi := GroupInfo{Version: 1, Status: statusEnabled, Members: []string{user}, MemberGrants: map[string]time.Time{user: grant}, UpdatedAt: grant, RevokedBefore: origin.Add(2 * time.Minute)}
mustIAM(t, sys.store.saveIAMConfig(ctx, &gi, getGroupInfoPath(group)))
if sys.IsAllowed(args) {
t.Fatal("cached group mapping ignored the committed group boundary")
}
gd, err := sys.GetGroupDescription(group)
mustIAM(t, err)
if gd.Policy != "" {
t.Fatal("group listing exposed a revoked mapping")
}
}
type (
iamExpiryLockFailure struct {
ObjectLayer
path string
}
iamFailedExpiryLock struct{ RWLocker }
)
func (o *iamExpiryLockFailure) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket == minioMetaBucket && len(objects) == 1 && objects[0] == o.path+".revision-lock" {
return &iamFailedExpiryLock{RWLocker: lock}
}
return lock
}
func (l *iamFailedExpiryLock) GetLock(context.Context, *dynamicTimeout) (LockContext, error) {
return LockContext{}, errIAMInjectedWrite
}
func TestIAMExpiredCredentialCleanupDoesNotBlockLoading(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "healthy-user", madmin.AddOrUpdateUserReq{SecretKey: "healthy-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "healthy-user", "readwrite", regUser, false)
mustIAM(t, err)
c, _, err := sys.NewServiceAccount(ctx, "healthy-user", nil, newServiceAccountOpts{accessKey: "expired-service", secretKey: "expired-service-password"})
mustIAM(t, err)
c.Expiration = UTCNow().Add(-time.Hour)
path := getUserIdentityPath(c.AccessKey, svcUser)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: c, UpdatedAt: UTCNow()}, path))
// A cold loader sees the existing version but cannot acquire the cleanup
// write lock. Healthy users must still load; the expired one stays denied.
fresh := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(&iamExpiryLockFailure{ObjectLayer: obj, path: path}, MinIOUsersSysType)}
mustIAM(t, fresh.LoadIAMCache(ctx, true))
if _, ok := fresh.GetUser("healthy-user"); !ok {
t.Fatal("cleanup failure prevented healthy IAM state from loading")
}
if _, ok := fresh.GetUser(c.AccessKey); ok {
t.Fatal("cleanup failure admitted an expired service account")
}
r, err := loadIAMRevision(ctx, fresh, path)
mustIAM(t, err)
if r.Deleted || !r.Credentials.IsExpired() {
t.Fatal("failed cleanup lost the existing expired revision")
}
}
+220
View File
@@ -0,0 +1,220 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"fmt"
"maps"
"strings"
"sync"
"time"
"github.com/minio/minio/internal/auth"
)
// This index is rebuilt by the existing IAM loaders and updated by successful
// storage operations. It avoids a second full IAM walk during every heal pass.
// It is an optimization of the durable records, never a reason to delete them.
// The index contains no secrets or grants.
type iamParentRevision struct {
deleted bool
before time.Time
}
type iamRevisionIndex struct {
mu sync.RWMutex
items map[string]iamRevision
parents map[string]iamParentRevision
floors map[string]time.Time
generation uint64
}
func (idx *iamRevisionIndex) observe(path string, data []byte) {
if !strings.HasPrefix(path, iamConfigPrefix+"/") {
return
}
var r iamRevision
if json.Unmarshal(data, &r) != nil {
return // The caller reports malformed data using its normal decoder.
}
r.Credentials = auth.Credentials{ParentUser: r.Credentials.ParentUser, Expiration: r.Credentials.Expiration}
idx.mu.Lock()
defer idx.mu.Unlock()
if strings.HasPrefix(path, iamConfigUsersPrefix) {
// Keep a compact name-keyed view for the authentication hot path;
// constructing a config path on every S3 request allocates needlessly.
defer func() {
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile)
if current, ok := idx.items[path]; ok {
if idx.parents == nil {
idx.parents = make(map[string]iamParentRevision)
}
idx.parents[name] = iamParentRevision{deleted: current.Deleted, before: current.RevokedBefore}
} else {
delete(idx.parents, name)
}
}()
}
if floor, ok := idx.floors[path]; ok && r.timestamp().Before(floor) {
return
}
if previous, ok := idx.items[path]; ok {
// A concurrent read that began before a write must not roll it back.
if previous.timestamp().After(r.timestamp()) || (previous.Deleted && !r.Deleted && !r.timestamp().After(previous.timestamp())) {
return
}
if previous.RevokedBefore.After(r.RevokedBefore) {
r.RevokedBefore = previous.RevokedBefore
}
if previous.timestamp().Equal(r.timestamp()) && previous.Deleted == r.Deleted && previous.RevokedBefore.Equal(r.RevokedBefore) {
return
}
}
if r.Deleted && !r.ExpiresAt.IsZero() && UTCNow().After(r.ExpiresAt) {
if _, tracked := idx.items[path]; tracked {
delete(idx.items, path)
idx.generation++
}
delete(idx.floors, path)
return
}
if !r.Deleted && r.RevokedBefore.IsZero() {
_, tracked := idx.items[path]
_, hasFloor := idx.floors[path]
if tracked || hasFloor {
if idx.floors == nil {
idx.floors = make(map[string]time.Time)
}
idx.floors[path] = r.timestamp()
}
if tracked {
delete(idx.items, path)
idx.generation++
}
return
}
if idx.items == nil {
idx.items = make(map[string]iamRevision)
}
idx.items[path] = r
delete(idx.floors, path)
idx.generation++
}
func (idx *iamRevisionIndex) get(path string) iamRevision {
if idx == nil {
return iamRevision{}
}
idx.mu.RLock()
defer idx.mu.RUnlock()
return idx.items[path]
}
func (idx *iamRevisionIndex) snapshot() map[string]iamRevision {
idx.mu.Lock()
defer idx.mu.Unlock()
for path, r := range idx.items {
if r.Deleted && !r.ExpiresAt.IsZero() && UTCNow().After(r.ExpiresAt) {
delete(idx.items, path)
delete(idx.floors, path)
idx.generation++
}
}
return maps.Clone(idx.items)
}
func (idx *iamRevisionIndex) count() int {
idx.mu.RLock()
defer idx.mu.RUnlock()
return len(idx.items)
}
// A process-local generation plus the protocol's instance ID is sufficient
// for acknowledgements. Avoid hashing the entire index on every IAM write.
func (idx *iamRevisionIndex) digest() string {
idx.mu.RLock()
defer idx.mu.RUnlock()
return fmt.Sprintf("%x:%x", idx.generation, len(idx.items))
}
func (idx *iamRevisionIndex) forget(path string) {
idx.mu.Lock()
if _, ok := idx.items[path]; ok {
delete(idx.items, path)
idx.generation++
}
delete(idx.floors, path)
if strings.HasPrefix(path, iamConfigUsersPrefix) {
delete(idx.parents, strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile))
}
idx.mu.Unlock()
}
func (c *iamCache) userRevocation(user string) iamRevision {
r := c.revisions.parentRevision(user)
if u, ok := c.iamUsersMap[user]; ok && u.RevokedBefore.After(r.RevokedBefore) {
r.RevokedBefore = u.RevokedBefore
}
return r
}
func (c *iamCache) groupMemberAllowed(member string, grantedAt, groupBoundary time.Time) bool {
r := c.userRevocation(member)
return !r.Deleted && (r.RevokedBefore.IsZero() || grantedAt.After(r.RevokedBefore)) && (groupBoundary.IsZero() || grantedAt.After(groupBoundary))
}
func iamMappingParentPath(path string) string {
kind, name, ok := strings.Cut(strings.TrimPrefix(path, iamConfigPolicyDBPrefix), "/")
if !ok {
return ""
}
name = strings.TrimSuffix(name, ".json")
switch kind {
case "users", "sts-users":
return getUserIdentityPath(name, regUser)
case "service-accounts":
return getUserIdentityPath(name, svcUser)
case "groups":
return getGroupInfoPath(name)
}
return ""
}
func (idx *iamRevisionIndex) mappingAllowed(path string, mp MappedPolicy) bool {
if mp.Deleted || idx.get(path).Deleted {
return false
}
r := idx.get(iamMappingParentPath(path))
return !r.Deleted && (r.RevokedBefore.IsZero() || mp.UpdatedAt.After(r.RevokedBefore))
}
// Apply the persisted commit boundary even before dependent cache cleanup has
// completed. The map namespace is part of the authorization record's identity.
func (c *iamCache) cachedMappedPolicy(name string, userType IAMUserType, isGroup bool) (MappedPolicy, bool) {
var mp MappedPolicy
var ok bool
switch {
case isGroup:
mp, ok = c.iamGroupPolicyMap.Load(name)
case userType == stsUser:
mp, ok = c.iamSTSPolicyMap.Load(name)
default:
mp, ok = c.iamUserPolicyMap.Load(name)
}
if !ok || !c.revisions.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), mp) {
return MappedPolicy{}, false
}
return mp, true
}
func (idx *iamRevisionIndex) parentRevision(user string) iamRevision {
if idx == nil {
return iamRevision{}
}
idx.mu.RLock()
p := idx.parents[user]
idx.mu.RUnlock()
return iamRevision{Deleted: p.deleted, RevokedBefore: p.before}
}
+305
View File
@@ -0,0 +1,305 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"crypto/sha256"
"fmt"
"os"
"strings"
"sync"
"testing"
"time"
"github.com/minio/madmin-go/v3"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/concurrency"
"go.etcd.io/etcd/client/v3/namespace"
)
type iamRevisionLockObserver struct {
ObjectLayer
path string
waiting chan struct{}
once sync.Once
}
func (o *iamRevisionLockObserver) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket == minioMetaBucket && len(objects) == 1 && objects[0] == o.path {
return &iamRevisionObservedLock{RWLocker: lock, observe: func() { o.once.Do(func() { close(o.waiting) }) }}
}
return lock
}
type iamRevisionObservedLock struct {
RWLocker
observe func()
}
func (l *iamRevisionObservedLock) GetLock(ctx context.Context, timeout *dynamicTimeout) (LockContext, error) {
l.observe()
return l.RWLocker.GetLock(ctx, timeout)
}
type iamRevisionWatchObserver struct {
etcd.Watcher
waiting chan struct{}
once sync.Once
}
func (w *iamRevisionWatchObserver) Watch(ctx context.Context, key string, opts ...etcd.OpOption) etcd.WatchChan {
w.once.Do(func() { close(w.waiting) })
return w.Watcher.Watch(ctx, key, opts...)
}
// Simulate an unavailable cleanup RPC. Mutex.Lock calls Delete after its wait
// is canceled; that RPC must inherit a deadline too, not Client.Ctx() forever.
type iamRevisionCleanupBlocker struct {
etcd.KV
release chan struct{}
}
func (b *iamRevisionCleanupBlocker) Delete(ctx context.Context, key string, opts ...etcd.OpOption) (*etcd.DeleteResponse, error) {
if strings.Contains(key, "/iam-revision-locks/") {
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-b.release:
}
}
return b.KV.Delete(ctx, key, opts...)
}
type iamRevisionReadBlocker struct {
IAMStorageAPI
path string
after int
waiting chan struct{}
}
func (b *iamRevisionReadBlocker) loadIAMConfig(ctx context.Context, item any, path string) error {
if path == b.path {
b.after--
if b.after == 0 {
close(b.waiting)
<-ctx.Done()
return ctx.Err()
}
}
return b.IAMStorageAPI.loadIAMConfig(ctx, item, path)
}
func TestIAMRevisionReadDoesNotBlockAuthentication(t *testing.T) {
for _, stage := range []struct {
name string
offset time.Duration
}{{"deletion", time.Minute}, {"retained_revocation", -time.Minute}} {
t.Run(stage.name, func(t *testing.T) {
resetTestGlobals()
t.Cleanup(resetTestGlobals)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() {
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
})
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
const user = "read-blocked-parent"
created, err := store.AddUser(ctx, user, madmin.AddOrUpdateUserReq{SecretKey: "original-password", Status: madmin.AccountEnabled})
if err != nil {
t.Fatal(err)
}
blocked := &iamRevisionReadBlocker{IAMStorageAPI: store.IAMStorageAPI, path: getUserIdentityPath(user, regUser), after: 1, waiting: make(chan struct{})}
store.IAMStorageAPI = blocked
done := make(chan error, 1)
go func() {
done <- store.DeleteUser(withIAMReplicationTime(ctx, created.Add(stage.offset)), user, regUser)
}()
defer func() { cancel(); <-done }()
select {
case <-blocked.waiting:
case <-time.After(5 * time.Second):
t.Fatal("revision read was not attempted")
}
read := make(chan bool, 1)
go func() {
u, ok := store.GetUser(user)
read <- ok && u.Credentials.SecretKey == "original-password"
}()
select {
case ok := <-read:
if !ok {
t.Fatal("pending revision read changed the cached identity")
}
case <-time.After(time.Second):
t.Fatal("revision read blocked cached authentication")
}
})
}
}
func TestIAMRevisionLockContention(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if backend == "etcd" && endpoint == "" {
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
for _, outcome := range []string{"release", "cancel", "default_timeout"} {
t.Run(outcome, func(t *testing.T) {
resetTestGlobals()
t.Cleanup(resetTestGlobals)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
oldTimeout := defaultContextTimeout
defaultContextTimeout = 2 * time.Second
t.Cleanup(func() { defaultContextTimeout = oldTimeout })
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
const user = "contended-user"
path := getUserIdentityPath(user, regUser)
waiting := make(chan struct{})
var store *IAMStoreSys
var hold func() func()
unblockCleanup := func() {}
if backend == "object" {
disks, err := getRandomDisks(1)
must(err)
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
must(err)
t.Cleanup(func() {
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
})
observed := &iamRevisionLockObserver{ObjectLayer: obj, path: path + ".revision-lock", waiting: waiting}
store = &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
hold = func() func() {
lock := obj.NewNSLock(minioMetaBucket, observed.path)
lc, err := lock.GetLock(ctx, newDynamicTimeout(time.Second, time.Second))
must(err)
store.IAMStorageAPI.(*IAMObjectStore).objAPI = observed
return func() { lock.Unlock(lc) }
}
} else {
client, err := etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: time.Second})
must(err)
t.Cleanup(func() { client.Close() })
prefix := fmt.Sprintf("/silo-lock-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(client.KV, prefix)
client.Watcher = namespace.NewWatcher(client.Watcher, prefix)
store = &IAMStoreSys{IAMStorageAPI: newIAMEtcdStore(client, MinIOUsersSysType)}
hold = func() func() {
session, err := concurrency.NewSession(client, concurrency.WithContext(ctx))
must(err)
lock := concurrency.NewMutex(session, fmt.Sprintf("%s/iam-revision-locks/%x", minioConfigPrefix, sha256.Sum256([]byte(path))))
must(lock.Lock(ctx))
client.Watcher = &iamRevisionWatchObserver{Watcher: client.Watcher, waiting: waiting}
blocker := &iamRevisionCleanupBlocker{KV: client.KV, release: make(chan struct{})}
client.KV = blocker
unblockCleanup = sync.OnceFunc(func() { close(blocker.release) })
t.Cleanup(unblockCleanup)
return func() { session.Close() }
}
}
request := func(secret string) madmin.AddOrUpdateUserReq {
return madmin.AddOrUpdateUserReq{SecretKey: secret, Status: madmin.AccountEnabled}
}
_, err := store.AddUser(ctx, user, request("original-password"))
must(err)
release := sync.OnceFunc(hold())
t.Cleanup(release)
writeCtx, cancelWrite := context.WithCancel(ctx)
defer cancelWrite()
first, second := make(chan error, 1), make(chan error, 1)
var writers sync.WaitGroup
t.Cleanup(func() {
cancelWrite()
unblockCleanup()
release()
writers.Wait()
})
writers.Go(func() {
_, err := store.AddUser(writeCtx, user, request("first-password"))
first <- err
})
select {
case <-waiting:
case <-time.After(5 * time.Second):
t.Fatal("writer did not attempt the held revision lock")
}
// A second writer must queue without taking the cache's RWMutex:
// Go's writer preference would otherwise block every new reader.
writers.Go(func() {
_, err := store.AddUser(ctx, user, request("second-password"))
second <- err
})
select {
case err := <-second:
t.Fatalf("second writer bypassed the first: %v", err)
case <-time.After(50 * time.Millisecond):
}
read := make(chan UserIdentity, 1)
go func() {
u, _ := store.GetUser(user)
read <- u
}()
select {
case u := <-read:
if u.Credentials.SecretKey != "original-password" {
t.Fatal("pending write changed the cached credential")
}
case <-time.After(time.Second):
t.Fatal("distributed lock contention blocked cached authentication")
}
switch outcome {
case "release":
release()
case "cancel":
cancelWrite()
}
select {
case err := <-first:
if outcome == "release" {
must(err)
} else if err == nil {
t.Fatal("canceled or timed-out write succeeded")
}
case <-time.After(5 * time.Second):
t.Fatal("lock wait or cancellation cleanup exceeded its deadline")
}
release()
select {
case err := <-second:
must(err)
case <-time.After(5 * time.Second):
t.Fatal("queued writer did not recover after the first completed")
}
cached, ok := store.GetUser(user)
if !ok || cached.Credentials.SecretKey != "second-password" {
t.Fatal("cached write order was lost")
}
var persisted UserIdentity
must(store.loadIAMConfig(ctx, &persisted, path))
if persisted.Credentials.SecretKey != cached.Credentials.SecretKey || !persisted.UpdatedAt.Equal(cached.UpdatedAt) {
t.Fatal("persistent and cached revisions differ")
}
})
}
})
}
}
+749
View File
@@ -0,0 +1,749 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"crypto/sha256"
"errors"
"fmt"
"net/http"
"sort"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/concurrency"
)
var errIAMStaleUpdate = errors.New("IAM update predates a stored revision or revocation")
// The parent is still live; callers must not broadcast a user deletion when
// only its revocation boundary was retained.
var errIAMRevocationRetained = errors.New("IAM revocation recorded without deleting the record")
// A revocation advances the boundary even when a newer identity already
// exists. Keep this operation distinct from replacing/deleting that identity.
type iamUserRevocation struct {
UserIdentity
retained bool
}
type iamGroupRevocation struct {
GroupInfo
retained bool
requireEmpty bool
}
// Natural expiration is distinct from revoking a live credential. An expired
// immutable STS token can be removed; a reusable service-account key retains
// its revision so an older non-expiring credential cannot return.
type iamExpireIdentity struct{}
// The authoritative revocation is durable even if dependent cleanup fails.
// Callers must publish it to sibling caches before returning the error.
type iamCommittedCleanupError struct {
err error
retained bool
}
func (e *iamCommittedCleanupError) Error() string {
return "IAM revocation committed; cleanup failed: " + e.err.Error()
}
func (e *iamCommittedCleanupError) Unwrap() error { return e.err }
type iamReplicationTimeKey struct{}
func withIAMReplicationTime(ctx context.Context, at time.Time) context.Context {
return context.WithValue(ctx, iamReplicationTimeKey{}, at)
}
func iamReplicationTime(ctx context.Context) (time.Time, bool) {
at, ok := ctx.Value(iamReplicationTimeKey{}).(time.Time)
return at, ok
}
func iamReplicationError(err error) error {
if errors.Is(err, errIAMStaleUpdate) {
// Retrying an obsolete event cannot change the result.
return nil
}
return wrapSRErr(err)
}
// Deletions occupy the original IAM config path. They contain no secret or
// grant and are hidden by the normal loaders, but remain available to heal
// and to timestamp comparisons after a restart. Do not age them out: a peer
// can be offline indefinitely.
type iamRevision struct {
UpdatedAt time.Time `json:"updatedAt"`
UpdateDate time.Time `json:"UpdateDate"`
Deleted bool `json:"deleted"`
RevokedBefore time.Time `json:"revokedBefore"`
ExpiresAt time.Time `json:"expiresAt,omitempty"`
Credentials auth.Credentials `json:"credentials"`
}
func (r iamRevision) timestamp() time.Time {
if r.UpdateDate.After(r.UpdatedAt) {
return r.UpdateDate
}
return r.UpdatedAt
}
func loadIAMRevision(ctx context.Context, store IAMStorageAPI, path string) (iamRevision, error) {
var r iamRevision
err := store.loadIAMConfig(ctx, &r, path)
if errors.Is(err, errConfigNotFound) {
err = nil
}
return r, err
}
func (store *IAMStoreSys) checkIAMRevision(ctx context.Context, path string, deleting bool) error {
at, replicated := iamReplicationTime(ctx)
if !replicated {
return nil
}
return store.withIAMStorage(ctx, func(ctx context.Context) error {
r, err := loadIAMRevision(ctx, store.IAMStorageAPI, path)
if err != nil {
return err
}
if r.timestamp().After(at) || (r.Deleted && !deleting && !at.After(r.timestamp())) {
return errIAMStaleUpdate
}
return nil
})
}
// This signed claim records the parent's revocation boundary at issuance.
// Unlike UpdatedAt, it cannot advance when an offline site edits an old child.
// It travels in the existing service-account Claims and STS SessionToken fields.
const iamParentRevocationClaim = "siloParentRevocation"
func setIAMParentRevocationClaim(ctx context.Context, store IAMStorageAPI, parent string, claims map[string]any) error {
delete(claims, iamParentRevocationClaim)
if parent == "" || parent == globalActiveCred.AccessKey {
return nil
}
r, err := loadIAMRevision(ctx, store, getUserIdentityPath(parent, regUser))
if err != nil {
return err
}
if r.Deleted {
return errIAMStaleUpdate
}
if !r.RevokedBefore.IsZero() {
claims[iamParentRevocationClaim] = r.RevokedBefore.Format(time.RFC3339Nano)
}
return nil
}
func iamCredentialSurvivesRevocation(cred auth.Credentials, at time.Time) bool {
if at.IsZero() {
return true
}
s, _ := cred.Claims[iamParentRevocationClaim].(string)
issuedAfter, err := time.Parse(time.RFC3339Nano, s)
return err == nil && !issuedAfter.Before(at)
}
// Parent revocations delete old children even if an offline peer has edited
// them later. Preserve children that prove issuance after this revocation.
func iamChildDeletionContext(ctx context.Context, child UserIdentity) (context.Context, bool) {
if at, replicated := iamReplicationTime(ctx); replicated {
if !at.IsZero() && iamCredentialSurvivesRevocation(child.Credentials, at) {
return ctx, false
}
if child.UpdatedAt.After(at) {
ctx = withIAMReplicationTime(ctx, child.UpdatedAt)
}
}
return ctx, true
}
// A delayed service account or STS event must not outlive deletion of its
// built-in parent. The caller must populate Claims from the verified token.
func checkIAMParentRevision(ctx context.Context, store IAMStorageAPI, cred auth.Credentials) error {
parent := cred.ParentUser
if parent == "" || parent == globalActiveCred.AccessKey {
return nil
}
r, err := loadIAMRevision(ctx, store, getUserIdentityPath(parent, regUser))
if err != nil {
return err
}
if r.Deleted || !iamCredentialSurvivesRevocation(cred, r.RevokedBefore) {
return errIAMStaleUpdate
}
return nil
}
// Called with the IAM writer mutex and cache lock held. Persistence only
// touches the caller's record, not the cache. Keep writers serialized while
// allowing cached authentication reads throughout storage and lock waits.
func (store *IAMStoreSys) withIAMStorage(ctx context.Context, fn func(context.Context) error) error {
store.IAMStorageAPI.unlock()
defer store.IAMStorageAPI.lock()
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
return fn(ctx)
}
func (store *IAMStoreSys) saveIAMRevision(ctx context.Context, path string, item any, opts ...options) error {
return store.withIAMStorage(ctx, func(ctx context.Context) error {
return saveIAMRevision(ctx, store.IAMStorageAPI, path, item, opts...)
})
}
func (store *IAMStoreSys) checkIAMParentRevision(ctx context.Context, cred auth.Credentials) error {
return store.withIAMStorage(ctx, func(ctx context.Context) error {
return checkIAMParentRevision(ctx, store.IAMStorageAPI, cred)
})
}
// Update the caller's record with the persisted revision before it is cached.
func saveIAMRevision(ctx context.Context, store IAMStorageAPI, path string, item any, opts ...options) error {
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
// Serialize compare-and-write across nodes, as well as goroutines. Use a
// separate lock name so saving the config does not reacquire this lock.
switch s := store.(type) {
case *IAMObjectStore:
lock := s.objAPI.NewNSLock(minioMetaBucket, path+".revision-lock")
lc, err := lock.GetLock(ctx, globalOperationTimeout)
if err != nil {
return err
}
defer lock.Unlock(lc)
ctx = lc.Context()
case *IAMEtcdStore:
// Mutex.Lock also uses Client.Ctx() for cleanup after cancellation.
// Borrow the existing services with the operation's bounded context;
// never close this facade, which does not own those services.
client := etcd.NewCtxClient(ctx, etcd.WithZapLogger(s.client.GetLogger()))
client.KV, client.Lease, client.Watcher = s.client.KV, s.client.Lease, s.client.Watcher
session, err := concurrency.NewSession(client, concurrency.WithContext(ctx))
if err != nil {
return err
}
defer func() {
session.Orphan()
// A canceled operation must still release its lease when etcd is
// reachable. If it is unavailable, stop waiting and let it expire.
cleanupCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), defaultContextTimeout)
defer cancel()
_, _ = s.client.Revoke(cleanupCtx, session.Lease())
}()
lock := concurrency.NewMutex(session, fmt.Sprintf("%s/iam-revision-locks/%x", minioConfigPrefix, sha256.Sum256([]byte(path))))
if err = lock.Lock(ctx); err != nil {
return err
}
// Revoking the session lease releases the lock, including on cancellation.
}
previous, err := loadIAMRevision(ctx, store, path)
if err != nil {
return err
}
if _, expiring := item.(*iamExpireIdentity); expiring {
sts := strings.HasPrefix(path, iamConfigSTSPrefix)
if previous.Deleted {
if sts && !previous.ExpiresAt.IsZero() && UTCNow().After(previous.ExpiresAt) {
return expireIAMSTSConfig(ctx, store, path)
}
return nil
}
if previous.timestamp().IsZero() || !previous.Credentials.IsExpired() {
return nil
}
if sts {
return expireIAMSTSConfig(ctx, store, path)
}
item = &UserIdentity{Version: 1, Deleted: true}
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
var revocation *iamUserRevocation
if op, ok := item.(*iamUserRevocation); ok {
revocation = op
op.UserIdentity = UserIdentity{Version: 1, Deleted: true}
if origin, replicated := iamReplicationTime(ctx); replicated && previous.timestamp().After(origin) {
if previous.Deleted || !origin.After(previous.RevokedBefore) {
return errIAMStaleUpdate
}
op.retained = true
op.UserIdentity = UserIdentity{Version: 1, Credentials: previous.Credentials, UpdatedAt: previous.timestamp(), RevokedBefore: origin}
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
item = &op.UserIdentity
}
var groupRevocation *iamGroupRevocation
if op, ok := item.(*iamGroupRevocation); ok {
groupRevocation = op
var group GroupInfo
if err := store.loadIAMConfig(ctx, &group, path); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
if op.requireEmpty && !group.Deleted {
for _, member := range group.Members {
r := store.revisionIndex().get(getUserIdentityPath(member, regUser))
at := group.MemberGrants[member]
if !r.Deleted && (r.RevokedBefore.IsZero() || at.After(r.RevokedBefore)) && (group.RevokedBefore.IsZero() || at.After(group.RevokedBefore)) {
return errGroupNotEmpty
}
}
}
op.GroupInfo = GroupInfo{Version: 1, Deleted: true}
if origin, replicated := iamReplicationTime(ctx); replicated && previous.timestamp().After(origin) {
if previous.Deleted || !origin.After(previous.RevokedBefore) {
return errIAMStaleUpdate
}
op.retained = true
op.GroupInfo = group
op.RevokedBefore = origin
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
item = &op.GroupInfo
}
var at *time.Time
var deleted bool
switch v := item.(type) {
case *UserIdentity:
at, deleted = &v.UpdatedAt, v.Deleted
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(v.RevokedBefore) {
v.RevokedBefore = boundary
}
if previous.RevokedBefore.After(v.RevokedBefore) {
v.RevokedBefore = previous.RevokedBefore
}
case *GroupInfo:
at, deleted = &v.UpdatedAt, v.Deleted
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(v.RevokedBefore) {
v.RevokedBefore = boundary
}
if !deleted {
var group GroupInfo
if err := store.loadIAMConfig(ctx, &group, path); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
mergeIAMGroupMutation(ctx, group, v)
}
if previous.RevokedBefore.After(v.RevokedBefore) {
v.RevokedBefore = previous.RevokedBefore
}
case *MappedPolicy:
at, deleted = &v.UpdatedAt, v.Deleted
case *PolicyDoc:
at, deleted = &v.UpdateDate, v.Deleted
default:
return errInvalidArgument
}
if strings.HasPrefix(path, iamConfigSTSPrefix) && previous.Deleted && !deleted {
// STS access keys identify immutable tokens, not reusable user names.
return errIAMStaleUpdate
}
if origin, replicated := iamReplicationTime(ctx); replicated {
*at = origin
if previous.timestamp().After(origin) || (previous.Deleted && !deleted && !origin.After(previous.timestamp())) {
return errIAMStaleUpdate
}
if strings.HasPrefix(path, iamConfigServiceAccountsPrefix) && !deleted && previous.Credentials.AccessKey != "" && previous.timestamp().Equal(origin) {
// Duplicate service snapshots are acknowledgements, not new creates
// or edits. Reload the winner without writing, so even a stale
// sibling cache is refreshed by the retry before acknowledging it.
return store.loadIAMConfig(ctx, item, path)
}
if previous.Deleted && deleted && !origin.After(previous.timestamp()) {
// An already-applied tombstone needs no further persistent write.
if v, ok := item.(*UserIdentity); ok {
v.RevokedBefore = previous.RevokedBefore
}
if v, ok := item.(*GroupInfo); ok {
v.RevokedBefore = previous.RevokedBefore
}
return nil
}
} else {
if previous.Deleted && deleted {
// A peer notification without an originating revision must not
// advance a tombstone past a subsequent deliberate recreation.
*at = previous.timestamp()
if v, ok := item.(*UserIdentity); ok {
v.RevokedBefore = previous.RevokedBefore
}
if v, ok := item.(*GroupInfo); ok {
v.RevokedBefore = previous.RevokedBefore
}
return nil
}
if at.IsZero() {
*at = UTCNow()
}
if !at.After(previous.timestamp()) {
*at = previous.timestamp().Add(time.Nanosecond)
}
}
if v, ok := item.(*UserIdentity); ok {
if deleted {
// Retain only the parent name for root-account exclusion during heal.
v.Credentials = auth.Credentials{ParentUser: previous.Credentials.ParentUser}
v.RevokedBefore = *at
if strings.HasPrefix(path, iamConfigSTSPrefix) && !previous.Credentials.Expiration.IsZero() && !previous.Credentials.Expiration.Equal(timeSentinel) {
// The signed STS token cannot authorize beyond this time, even
// if an offline site replays it with a newer event timestamp.
v.ExpiresAt = previous.Credentials.Expiration.Add(globalMaxSkewTime)
opts = []options{{ttl: max(1, int64(time.Until(v.ExpiresAt).Seconds())+1)}}
}
} else {
if v.Credentials.SessionToken != "" && v.Credentials.Claims == nil {
claims, err := extractJWTClaims(*v)
if err != nil {
return err
}
v.Credentials.Claims = claims.Map()
}
if err = checkIAMParentRevision(ctx, store, v.Credentials); err != nil {
return err
}
}
}
if v, ok := item.(*GroupInfo); ok && deleted {
v.RevokedBefore = *at
v.Members, v.MemberGrants = nil, nil
}
if _, ok := item.(*MappedPolicy); ok && !deleted {
if parentPath := iamMappingParentPath(path); parentPath != "" {
parent, err := loadIAMRevision(ctx, store, parentPath)
if err != nil {
return err
}
if parent.Deleted {
return errIAMStaleUpdate
}
if !parent.RevokedBefore.IsZero() && !at.After(parent.RevokedBefore) {
if _, replicated := iamReplicationTime(ctx); replicated {
return errIAMStaleUpdate
}
*at = parent.RevokedBefore.Add(time.Nanosecond)
}
}
}
if err := store.saveIAMConfig(ctx, item, path, opts...); err != nil {
return err
}
if revocation != nil && revocation.retained {
return errIAMRevocationRetained
}
if groupRevocation != nil && groupRevocation.retained {
return errIAMRevocationRetained
}
return nil
}
func (iamOS *IAMObjectStore) listIAMConfigPaths(ctx context.Context) ([]string, error) {
ctx, cancel := context.WithCancel(ctx)
defer cancel()
var paths []string
for item := range listIAMConfigItems(ctx, iamOS.objAPI, iamConfigPrefix+"/") {
if item.Err != nil {
return nil, item.Err
}
paths = append(paths, iamConfigPrefix+"/"+item.Item)
}
return paths, nil
}
func (ies *IAMEtcdStore) listIAMConfigPaths(ctx context.Context) ([]string, error) {
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
r, err := ies.client.Get(ctx, iamConfigPrefix+"/", etcd.WithPrefix(), etcd.WithKeysOnly())
if err != nil {
return nil, err
}
paths := make([]string, 0, len(r.Kvs))
for _, kv := range r.Kvs {
paths = append(paths, string(kv.Key))
}
return paths, nil
}
func iamDeletionItem(path string, r iamRevision) (item madmin.SRIAMItem, ok bool) {
if (strings.HasPrefix(path, iamConfigUsersPrefix) || strings.HasPrefix(path, iamConfigGroupsPrefix)) && !r.RevokedBefore.IsZero() {
// Recreating a parent does not cancel its older revocation of derived
// credentials. Replay this boundary even after the parent is live again.
r.Deleted = true
r.UpdatedAt, r.UpdateDate = r.RevokedBefore, time.Time{}
}
if !r.Deleted {
return item, false
}
item.UpdatedAt = r.timestamp()
switch {
case strings.HasPrefix(path, iamConfigUsersPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile)
item.Type = madmin.SRIAMItemIAMUser
item.IAMUser = &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}
case strings.HasPrefix(path, iamConfigServiceAccountsPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigServiceAccountsPrefix), "/"+iamIdentityFile)
if name == siteReplicatorSvcAcc || r.Credentials.ParentUser == globalActiveCred.AccessKey {
return item, false
}
item.Type = madmin.SRIAMItemSvcAcc
item.SvcAccChange = &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: name}}
case strings.HasPrefix(path, iamConfigGroupsPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigGroupsPrefix), "/"+iamGroupMembersFile)
item.Type = madmin.SRIAMItemGroupInfo
item.GroupInfo = &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: name, IsRemove: true}}
case strings.HasPrefix(path, iamConfigPoliciesPrefix):
item.Type = madmin.SRIAMItemPolicy
item.Name = strings.TrimSuffix(strings.TrimPrefix(path, iamConfigPoliciesPrefix), "/"+iamPolicyFile)
case strings.HasPrefix(path, iamConfigPolicyDBPrefix):
prefix, name, found := strings.Cut(strings.TrimPrefix(path, iamConfigPolicyDBPrefix), "/")
if !found {
return item, false
}
typ := regUser
switch prefix {
case "sts-users":
typ = stsUser
case "service-accounts":
typ = svcUser
}
item.Type = madmin.SRIAMItemPolicyMapping
item.PolicyMapping = &madmin.SRPolicyMapping{UserOrGroup: strings.TrimSuffix(name, ".json"), UserType: int(typ), IsGroup: prefix == "groups"}
default:
// Expired STS credentials are not replayed. Parent revocations and
// their retained timestamp reject delayed copies of derived tokens.
return item, false
}
return item, true
}
func iamDeletionPath(item madmin.SRIAMItem) string {
switch item.Type {
case madmin.SRIAMItemIAMUser:
if item.IAMUser != nil && item.IAMUser.IsDeleteReq {
return getUserIdentityPath(item.IAMUser.AccessKey, regUser)
}
case madmin.SRIAMItemSvcAcc:
if item.SvcAccChange != nil && item.SvcAccChange.Delete != nil {
return getUserIdentityPath(item.SvcAccChange.Delete.AccessKey, svcUser)
}
case madmin.SRIAMItemGroupInfo:
if item.GroupInfo != nil && item.GroupInfo.UpdateReq.IsRemove && len(item.GroupInfo.UpdateReq.Members) == 0 {
return getGroupInfoPath(item.GroupInfo.UpdateReq.Group)
}
case madmin.SRIAMItemPolicy:
if len(item.Policy) == 0 {
return getPolicyDocPath(item.Name)
}
case madmin.SRIAMItemPolicyMapping:
if p := item.PolicyMapping; p != nil && p.Policy == "" {
return getMappedPolicyPath(p.UserOrGroup, IAMUserType(p.UserType), p.IsGroup)
}
}
return ""
}
func (c *SiteReplicationSys) healIAMDeletions(ctx context.Context) (err error) {
started := time.Now()
defer func() {
c.iamRevisionMetrics.healDurationMillis.Store(time.Since(started).Milliseconds())
if err != nil {
c.iamRevisionMetrics.healFailures.Add(1)
} else {
c.iamRevisionMetrics.healLastSuccess.Store(time.Now().Unix())
}
}()
c.iamHealMu.Lock()
defer c.iamHealMu.Unlock()
c.RLock()
defer c.RUnlock()
if !c.enabled {
return nil
}
snapshot := globalIAMSys.store.revisionIndex().snapshot()
paths := make([]string, 0, len(snapshot))
for path := range snapshot {
paths = append(paths, path)
}
sort.Strings(paths)
byType := make(map[string][]iamReplicationItem)
for _, path := range paths {
r := snapshot[path]
item, ok := iamDeletionItem(path, r)
if !ok {
continue
}
out := iamReplicationItem{SRIAMItem: item}
if item.Type == madmin.SRIAMItemIAMUser && !r.Deleted {
out.Type, out.IAMUser = iamUserBoundaryType, nil
out.UserRevocation = &iamUserBoundary{User: item.IAMUser.AccessKey, Before: r.RevokedBefore}
}
if item.Type == madmin.SRIAMItemGroupInfo && !r.Deleted {
out.Type, out.GroupInfo = iamGroupBoundaryType, nil
out.GroupRevocation = &iamGroupBoundary{Group: item.GroupInfo.UpdateReq.Group, Before: r.RevokedBefore}
}
byType[item.Type] = append(byType[item.Type], out)
}
var items []iamReplicationItem
for _, typ := range []string{madmin.SRIAMItemPolicyMapping, madmin.SRIAMItemIAMUser, madmin.SRIAMItemSvcAcc, madmin.SRIAMItemGroupInfo, madmin.SRIAMItemPolicy} {
items = append(items, byType[typ]...)
}
if len(items) == 0 {
return nil
}
if c.iamRevisionProgress == nil {
c.iamRevisionProgress = make(map[string]iamRevisionProgress)
}
for id := range c.iamRevisionProgress {
if _, present := c.state.Peers[id]; !present {
delete(c.iamRevisionProgress, id)
}
}
var progressMu sync.Mutex
cerr := c.concDo(nil, func(id string, p madmin.PeerInfo) error {
// Bound each pass, but retain acknowledgements independently of the
// pass deadline or unrelated changes at either site.
peerCtx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
client, err := c.getAdminClient(peerCtx, id)
if err != nil {
return err
}
remote, err := executeIAMRevisionRequest(peerCtx, client, http.MethodGet, nil)
if err != nil {
return err
}
progressMu.Lock()
progress := c.iamRevisionProgress[id]
progressMu.Unlock()
progress.observePeer(remote)
defer func() {
progressMu.Lock()
c.iamRevisionProgress[id] = progress
progressMu.Unlock()
}()
var pending []iamReplicationItem
for _, item := range items {
path, version := iamReplicationMarker(item)
if progress.Acknowledged[path] != version {
pending = append(pending, item)
}
}
// Acknowledgements are only a replay optimization, never GC proof.
for path := range progress.Acknowledged {
if _, retained := snapshot[path]; !retained {
delete(progress.Acknowledged, path)
}
}
var failures []error
for next := 0; next < len(pending); {
end := min(next+maxIAMRevisionBatch, len(pending))
batch := pending[next:end]
remote, err = executeIAMRevisionRequest(peerCtx, client, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: batch})
if err != nil {
var batchErr *iamRevisionBatchError
if !errors.As(err, &batchErr) {
return errors.Join(append(failures, err)...)
}
failures = append(failures, err)
} else {
progress.observePeer(remote)
for _, item := range batch {
path, version := iamReplicationMarker(item)
progress.Acknowledged[path] = version
}
}
next = end
}
return errors.Join(failures...)
}, "IAM revision convergence")
return errors.Unwrap(cerr)
}
func iamReplicationMarker(item iamReplicationItem) (path, version string) {
path = iamDeletionPath(item.SRIAMItem)
if item.UserRevocation != nil {
path = getUserIdentityPath(item.UserRevocation.User, regUser)
}
if item.GroupRevocation != nil {
path = getGroupInfoPath(item.GroupRevocation.Group)
}
return path, item.Type + ":" + item.UpdatedAt.UTC().Format(time.RFC3339Nano)
}
func (store *IAMStoreSys) savePolicyDoc(ctx context.Context, policyName string, p *PolicyDoc) error {
return store.saveIAMRevision(ctx, getPolicyDocPath(policyName), p)
}
func (store *IAMStoreSys) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp *MappedPolicy, opts ...options) error {
return store.saveIAMRevision(ctx, getMappedPolicyPath(name, userType, isGroup), mp, opts...)
}
func (store *IAMStoreSys) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u *UserIdentity, opts ...options) error {
return store.saveIAMRevision(ctx, getUserIdentityPath(name, userType), u, opts...)
}
func (store *IAMStoreSys) saveGroupInfo(ctx context.Context, name string, gi *GroupInfo) error {
return store.saveIAMRevision(ctx, getGroupInfoPath(name), gi)
}
func (store *IAMStoreSys) deletePolicyDoc(ctx context.Context, name string) error {
return store.saveIAMRevision(ctx, getPolicyDocPath(name), &PolicyDoc{Version: 1, Deleted: true})
}
func (store *IAMStoreSys) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
return store.saveIAMRevision(ctx, getMappedPolicyPath(name, userType, isGroup), &MappedPolicy{Version: 1, Deleted: true})
}
func (store *IAMStoreSys) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
return store.saveIAMRevision(ctx, getUserIdentityPath(name, userType), &UserIdentity{Version: 1, Deleted: true})
}
// Called under the identity's distributed revision lock, after verifying that
// its immutable STS token (or early-revocation retention) has expired. Only the
// old token-key mapping is removed; the reusable parent mapping is unaffected.
func expireIAMSTSConfig(ctx context.Context, store IAMStorageAPI, path string) error {
key := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigSTSPrefix), "/"+iamIdentityFile)
if err := store.deleteIAMConfig(ctx, getMappedPolicyPath(key, stsUser, false)); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
return store.deleteIAMConfig(ctx, path)
}
type (
iamExpirationCleanupKey struct{}
iamExpirationCleanupState struct{ failed atomic.Bool }
)
func withIAMExpirationCleanup(ctx context.Context) context.Context {
if _, ok := ctx.Value(iamExpirationCleanupKey{}).(*iamExpirationCleanupState); ok {
return ctx
}
return context.WithValue(ctx, iamExpirationCleanupKey{}, &iamExpirationCleanupState{})
}
func bestEffortIAMExpiration(ctx context.Context, store IAMStorageAPI, path string) {
state, _ := ctx.Value(iamExpirationCleanupKey{}).(*iamExpirationCleanupState)
if ctx.Err() != nil || (state != nil && state.failed.Load()) {
return
}
ctx, cancel := context.WithTimeout(ctx, time.Second)
defer cancel()
// Failure leaves the expired record and its existing version intact.
// Stop optional reclamation for this load, while still loading healthy
// users. Healthy cleanup has no per-scan quota that could build a backlog.
if err := saveIAMRevision(ctx, store, path, &iamExpireIdentity{}); err != nil {
if state != nil {
state.failed.Store(true)
}
iamLogIf(ctx, err)
}
}
+330
View File
@@ -0,0 +1,330 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
)
// Count physical saves: comparing timestamps alone would miss identical
// tombstones being rewritten on every heal pass.
type iamRevisionWriteCounter struct {
IAMStorageAPI
data []byte
writes int
}
func (s *iamRevisionWriteCounter) loadIAMConfig(_ context.Context, item any, _ string) error {
return json.Unmarshal(s.data, item)
}
func (s *iamRevisionWriteCounter) saveIAMConfig(_ context.Context, item any, _ string, _ ...options) error {
data, err := json.Marshal(item)
if err == nil {
s.data = data
s.writes++
}
return err
}
func TestIAMRevocationTombstoneReplayIsIdempotent(t *testing.T) {
at := time.Date(2026, 9, 14, 12, 0, 0, 0, time.UTC)
for _, record := range []struct {
name string
new func(bool) any
}{
{"user", func(deleted bool) any { return &UserIdentity{Version: 1, Deleted: deleted} }},
{"group", func(deleted bool) any { return &GroupInfo{Version: 1, Deleted: deleted} }},
{"policy", func(deleted bool) any { return &PolicyDoc{Version: 1, Deleted: deleted} }},
{"mapping", func(deleted bool) any { return &MappedPolicy{Version: 1, Deleted: deleted} }},
} {
t.Run(record.name, func(t *testing.T) {
data, err := json.Marshal(iamRevision{Deleted: true, UpdatedAt: at, RevokedBefore: at})
if err != nil {
t.Fatal(err)
}
store := &iamRevisionWriteCounter{data: data}
ctx := context.Background()
for range 3 {
// Site heal carries the original timestamp. Sibling notifications
// have no timestamp; both must leave an applied deletion untouched.
for _, replay := range []context.Context{withIAMReplicationTime(ctx, at), ctx} {
if err := saveIAMRevision(replay, store, record.name, record.new(true)); err != nil {
t.Fatal(err)
}
}
}
if store.writes != 0 || string(store.data) != string(data) {
t.Fatalf("replayed tombstone changed storage: writes=%d, record=%s", store.writes, store.data)
}
for _, deleted := range []bool{false, true} {
err := saveIAMRevision(withIAMReplicationTime(ctx, at.Add(-time.Second)), store, record.name, record.new(deleted))
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("older event accepted, deleted=%t: %v", deleted, err)
}
}
if err := saveIAMRevision(withIAMReplicationTime(ctx, at), store, record.name, record.new(false)); !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("equal-time recreation accepted: %v", err)
}
newer := at.Add(time.Minute)
if err := saveIAMRevision(withIAMReplicationTime(ctx, newer), store, record.name, record.new(true)); err != nil {
t.Fatal(err)
}
r, err := loadIAMRevision(ctx, store, record.name)
if err != nil || store.writes != 1 || !r.timestamp().Equal(newer) || !r.Deleted {
t.Fatalf("newer deletion did not advance storage: writes=%d, revision=%+v, error=%v", store.writes, r, err)
}
if err := saveIAMRevision(withIAMReplicationTime(ctx, newer.Add(time.Minute)), store, record.name, record.new(false)); err != nil {
t.Fatalf("newer recreation rejected: %v", err)
}
})
}
}
// Run with both object storage and etcd through TestIAMRevocation*Lifecycle.
// The object-store case uses the real peer RPC and deletion handler, so a
// spurious notification actually destroys the parent instead of only counting it.
func testIAMRevocationReplayAfterRecreation(ctx context.Context, t *testing.T, sys *IAMSys) {
t.Helper()
peer := &globalSiteReplicationSys
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user := "heal-recreated-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
deleted, recreated := origin.Add(time.Minute), origin.Add(3*time.Minute)
create := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, at))
}
revoke := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, at))
}
create(origin)
revoke(deleted)
create(recreated)
child, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "heal-recreated-child", secretKey: "valid-service-password",
})
must(err)
tg, err := grid.SetupTestGrid(2)
must(err)
t.Cleanup(tg.Cleanup)
var deletes atomic.Int32
server := &peerRESTServer{}
must(deleteUserRPC.Register(tg.Managers[1], func(req *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
deletes.Add(1)
return server.DeleteUserHandler(req)
}))
// Future user updates still use the normal peer reload notification.
must(loadUserRPC.Register(tg.Managers[1], server.LoadUserHandler))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
must(err)
previousNotifications := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection {
return tg.Managers[0].Connection(tg.Hosts[1])
},
}}}
t.Cleanup(func() { globalNotificationSys = previousNotifications })
assertLive := func(key string) {
t.Helper()
if _, ok := sys.GetUser(ctx, key); !ok {
t.Fatalf("live credential %s lost during deletion replay", key)
}
}
assertNoDelete := func() {
t.Helper()
if n := deletes.Load(); n != 0 {
t.Fatalf("retained revocation sent %d destructive sibling notifications", n)
}
}
for range 3 {
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
item, ok := iamDeletionItem(getUserIdentityPath(user, regUser), r)
if !ok || item.IAMUser == nil || !item.UpdatedAt.Equal(deleted) {
t.Fatal("recreated user lost its durable revocation replay")
}
must(peer.PeerIAMUserChangeHandler(ctx, item.IAMUser, item.UpdatedAt))
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
assertLive(child.AccessKey)
assertNoDelete()
}
// A divergent site sends a previously unseen revocation between our old
// boundary and recreation. Retain it and revoke old children, but never
// turn it into an unversioned delete of the recreated parent.
delayed := deleted.Add(time.Minute)
revoke(delayed)
assertNoDelete()
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
if _, ok := sys.GetUser(ctx, child.AccessKey); ok {
t.Fatal("child from before the delayed revocation remains usable")
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
if r.Deleted || !r.RevokedBefore.Equal(delayed) || !r.timestamp().Equal(recreated) {
t.Fatalf("retained revocation damaged the recreated identity: %+v", r)
}
// A genuinely newer deletion must still reach siblings and remove the
// parent plus credentials issued under its latest revocation boundary.
fresh, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "heal-fresh-child", secretKey: "valid-service-password",
})
must(err)
latest := recreated.Add(time.Minute)
revoke(latest)
wantDeletes := int32(1)
if sys.HasWatcher() {
wantDeletes = 0
}
if n := deletes.Load(); n != wantDeletes {
t.Fatalf("new deletion notifications=%d, want %d", n, wantDeletes)
}
for _, key := range []string{user, fresh.AccessKey} {
if _, ok := sys.GetUser(ctx, key); ok {
t.Fatalf("newer deletion left credential %s usable", key)
}
}
// Exercise the actual sibling handler again against the already persisted
// tombstone. Its context has no revision; it must not re-stamp the record.
_, remoteErr := server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
r, err = loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
if !r.Deleted || !r.timestamp().Equal(latest) {
t.Fatalf("sibling re-stamped the tombstone: got %s, want %s", r.timestamp(), latest)
}
create(latest.Add(time.Minute))
_, remoteErr = server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
assertLive(user)
revoke(latest)
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
}
// Counts what a retained revocation actually sends to sibling nodes.
func TestIAMRevocationRetainedReloadsSibling(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
t.Fatal(err)
}
initAllSubsystems(ctx)
globalIAMSys.Init(ctx, obj, nil, 2*time.Second)
defer os.RemoveAll(disks[0])
defer obj.Shutdown(ctx)
defer resetTestGlobals()
sys, peer := globalIAMSys, &globalSiteReplicationSys
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user := "retained-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
deleted, recreated := origin.Add(time.Minute), origin.Add(3*time.Minute)
create := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, at))
}
revoke := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, at))
}
create(origin)
revoke(deleted)
create(recreated)
child, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "retained-child", secretKey: "valid-service-password",
})
must(err)
// A sibling shares persistent state but has an independent IAM cache.
sibling := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, sys.usersSysType)}
must(sibling.LoadIAMCache(ctx, false))
if _, ok := sibling.GetUser(child.AccessKey); !ok {
t.Fatal("sibling fixture did not load child")
}
tg, err := grid.SetupTestGrid(2)
must(err)
t.Cleanup(tg.Cleanup)
var deletes, loads atomic.Int32
server := &peerRESTServer{}
must(deleteUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
deletes.Add(1)
return server.DeleteUserHandler(r)
}))
must(loadUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
loads.Add(1)
// LoadUserHandler delegates to this same cache reload method.
if err := sibling.UserNotificationHandler(ctx, r.Get(peerRESTUser), regUser); err != nil {
return grid.NoPayload{}, grid.NewRemoteErr(err)
}
return grid.NoPayload{}, nil
}))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
must(err)
prev := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) },
}}}
t.Cleanup(func() { globalNotificationSys = prev })
delayed := deleted.Add(time.Minute)
revoke(delayed)
t.Logf("sibling notifications after a retained revocation: destructive=%d reload=%d", deletes.Load(), loads.Load())
if deletes.Load() != 0 {
t.Errorf("destructive sibling delete sent: %d", deletes.Load())
}
if loads.Load() == 0 {
t.Errorf("retained revocation did not notify the sibling")
}
if _, ok := sibling.GetUser(child.AccessKey); ok {
t.Error("sibling still resolves revoked child")
}
if _, ok := sibling.GetUser(user); !ok {
t.Error("sibling lost live parent")
}
if _, ok := sys.store.GetUser(user); !ok {
t.Error("live parent lost")
}
if _, ok := sys.store.GetUser(child.AccessKey); ok {
t.Error("revoked child still resolves on the receiving node")
}
}
+535
View File
@@ -0,0 +1,535 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/namespace"
)
// Exercise the persisted IAM store and the same peer handler used by site heal.
// A delete must survive a cache reload and an older create arriving afterwards.
func TestIAMRevocationRejectsOfflineUser(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "offline-revoked-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "test-password-valid", Status: madmin.AccountEnabled}
created, err := globalIAMSys.CreateUser(ctx, user, req)
if err != nil {
t.Fatal(err)
}
if err = globalIAMSys.DeleteUser(ctx, user, false); err != nil {
t.Fatal(err)
}
if err = globalIAMSys.store.LoadIAMCache(ctx, false); err != nil {
t.Fatal(err)
}
if err = globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, created); err != nil {
t.Fatal(err)
}
if _, err = globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Fatalf("revoked user restored by old peer event: %v", err)
}
}
func TestIAMRevocationHealingContinuesAfterPeerRejectsDelete(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
if _, err := globalIAMSys.CreateUser(ctx, "heal-sync", req); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.CreateUser(ctx, "heal-deleted", req); err != nil {
t.Fatal(err)
}
if err := globalIAMSys.DeleteUser(ctx, "heal-deleted", false); err != nil {
t.Fatal(err)
}
p, err := globalIAMSys.store.GetPolicy("readwrite")
if err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.SetPolicy(ctx, "heal-new-policy", p); err != nil {
t.Fatal(err)
}
var liveUpdates atomic.Int32
peer := func(id string, rejectDelete bool) *httptest.Server {
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
switch {
case strings.HasSuffix(r.URL.Path, "/metainfo"):
_ = json.NewEncoder(w).Encode(madmin.SRInfo{DeploymentID: id})
case r.URL.Path == "/minio/admin/v3/site-replication/peer/iam-revisions":
if r.Method == http.MethodGet {
_ = json.NewEncoder(w).Encode(iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: id, Digest: "fixture"})
return
}
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(http.StatusBadRequest)
return
}
for _, item := range batch.Items {
if rejectDelete && iamDeletionPath(item.SRIAMItem) != "" {
w.WriteHeader(http.StatusForbidden)
_, _ = w.Write([]byte(`{"Code":"AccessDenied","Message":"delete rejected"}`))
return
}
if id == "healthy" && item.Type == madmin.SRIAMItemPolicy && item.Name == "heal-new-policy" && len(item.Policy) > 0 {
liveUpdates.Add(1)
}
}
_ = json.NewEncoder(w).Encode(iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: id, Digest: "fixture"})
default:
t.Errorf("unexpected peer request %s", r.URL.Path)
w.WriteHeader(http.StatusNotFound)
}
}))
}
healthy, rejected := peer("healthy", false), peer("rejected", true)
defer healthy.Close()
defer rejected.Close()
c := &SiteReplicationSys{enabled: true, state: srState{
ServiceAccountAccessKey: "heal-sync",
Peers: map[string]madmin.PeerInfo{
globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()},
"healthy": {Name: "healthy", DeploymentID: "healthy", Endpoint: healthy.URL},
"rejected": {Name: "rejected", DeploymentID: "rejected", Endpoint: rejected.URL},
},
}}
if err := c.healIAMSystem(ctx, obj); err == nil {
t.Fatal("deletion failure was not reported")
}
if liveUpdates.Load() == 0 {
t.Fatal("one peer rejecting a deletion blocked unrelated live IAM healing to a healthy peer")
}
}
func TestIAMRevocationLifecycle(t *testing.T) {
testIAMRevocationLifecycle(t, nil)
}
func TestIAMRevocationEtcdLifecycle(t *testing.T) {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if endpoint == "" {
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
connection, err := etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: 5 * time.Second})
if err != nil {
t.Fatal(err)
}
defer connection.Close()
// The facade borrows the connection's services. Close the owning client,
// not namespace.Watcher while IAM's canceled watch loop is winding down.
ctx, cancel := context.WithCancel(connection.Ctx())
defer cancel()
client := etcd.NewCtxClient(ctx, etcd.WithZapLogger(connection.GetLogger()))
prefix := fmt.Sprintf("/silo-revocation-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(connection.KV, prefix)
client.Watcher = namespace.NewWatcher(connection.Watcher, prefix)
client.Lease = connection.Lease
testIAMRevocationLifecycle(t, client)
}
func testIAMRevocationLifecycle(t *testing.T, client *etcd.Client) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
disk := disks[0]
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err == nil {
initAllSubsystems(ctx)
globalIAMSys.Init(ctx, obj, client, 2*time.Second)
}
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
sys, peer := globalIAMSys, &globalSiteReplicationSys
must := func(t *testing.T, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
reload := func(t *testing.T) { t.Helper(); must(t, sys.store.LoadIAMCache(ctx, false)) }
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
createUser := func(t *testing.T, name string) {
t.Helper()
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, origin))
}
assertAbsent := func(t *testing.T, name string) {
t.Helper()
if _, ok := sys.GetUser(ctx, name); ok {
t.Fatalf("revoked credential %s is usable", name)
}
}
t.Run("replay after recreation", func(t *testing.T) {
testIAMRevocationReplayAfterRecreation(ctx, t, sys)
})
t.Run("origin timestamp and recreation", func(t *testing.T) {
name := "revocation-recreate"
createUser(t, name)
ui, ok := sys.store.GetUser(name)
if !ok || !ui.UpdatedAt.Equal(origin) {
t.Fatalf("origin time changed: %v", ui.UpdatedAt)
}
must(t, sys.DeleteUser(ctx, name, false))
reload(t)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, time.Time{}))
assertAbsent(t, name)
newTime := UTCNow().Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, newTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}, origin.Add(time.Second)))
reload(t)
ui, ok = sys.store.GetUser(name)
if !ok || !ui.UpdatedAt.Equal(newTime) || ui.RevokedBefore.IsZero() {
t.Fatalf("newer recreation lost, or deletion boundary missing: present=%v", ok)
}
})
t.Run("groups policies and mappings", func(t *testing.T) {
user, group, name := "revocation-member", "revocation-group", "revocation-policy"
createUser(t, user)
p, err := sys.store.GetPolicy("readwrite")
must(t, err)
must(t, peer.PeerAddPolicyHandler(ctx, name, &p, origin))
add := &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}
must(t, peer.PeerGroupInfoChangeHandler(ctx, add, origin))
for _, isGroup := range []bool{false, true} {
entity := user
if isGroup {
entity = group
}
mp := &madmin.SRPolicyMapping{UserOrGroup: entity, Policy: name, UserType: int(regUser), IsGroup: isGroup}
must(t, peer.PeerPolicyMappingHandler(ctx, mp, origin))
_, err = sys.PolicyDBSet(ctx, entity, "", regUser, isGroup)
must(t, err)
must(t, peer.PeerPolicyMappingHandler(ctx, mp, origin))
if _, ok := sys.store.GetMappedPolicy(entity, isGroup); ok {
t.Fatal("old grant restored")
}
}
// This receiver never saw the member-removal event preceding deletion.
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, IsRemove: true}}, UTCNow()))
must(t, sys.DeletePolicy(ctx, name, true))
reload(t)
must(t, peer.PeerGroupInfoChangeHandler(ctx, add, origin))
must(t, peer.PeerAddPolicyHandler(ctx, name, &p, origin))
if _, err = sys.GetGroupDescription(group); !errors.Is(err, errNoSuchGroup) {
t.Fatalf("group restored: %v", err)
}
if _, err = sys.store.GetPolicyDoc(name); !errors.Is(err, errNoSuchPolicy) {
t.Fatalf("policy restored: %v", err)
}
paths, err := sys.store.listIAMConfigPaths(ctx)
must(t, err)
found := make(map[string]bool)
for _, path := range paths {
r, err := loadIAMRevision(ctx, sys.store, path)
must(t, err)
if item, ok := iamDeletionItem(path, r); ok {
found[iamDeletionPath(item)] = true
if item.UpdatedAt.IsZero() {
t.Fatal("undated delete replay")
}
}
}
for _, path := range []string{getGroupInfoPath(group), getPolicyDocPath(name), getMappedPolicyPath(user, regUser, false), getMappedPolicyPath(group, regUser, true)} {
if !found[path] {
t.Errorf("deletion missing from heal: %s", path)
}
}
})
t.Run("parent revokes service accounts and STS", func(t *testing.T) {
parent := "revocation-parent"
createUser(t, parent)
svc, svcAt, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: "revocation-service", secretKey: "valid-service-password"})
must(t, err)
secret, err := getTokenSigningKey()
must(t, err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
sts.ParentUser = parent
_, err = sys.SetTempUser(withIAMReplicationTime(ctx, origin), sts.AccessKey, sts, "readwrite")
must(t, err)
must(t, sys.DeleteUser(ctx, parent, false))
reload(t)
assertAbsent(t, parent)
assertAbsent(t, svc.AccessKey)
assertAbsent(t, sts.AccessKey)
// Recreate the parent, then deliver old child events from the offline site.
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, UTCNow()))
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: svc.AccessKey, SecretKey: svc.SecretKey}}, svcAt))
must(t, peer.PeerSTSAccHandler(ctx, &madmin.SRSTSCredential{AccessKey: sts.AccessKey, SecretKey: sts.SecretKey, ParentUser: parent, SessionToken: sts.SessionToken, ParentPolicyMapping: "readwrite"}, origin))
reload(t)
assertAbsent(t, svc.AccessKey)
assertAbsent(t, sts.AccessKey)
// A freshly issued credential is still supported after deliberate recreation.
_, _, err = sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "new-service", secretKey: "valid-service-password"})
must(t, err)
if _, ok := sys.GetUser(ctx, "new-service"); !ok {
t.Fatal("fresh service account rejected")
}
})
t.Run("delete before first create", func(t *testing.T) {
name := "revocation-unseen"
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}, UTCNow()))
createUser(t, name)
assertAbsent(t, name)
svc := "unseen-service"
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: svc}}, UTCNow()))
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "revocation-recreate", AccessKey: svc, SecretKey: "valid-service-password"}}, origin))
assertAbsent(t, svc)
})
t.Run("recreation arrives before revocation", func(t *testing.T) {
parent := "reordered-parent"
createUser(t, parent)
child, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: "reordered-child", secretKey: "valid-service-password"})
must(t, err)
newTime, deleteTime := UTCNow(), origin.Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, newTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
assertAbsent(t, child.AccessKey)
reload(t)
assertAbsent(t, child.AccessKey)
u, ok := sys.GetUser(ctx, parent)
if !ok || !u.UpdatedAt.Equal(newTime) || !u.RevokedBefore.Equal(deleteTime) {
t.Fatal("reordered revocation damaged the new parent or lost its boundary")
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(parent, regUser))
must(t, err)
item, ok := iamDeletionItem(getUserIdentityPath(parent, regUser), r)
if !ok || !item.UpdatedAt.Equal(deleteTime) {
t.Fatal("recreation erased deletion replay")
}
})
t.Run("user cleanup does not supersede group deletion", func(t *testing.T) {
user, group := "cascade-user", "cascade-group"
createUser(t, user)
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}, origin))
// On the origin site the group was removed before the user, but the
// recovering receiver processes those independent events in reverse.
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, IsRemove: true}}, origin.Add(time.Minute)))
reload(t)
if _, err := sys.GetGroupDescription(group); !errors.Is(err, errNoSuchGroup) {
t.Fatalf("deleted group survived reordered cleanup: %v", err)
}
groups, err := sys.ListGroups(ctx)
must(t, err)
for _, name := range groups {
if name == group {
t.Fatal("deleted group listed")
}
}
})
t.Run("parent revocation covers later updates to existing children", func(t *testing.T) {
parent, key := "late-update-parent", "late-update-child"
createUser(t, parent)
_, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: key, secretKey: "valid-service-password"})
must(t, err)
// This site missed the deletion and subsequently edited an old child.
_, err = sys.UpdateServiceAccount(withIAMReplicationTime(ctx, origin.Add(2*time.Minute)), key, updateServiceAccountOpts{description: "edited while the peer was offline"})
must(t, err)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, origin.Add(time.Minute)))
assertAbsent(t, key)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, origin.Add(3*time.Minute)))
reload(t)
assertAbsent(t, key)
})
t.Run("old generation cannot return with a newer event timestamp", func(t *testing.T) {
parent := "generation-parent"
createUser(t, parent)
deleteTime := origin.Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, origin.Add(2*time.Minute)))
// Another offline site issued this child under the original parent,
// after this site's delete/recreate. Wall-clock ordering cannot identify it.
late := origin.Add(3 * time.Minute)
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: "old-gen-service", SecretKey: "valid-service-password"}}, late))
secret, err := getTokenSigningKey()
must(t, err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
must(t, peer.PeerSTSAccHandler(ctx, &madmin.SRSTSCredential{AccessKey: sts.AccessKey, SecretKey: sts.SecretKey, ParentUser: parent, SessionToken: sts.SessionToken}, late))
reload(t)
assertAbsent(t, "old-gen-service")
assertAbsent(t, sts.AccessKey)
// A local issuer knows the new boundary and signs it into both kinds
// of child. Untrusted inherited claims cannot select that boundary.
child, _, err := sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "new-gen-service", secretKey: "valid-service-password", claims: map[string]any{iamParentRevocationClaim: "forged"}})
must(t, err)
newClaims := map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}
must(t, setIAMParentRevocationClaim(ctx, sys.store, parent, newClaims))
fresh, err := auth.GetNewCredentialsWithMetadata(newClaims, secret)
must(t, err)
fresh.ParentUser = parent
_, err = sys.SetTempUser(ctx, fresh.AccessKey, fresh, "")
must(t, err)
reload(t)
// A periodic reload retains the STS cache. Explicitly clear it to
// exercise the cold credential load performed after process restart.
cache := sys.store.lock()
cache.iamSTSAccountsMap = make(map[string]UserIdentity)
sys.store.unlock()
for _, key := range []string{child.AccessKey, fresh.AccessKey} {
u, ok := sys.GetUser(ctx, key)
if !ok || !iamCredentialSurvivesRevocation(u.Credentials, deleteTime) {
t.Fatalf("new-generation credential %s rejected", key)
}
}
})
t.Run("late revocation preserves proven new-generation children", func(t *testing.T) {
for _, recreateFirst := range []bool{false, true} {
parent := fmt.Sprintf("gen-parent-%t", recreateFirst)
createUser(t, parent)
deleteTime, createTime := origin.Add(time.Minute), origin.Add(2*time.Minute)
if recreateFirst {
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, createTime))
}
key := fmt.Sprintf("gen-child-%t", recreateFirst)
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: key, SecretKey: "valid-service-password", Claims: map[string]any{iamParentRevocationClaim: deleteTime.Format(time.RFC3339Nano)}}}, createTime.Add(time.Second)))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
if !recreateFirst {
assertAbsent(t, key)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, createTime))
}
reload(t)
if _, ok := sys.GetUser(ctx, key); !ok {
t.Fatal("late revocation deleted a child issued by the recreated parent")
}
}
})
t.Run("cold loading preserves site-signed STS", func(t *testing.T) {
parent := "cold-sts-parent"
createUser(t, parent)
secret := "site-signing-key-valid"
_, _, err := sys.NewServiceAccount(ctx, globalActiveCred.AccessKey, nil, newServiceAccountOpts{
accessKey: siteReplicatorSvcAcc, secretKey: secret, allowSiteReplicatorAccount: true,
})
must(t, err)
setReplication := func(enabled bool) {
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = enabled
globalSiteReplicationSys.Unlock()
globalSiteReplicatorCred.Set("")
}
setReplication(true)
defer setReplication(false)
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
cred.ParentUser = parent
_, err = sys.SetTempUser(ctx, cred.AccessKey, cred, "")
must(t, err)
// IAM can load before the site replication manager during startup.
// A signing key that is not available yet must not delete live tokens.
setReplication(false)
for range 3 {
unverified := make(map[string]UserIdentity)
_ = sys.store.loadUser(ctx, cred.AccessKey, stsUser, unverified)
if _, ok := unverified[cred.AccessKey]; ok {
t.Fatal("accepted STS before the signing key became available")
}
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(cred.AccessKey, stsUser))
must(t, err)
if r.Credentials.SessionToken == "" {
t.Fatal("cold IAM load physically deleted a non-expired site-signed STS credential")
}
setReplication(true)
loaded := make(map[string]UserIdentity)
must(t, sys.store.loadUser(ctx, cred.AccessKey, stsUser, loaded))
if _, ok := loaded[cred.AccessKey]; !ok {
t.Fatal("STS credential did not recover when the signing key became available")
}
})
t.Run("unverifiable STS stay denied and expired STS are removed", func(t *testing.T) {
parent := "invalid-sts-parent"
createUser(t, parent)
// Keep the etcd watcher from cleaning half of the fixture before the
// second record is seeded; this subtest exercises the loader directly.
sys.store.lock()
defer sys.store.unlock()
for _, expired := range []bool{false, true} {
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, "unavailable-test-signing-key")
must(t, err)
cred.ParentUser = parent
if expired {
cred.Expiration = UTCNow().Add(-time.Minute)
}
identityPath := getUserIdentityPath(cred.AccessKey, stsUser)
mappingPath := getMappedPolicyPath(cred.AccessKey, stsUser, false)
// Seed disk directly to exercise loading, including existing records
// whose key is unknown. The write API should not accept such tokens.
must(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: cred, UpdatedAt: UTCNow()}, identityPath))
must(t, sys.store.saveIAMConfig(ctx, &MappedPolicy{Version: 1, Policies: "readwrite"}, mappingPath))
loaded := make(map[string]UserIdentity)
_ = sys.store.loadUser(ctx, cred.AccessKey, stsUser, loaded)
if _, ok := loaded[cred.AccessKey]; ok {
t.Fatalf("invalid STS accepted, expired=%t", expired)
}
for _, path := range []string{identityPath, mappingPath} {
var record map[string]any
err := sys.store.loadIAMConfig(ctx, &record, path)
if expired {
if !errors.Is(err, errConfigNotFound) {
t.Fatalf("expired STS data not cleaned up at %s: %v", path, err)
}
} else {
must(t, err)
}
}
}
})
}
+239
View File
@@ -0,0 +1,239 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
)
// The receiver missed a deletion and still has the previous service key.
// A later full snapshot must replace it, including when the owner changed.
func TestIAMServiceAccountRecreation(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
for _, parent := range []string{"old-owner", "new-owner"} {
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), parent, madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
}
const key = "reusable-service"
old := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "old-owner", AccessKey: key, SecretKey: "old-service-password"}}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
oldIdentity, _ := sys.store.GetUser(key)
_, err := sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), key, "readwrite", svcUser, false)
mustIAM(t, err)
boundary, newer := origin.Add(time.Minute), origin.Add(2*time.Minute)
fresh := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemSvcAcc, UpdatedAt: newer, SvcAccChange: &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "new-owner", AccessKey: key, SecretKey: "new-service-password", Status: auth.AccountOff}}}, RevokedBefore: boundary}
mustIAM(t, applyIAMReplicationItem(ctx, fresh))
// A sibling may have missed the notification of the committed
// replacement. An equal-version retry must refresh that cache too.
staleCache := sys.store.lock()
staleCache.iamUsersMap[key] = oldIdentity
sys.store.unlock()
mustIAM(t, applyIAMReplicationItem(ctx, fresh)) // duplicate delivery is acknowledged
if current, _ := sys.store.GetUser(key); current.Credentials.SecretKey != fresh.SvcAccChange.Create.SecretKey {
t.Fatal("duplicate snapshot acknowledged without refreshing the stale cache")
}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: key}}, boundary))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
u, ok := sys.store.GetUser(key)
if !ok || u.Credentials.SecretKey != fresh.SvcAccChange.Create.SecretKey || u.Credentials.ParentUser != "new-owner" || u.Credentials.Status != auth.AccountOff || !u.UpdatedAt.Equal(newer) || !u.RevokedBefore.Equal(boundary) {
t.Fatal("recreation did not retain the new identity, disabled status, source version and revocation")
}
if _, ok := sys.GetUser(ctx, key); ok {
t.Fatal("replicated disabled service can authenticate")
}
cache := sys.store.rlock()
_, mapped := cache.cachedMappedPolicy(key, svcUser, false)
sys.store.runlock()
if mapped {
t.Fatal("recreated service inherited an older mapping")
}
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), key, "readwrite", svcUser, false)
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("old service mapping replay was accepted: %v", err)
}
_, _, err = sys.NewServiceAccount(ctx, "new-owner", nil, newServiceAccountOpts{accessKey: key, secretKey: "local-service-password"})
if !errors.Is(err, errIAMServiceAccountNotAllowed) {
t.Fatalf("local duplicate creation must remain rejected: %v", err)
}
// Outbound snapshots must carry the retained service boundary too.
out, err := globalSiteReplicationSys.replicationItem(ctx, fresh.SRIAMItem)
mustIAM(t, err)
if !out.RevokedBefore.Equal(boundary) {
t.Fatal("outbound service snapshot lost its revocation")
}
})
}
}
func TestIAMServiceAccountReplicationRejectsOtherCredentialKinds(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "builtin-collision", madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
secret, err := getTokenSigningKey()
mustIAM(t, err)
token, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: "builtin-collision"}, secret)
mustIAM(t, err)
token.ParentUser = "builtin-collision"
_, err = sys.SetTempUser(ctx, token.AccessKey, token, "")
mustIAM(t, err)
for _, key := range []string{"builtin-collision", token.AccessKey} {
_, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, UTCNow().Add(time.Minute)), "another-owner", nil, newServiceAccountOpts{accessKey: key, secretKey: "valid-service-password"})
if !errors.Is(err, errIAMServiceAccountNotAllowed) {
t.Fatalf("service replication replaced another credential kind: %v", err)
}
}
}
// SR configuration can be temporarily unreadable even though a service token
// is signed with its own valid secret. Do not acknowledge a failed cache load.
func TestIAMServiceAccountRetryReportsClaimLoadFailure(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
globalSiteReplicatorCred.RLock()
previousSigningKey := globalSiteReplicatorCred.secretKey
globalSiteReplicatorCred.RUnlock()
globalSiteReplicatorCred.Set("")
t.Cleanup(func() { globalSiteReplicatorCred.Set(previousSigningKey) })
_, err := sys.CreateUser(ctx, "retry-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
opts := newServiceAccountOpts{accessKey: "retry-service", secretKey: "valid-service-password"}
_, _, err = sys.NewServiceAccount(ctx, "retry-owner", nil, opts)
mustIAM(t, err)
old, _ := sys.store.GetUser(opts.accessKey)
opts.secretKey = "replacement-service-password"
at, err := sys.UpdateServiceAccount(ctx, opts.accessKey, updateServiceAccountOpts{secretKey: opts.secretKey})
mustIAM(t, err)
cache := sys.store.lock()
cache.iamUsersMap[opts.accessKey] = old // Missed sibling notification.
sys.store.unlock()
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = true // No site-replicator credential is installed.
globalSiteReplicationSys.Unlock()
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, at), "retry-owner", nil, opts)
if err == nil {
t.Fatal("acknowledged service retry despite failed claims loading")
}
if _, ok := sys.store.GetUser(opts.accessKey); ok {
t.Fatal("failed cache refresh retained the superseded service secret")
}
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = false
globalSiteReplicationSys.Unlock()
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, at), "retry-owner", nil, opts)
mustIAM(t, err)
}
// A delayed snapshot still has its original absolute expiration. Reapplying
// the local minimum issuance lifetime would leave the old unexpired key alive.
func TestIAMServiceAccountReplicationPreservesExpiration(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
for _, action := range []string{"create", "update"} {
for _, expired := range []bool{false, true} {
name := backend + "/" + action + "/near_expiry"
if expired {
name = backend + "/" + action + "/expired"
}
t.Run(name, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
_, err := sys.CreateUser(ctx, "expiry-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
old := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "expiry-owner", AccessKey: "expiry-service", SecretKey: "old-service-password"}}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
expires := UTCNow().Add(time.Minute)
if expired {
expires = UTCNow().Add(-time.Minute)
}
change := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "expiry-owner", AccessKey: "expiry-service", SecretKey: "new-service-password", Expiration: &expires}}
if action == "update" {
change = &madmin.SRSvcAccChange{Update: &madmin.SRSvcAccUpdate{AccessKey: "expiry-service", SecretKey: "new-service-password", Expiration: &expires}}
}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, change, origin.Add(2*time.Minute)))
u, ok := sys.store.GetUser("expiry-service")
if !ok || u.Credentials.SecretKey != "new-service-password" || !u.Credentials.Expiration.Equal(expires) {
t.Fatal("delayed snapshot lost its new secret or absolute expiration")
}
_, allowed := sys.GetUser(ctx, "expiry-service")
if allowed == expired {
t.Fatal("credential validity disagrees with its absolute expiration")
}
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath("expiry-service", svcUser))
mustIAM(t, err)
if r.Credentials.SecretKey == old.Create.SecretKey || (expired && !r.Deleted) {
t.Fatal("old non-expiring credential returned after reload")
}
_, _, err = sys.NewServiceAccount(ctx, "expiry-owner", nil, newServiceAccountOpts{accessKey: "local-expiry", secretKey: "valid-service-password", expiration: &expires})
if !errors.Is(err, errInvalidSvcAcctExpiration) {
t.Fatalf("local issuance lifetime check changed: %v", err)
}
})
}
}
}
}
// Status-only summaries intentionally omit secrets. Different revisions must
// still trigger live healing, and disabled identities must be eligible sources.
func TestIAMServiceAccountHealingNewerSnapshot(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
_, err := sys.CreateUser(ctx, "heal-svc-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), "heal-svc-owner", nil, newServiceAccountOpts{accessKey: "heal-service", secretKey: "valid-service-password"})
mustIAM(t, err)
at, err := sys.UpdateServiceAccount(ctx, "heal-service", updateServiceAccountOpts{status: auth.AccountOff})
mustIAM(t, err)
var sent atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
return
}
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" || r.Method != http.MethodPut {
t.Errorf("unexpected request %s %s", r.Method, r.URL.Path)
w.WriteHeader(http.StatusNotFound)
return
}
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(http.StatusBadRequest)
return
}
for _, item := range batch.Items {
if item.SvcAccChange != nil && item.SvcAccChange.Create != nil && item.SvcAccChange.Create.AccessKey == "heal-service" && item.SvcAccChange.Create.Status == auth.AccountOff && item.UpdatedAt.Equal(at) {
sent.Add(1)
}
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "node", Instance: "boot", Digest: "fixture"}})
}))
defer server.Close()
peers := map[string]madmin.PeerInfo{globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()}, "remote": {Name: "remote", DeploymentID: "remote", Endpoint: server.URL}}
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "heal-svc-owner", Peers: peers}}
local := madmin.UserInfo{Status: madmin.AccountStatus(auth.AccountOff), UpdatedAt: at}
remote := local
remote.UpdatedAt = origin
if isUserInfoReplicated(2, 2, []madmin.UserInfo{local, remote}) {
t.Fatal("status-only summaries concealed different service revisions")
}
info := srStatusInfo{Sites: peers, UserStats: map[string]map[string]srUserStatsSummary{"heal-service": {globalDeploymentID(): {userInfo: srUserInfo{UserInfo: local}}, "remote": {SRUserStatsSummary: madmin.SRUserStatsSummary{UserInfoMismatch: true}, userInfo: srUserInfo{UserInfo: remote}}}}}
mustIAM(t, c.healUsers(ctx, obj, "heal-service", info))
if sent.Load() == 0 {
t.Fatal("disabled newer service snapshot was not healed")
}
}
+383 -221
View File
File diff suppressed because it is too large Load Diff
+65 -21
View File
@@ -162,6 +162,15 @@ func (sys *IAMSys) LoadUser(ctx context.Context, objAPI ObjectLayer, accessKey s
return sys.store.UserNotificationHandler(ctx, accessKey, userType)
}
// LoadUserAfterDelete reloads a parent's identity and cached dependents after a
// sibling committed a deletion. Each record may already have been recreated.
func (sys *IAMSys) LoadUserAfterDelete(ctx context.Context, accessKey string) error {
if !sys.Initialized() {
return errServerNotInitialized
}
return sys.store.UserDeletionNotificationHandler(ctx, accessKey)
}
// LoadServiceAccount - reloads a specific service account from backend disks or etcd.
func (sys *IAMSys) LoadServiceAccount(ctx context.Context, accessKey string) error {
if !sys.Initialized() {
@@ -596,10 +605,22 @@ func (sys *IAMSys) DeletePolicy(ctx context.Context, policyName string, notifyPe
return errServerNotInitialized
}
for _, v := range policy.DefaultPolicies {
if v.Name == policyName {
if err := checkConfig(ctx, globalObjectAPI, getPolicyDocPath(policyName)); err != nil && err == errConfigNotFound {
return fmt.Errorf("inbuilt policy `%s` not allowed to be deleted", policyName)
if _, replicated := iamReplicationTime(ctx); !replicated && notifyPeers {
for _, v := range policy.DefaultPolicies {
if v.Name == policyName {
var err error
if objectStore, ok := sys.store.IAMStorageAPI.(*IAMObjectStore); ok {
err = checkConfig(ctx, objectStore.objAPI, getPolicyDocPath(policyName))
} else {
var r iamRevision
err = sys.store.loadIAMConfig(ctx, &r, getPolicyDocPath(policyName))
}
if errors.Is(err, errConfigNotFound) {
return fmt.Errorf("inbuilt policy `%s` not allowed to be deleted", policyName)
}
if err != nil {
return err
}
}
}
}
@@ -705,20 +726,30 @@ func (sys *IAMSys) DeleteUser(ctx context.Context, accessKey string, notifyPeers
return errServerNotInitialized
}
if err := sys.store.DeleteUser(ctx, accessKey, regUser); err != nil {
err := sys.store.DeleteUser(ctx, accessKey, regUser)
var cleanupErr *iamCommittedCleanupError
retained := errors.Is(err, errIAMRevocationRetained)
if errors.As(err, &cleanupErr) {
retained = cleanupErr.retained
} else if err != nil && !retained {
return err
}
// Notify all other MinIO peers to delete user.
// Publish the committed state even when dependent cleanup must be retried.
if notifyPeers && !sys.HasWatcher() {
for _, nerr := range globalNotificationSys.DeleteUser(ctx, accessKey) {
if nerr.Err != nil {
logger.GetReqInfo(ctx).SetTags("peerAddress", nerr.Host.String())
iamLogIf(ctx, nerr.Err)
if retained {
sys.notifyForUser(ctx, accessKey, false)
} else {
for _, nerr := range globalNotificationSys.DeleteUser(ctx, accessKey) {
if nerr.Err != nil {
logger.GetReqInfo(ctx).SetTags("peerAddress", nerr.Host.String())
iamLogIf(ctx, nerr.Err)
}
}
}
}
if cleanupErr != nil {
return cleanupErr
}
return nil
}
@@ -1052,6 +1083,7 @@ type newServiceAccountOpts struct {
sessionPolicy *policy.Policy
accessKey string
secretKey string
status string // Used by replication snapshots; local creates default to enabled.
name, description string
expiration *time.Time
allowSiteReplicatorAccount bool // allow creating internal service account for site-replication.
@@ -1116,6 +1148,11 @@ func (sys *IAMSys) NewServiceAccount(ctx context.Context, parentUser string, gro
m[k] = v
}
}
if _, replicated := iamReplicationTime(ctx); !replicated {
if err := setIAMParentRevocationClaim(ctx, sys.store, parentUser, m); err != nil {
return auth.Credentials{}, time.Time{}, err
}
}
var accessKey, secretKey string
var err error
@@ -1134,12 +1171,19 @@ func (sys *IAMSys) NewServiceAccount(ctx context.Context, parentUser string, gro
cred.ParentUser = parentUser
cred.Groups = groups
cred.Status = string(auth.AccountOn)
switch opts.status {
case "", auth.AccountOn, string(madmin.AccountEnabled):
case auth.AccountOff, string(madmin.AccountDisabled):
cred.Status = auth.AccountOff
default:
return auth.Credentials{}, time.Time{}, errInvalidArgument
}
cred.Name = opts.name
cred.Description = opts.description
if opts.expiration != nil {
expirationInUTC := opts.expiration.UTC()
if err := validateSvcExpirationInUTC(expirationInUTC); err != nil {
if err := validateSvcExpirationInUTC(ctx, expirationInUTC); err != nil {
return auth.Credentials{}, time.Time{}, err
}
cred.Expiration = expirationInUTC
@@ -1370,7 +1414,7 @@ func (sys *IAMSys) DeleteServiceAccount(ctx context.Context, accessKey string, n
}
sa, ok := sys.store.GetUser(accessKey)
if !ok || !sa.Credentials.IsServiceAccount() {
if _, replicated := iamReplicationTime(ctx); (!ok || !sa.Credentials.IsServiceAccount()) && !replicated {
return nil
}
@@ -1474,8 +1518,8 @@ func (sys *IAMSys) purgeExpiredCredentialsForExternalSSO(ctx context.Context) {
}
}
// We ignore any errors
_ = sys.store.DeleteUsers(ctx, expiredUsers)
// Keep failed revocations visible so the next purge can retry.
iamLogIf(ctx, sys.store.DeleteUsers(ctx, expiredUsers))
}
// purgeExpiredCredentialsForLDAP - validates if local credentials are still
@@ -1503,8 +1547,8 @@ func (sys *IAMSys) purgeExpiredCredentialsForLDAP(ctx context.Context) {
return
}
// We ignore any errors
_ = sys.store.DeleteUsers(ctx, expiredUsers)
// Keep failed revocations visible so the next purge can retry.
iamLogIf(ctx, sys.store.DeleteUsers(ctx, expiredUsers))
}
// updateGroupMembershipsForLDAP - updates the list of groups associated with the credential.
@@ -1925,12 +1969,12 @@ func (sys *IAMSys) RemoveUsersFromGroup(ctx context.Context, group string, membe
}
updatedAt, err = sys.store.RemoveUsersFromGroup(ctx, group, members)
if err != nil {
var cleanupErr *iamCommittedCleanupError
if err != nil && !errors.As(err, &cleanupErr) {
return updatedAt, err
}
sys.notifyForGroup(ctx, group)
return updatedAt, nil
return updatedAt, err
}
// SetGroupStatus - enable/disabled a group
+14
View File
@@ -34,6 +34,10 @@ const (
sinceLastSyncMillis = "since_last_sync_millis"
syncFailures = "sync_failures"
syncSuccesses = "sync_successes"
revocationRecords = "revocation_records"
revocationHealFailures = "revocation_heal_failures"
revocationHealDurationMillis = "revocation_heal_duration_millis"
revocationHealLastSuccess = "revocation_heal_last_success_timestamp_seconds"
)
var (
@@ -47,10 +51,20 @@ var (
sinceLastSyncMillisMD = NewCounterMD(sinceLastSyncMillis, "Time (in milliseconds) since last successful IAM data sync.")
syncFailuresMD = NewCounterMD(syncFailures, "Number of failed IAM data syncs since server start.")
syncSuccessesMD = NewCounterMD(syncSuccesses, "Number of successful IAM data syncs since server start.")
revocationRecordsMD = NewGaugeMD(revocationRecords, "Retained IAM deletion records and revocation boundaries in this node's index.")
revocationHealFailuresMD = NewCounterMD(revocationHealFailures, "Failed IAM revocation convergence passes since server start.")
revocationHealDurationMillisMD = NewGaugeMD(revocationHealDurationMillis, "Duration of the last IAM revocation convergence pass in milliseconds.")
revocationHealLastSuccessMD = NewGaugeMD(revocationHealLastSuccess, "Unix timestamp of the last successful IAM revocation convergence pass.")
)
// loadClusterIAMMetrics - `MetricsLoaderFn` for cluster IAM metrics.
func loadClusterIAMMetrics(_ context.Context, m MetricValues, _ *metricsCache) error {
if globalIAMSys.Initialized() {
m.Set(revocationRecords, float64(globalIAMSys.store.revisionIndex().count()))
}
m.Set(revocationHealFailures, float64(globalSiteReplicationSys.iamRevisionMetrics.healFailures.Load()))
m.Set(revocationHealDurationMillis, float64(globalSiteReplicationSys.iamRevisionMetrics.healDurationMillis.Load()))
m.Set(revocationHealLastSuccess, float64(globalSiteReplicationSys.iamRevisionMetrics.healLastSuccess.Load()))
m.Set(lastSyncDurationMillis, float64(atomic.LoadUint64(&globalIAMSys.LastRefreshDurationMilliseconds)))
pluginAuthNMetrics := globalAuthNPlugin.Metrics()
m.Set(pluginAuthnServiceFailedRequestsMinute, float64(pluginAuthNMetrics.FailedRequests))
+4
View File
@@ -323,6 +323,10 @@ func newMetricGroups(r *prometheus.Registry) *metricsV3Collection {
sinceLastSyncMillisMD,
syncFailuresMD,
syncSuccessesMD,
revocationRecordsMD,
revocationHealFailuresMD,
revocationHealDurationMillisMD,
revocationHealLastSuccessMD,
},
loadClusterIAMMetrics,
)
+113
View File
@@ -0,0 +1,113 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"encoding/base64"
"net/http"
"reflect"
"strconv"
"strings"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
)
func TestPutOptsFromHeadersReplicationTimestamps(t *testing.T) {
stamp := time.Date(2026, 9, 15, 1, 2, 3, 123456789, time.UTC)
context := base64.StdEncoding.EncodeToString([]byte(`{"purpose":"tag-replication"}`))
for _, encryption := range []struct {
name string
headers map[string]string
}{
{name: "none"},
{name: "SSE-S3", headers: map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}},
{name: "SSE-KMS", headers: map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS}},
{name: "SSE-KMS-context", headers: map[string]string{
xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS, xhttp.AmzServerSideEncryptionKmsID: "tag-replication-key",
xhttp.AmzServerSideEncryptionKmsContext: context,
}},
{name: "SSE-C", headers: ssecKeyHeaders([]byte("01234567890123456789012345678901"), false)},
} {
t.Run(encryption.name, func(t *testing.T) {
for _, trusted := range []bool{false, true} {
t.Run("trusted="+strconv.FormatBool(trusted), func(t *testing.T) {
for _, tagging := range []struct {
name, header string
want time.Time
invalid bool
}{
{name: "absent"},
{name: "nanoseconds", header: stamp.Format(time.RFC3339Nano), want: stamp},
{name: "offset-whitespace", header: " " + stamp.In(time.FixedZone("UTC+8", 8*60*60)).Format(time.RFC3339Nano) + " ", want: stamp},
{name: "invalid", header: "not-a-timestamp", invalid: true},
} {
t.Run(tagging.name, func(t *testing.T) {
for _, metadata := range []map[string]string{nil, {"x-amz-meta-test": "kept"}} {
hdr := make(http.Header)
wantEncryption := make(http.Header)
for key, value := range encryption.headers {
hdr.Set(key, value)
wantEncryption.Set(key, value)
}
hdr.Set(xhttp.MinIOSourceTaggingTimestamp, tagging.header)
hdr.Set(xhttp.MinIOSourceMTime, stamp.Add(-time.Hour).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceObjectRetentionTimestamp, stamp.Add(-time.Minute).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceObjectLegalHoldTimestamp, stamp.Add(-time.Second).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceETag, "source-etag")
opts, err := putOptsFromHeaders(t.Context(), hdr, metadata, trusted)
if trusted && tagging.invalid {
if err == nil || !strings.Contains(err.Error(), xhttp.MinIOSourceTaggingTimestamp) {
t.Fatalf("malformed trusted timestamp: got %v", err)
}
continue
}
if err != nil {
t.Fatal(err)
}
wantTag, wantMTime, wantRetention, wantLegalhold, wantETag := time.Time{}, time.Time{}, time.Time{}, time.Time{}, ""
if trusted {
wantTag, wantMTime = tagging.want, stamp.Add(-time.Hour)
wantRetention, wantLegalhold, wantETag = stamp.Add(-time.Minute), stamp.Add(-time.Second), "source-etag"
}
if !opts.ReplicationSourceTaggingTimestamp.Equal(wantTag) {
t.Errorf("tag timestamp=%s, want %s", opts.ReplicationSourceTaggingTimestamp, wantTag)
}
if !opts.MTime.Equal(wantMTime) || !opts.ReplicationSourceRetentionTimestamp.Equal(wantRetention) ||
!opts.ReplicationSourceLegalholdTimestamp.Equal(wantLegalhold) || opts.PreserveETag != wantETag || opts.ReplicationRequest != trusted {
t.Error("other source fields did not preserve the replication trust boundary")
}
if opts.UserDefined == nil || (metadata != nil && !reflect.DeepEqual(opts.UserDefined, metadata)) {
t.Errorf("metadata=%v, want nonnil map preserving %v", opts.UserDefined, metadata)
}
gotEncryption := make(http.Header)
if opts.ServerSideEncryption != nil {
opts.ServerSideEncryption.Marshal(gotEncryption)
}
if !reflect.DeepEqual(gotEncryption, wantEncryption) {
t.Errorf("SSE headers=%v, want %v", gotEncryption, wantEncryption)
}
}
})
}
})
}
})
}
}
+3 -3
View File
@@ -452,11 +452,11 @@ func putOptsFromHeaders(ctx context.Context, hdr http.Header, metadata map[strin
MTime: mtime,
PreserveETag: etag,
ReplicationRequest: trustedReplication,
// The Object Lock timestamps order replicated retention and legal
// hold updates. Dropping them here would leave every update on an
// SSE-KMS destination unordered.
// These timestamps order replicated retention, legal hold and tagging
// updates on an SSE-KMS destination.
ReplicationSourceLegalholdTimestamp: lholdtimestmp,
ReplicationSourceRetentionTimestamp: retaintimestmp,
ReplicationSourceTaggingTimestamp: taggingtimestmp,
}
return op, nil
}
+143
View File
@@ -0,0 +1,143 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// TestAPICopyObjectReplicaTaggingTimestampUnderKMS covers signed replica COPY
// requests through encryption, metadata replacement and disk persistence. Both
// the single-disk and 16-disk fixtures are single-pool backends.
func TestAPICopyObjectReplicaTaggingTimestampUnderKMS(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPICopyObjectReplicaTaggingTimestampUnderKMS})
}
func testAPICopyObjectReplicaTaggingTimestampUnderKMS(obj ObjectLayer, instance, bucket string, router http.Handler, creds auth.Credentials, t *testing.T) {
// Ignore the host free-space percentage while retaining real disk I/O.
for _, pool := range obj.(*erasureServerPools).serverPools {
for _, set := range pool.sets {
original := set.getDisks
disks := append([]StorageAPI(nil), original()...)
for i := range disks {
disks[i] = tagTestCapacityDisk{StorageAPI: disks[i]}
}
set.getDisks = func() []StorageAPI { return disks }
defer func() { set.getDisks = original }()
}
}
oldKMS, oldAuto := GlobalKMS, globalAutoEncryption
GlobalKMS = kms.NewStub("replica-tags-key")
globalAutoEncryption = false
defer func() { GlobalKMS, globalAutoEncryption = oldKMS, oldAuto }()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
const tsKey = ReservedMetadataPrefixLower + TaggingTimestamp
stamp := time.Date(2026, 9, 15, 1, 0, 0, 123456789, time.UTC)
for _, mode := range []string{"none", "explicit-sse-s3", "explicit-kms", "auto-kms", "bucket-kms"} {
t.Run(instance+"/"+mode, func(t *testing.T) {
globalAutoEncryption = mode == "auto-kms"
if mode == "bucket-kms" {
sseXML := []byte(`<ServerSideEncryptionConfiguration xmlns="http://s3.amazonaws.com/doc/2006-03-01/"><Rule><ApplyServerSideEncryptionByDefault><SSEAlgorithm>aws:kms</SSEAlgorithm><KMSMasterKeyID>replica-tags-key</KMSMasterKeyID></ApplyServerSideEncryptionByDefault></Rule></ServerSideEncryptionConfiguration>`)
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketSSEConfig, sseXML); err != nil {
t.Fatal(err)
}
}
const data = "encrypted replica copy remains readable"
oi, err := obj.PutObject(t.Context(), bucket, mode, mustGetPutObjReader(t, bytes.NewReader([]byte(data)), int64(len(data)), "", ""), ObjectOptions{
Versioned: true, UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=old", tsKey: stamp.Format(time.RFC3339Nano)},
})
if err != nil {
t.Fatal(err)
}
for _, event := range []struct {
name, tags, wantTags string
delta, wantDelta time.Duration
missingTimestamp bool
}{
{"newer", "key=new", "key=new", 2, 2, false},
{"stale", "key=stale", "key=new", 1, 2, false},
{"duplicate", "key=new", "key=new", 2, 2, false},
{"newer-again", "key=latest", "key=latest", 3, 3, false},
{"missing-timestamp", "key=unordered", "key=latest", 0, 3, true},
} {
headers := map[string]string{
xhttp.AmzCopySource: "/" + bucket + "/" + mode + "?versionId=" + oi.VersionID,
xhttp.AmzMetadataDirective: "REPLACE", xhttp.AmzTagDirective: "REPLACE",
xhttp.AmzObjectTagging: event.tags, xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA", xhttp.MinIOSourceTaggingTimestamp: stamp.Add(event.delta).Format(time.RFC3339Nano),
xhttp.MinIOSourceMTime: oi.ModTime.Format(time.RFC3339Nano), xhttp.MinIOSourceETag: oi.ETag,
}
if event.missingTimestamp {
delete(headers, xhttp.MinIOSourceTaggingTimestamp)
}
if mode == "explicit-sse-s3" {
headers[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionAES
}
if mode == "explicit-kms" {
headers[xhttp.AmzServerSideEncryption] = "aws:kms"
headers[xhttp.AmzServerSideEncryptionKmsID] = "replica-tags-key"
}
req, err := newTestSignedRequestV4(http.MethodPut, "/"+bucket+"/"+mode+"?versionId="+oi.VersionID, 0, nil, creds.AccessKey, creds.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
if w.Code != http.StatusOK {
t.Fatalf("%s: COPY %d %s", event.name, w.Code, w.Body.String())
}
got, err := obj.GetObjectInfo(t.Context(), bucket, mode, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
t.Logf("%s: tags=%q timestamp=%q kms=%v", event.name, got.UserTags, got.UserDefined[tsKey], crypto.S3KMS.IsEncrypted(got.UserDefined))
if got.UserTags != event.wantTags || got.UserDefined[tsKey] != stamp.Add(event.wantDelta).Format(time.RFC3339Nano) {
t.Errorf("%s: incorrect persisted tags/timestamp", event.name)
}
wantKMS := mode == "explicit-kms" || mode == "auto-kms" || mode == "bucket-kms"
if crypto.S3KMS.IsEncrypted(got.UserDefined) != wantKMS || crypto.S3.IsEncrypted(got.UserDefined) != (mode == "explicit-sse-s3") {
t.Errorf("%s: unexpected destination encryption", event.name)
}
if got.VersionID != oi.VersionID {
t.Errorf("%s: version=%q, want %q", event.name, got.VersionID, oi.VersionID)
}
req, err = newTestSignedRequestV4(http.MethodGet, "/"+bucket+"/"+mode+"?versionId="+oi.VersionID, 0, nil, creds.AccessKey, creds.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
w = httptest.NewRecorder()
router.ServeHTTP(w, req)
if w.Code != http.StatusOK || w.Body.String() != data {
t.Fatalf("%s: GET %d %q", event.name, w.Code, w.Body.String())
}
}
})
}
}
+4 -1
View File
@@ -240,7 +240,10 @@ func checkPreconditionsPUT(ctx context.Context, w http.ResponseWriter, r *http.R
// updated. The predicate is the incoming request's restored SSE-C metadata,
// not what the destination happens to hold.
ssecReplica := isReplicaTrusted(r.Context()) && crypto.SSEC.IsEncrypted(opts.UserDefined)
if etagMatch && vidMatch && !ssecReplica {
// Matching content does not imply that its tag revision was delivered.
// Keep client preconditions above; relax only the internal duplicate check.
newerTags := isReplicaTrusted(r.Context()) && olderThan(objInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], opts.ReplicationSourceTaggingTimestamp)
if etagMatch && vidMatch && !ssecReplica && !newerTags {
writeHeaders()
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(ErrPreconditionFailed), r.URL)
return true
+30 -19
View File
@@ -1797,6 +1797,7 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
// source timestamp is newer than the stored one, and a stale update must
// leave the stored state in place instead of erasing it.
storedLock := storedObjectLockState(srcInfo.UserDefined)
storedTagTimestamp := srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]
srcInfo.UserDefined, err = getCpObjMetadataFromHeader(ctx, r, srcInfo.UserDefined, allowReplicationMetadata)
if err != nil {
@@ -1814,23 +1815,29 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
}
}
if objTags != "" {
lastTaggingTimestamp := srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]
if dstOpts.ReplicationRequest {
srcTimestamp := dstOpts.ReplicationSourceTaggingTimestamp
if !srcTimestamp.IsZero() {
ondiskTimestamp, err := time.Parse(time.RFC3339Nano, lastTaggingTimestamp)
// update tagging metadata only if replica timestamp is newer than what's on disk
if err != nil || (err == nil && !ondiskTimestamp.After(srcTimestamp)) {
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = srcTimestamp.UTC().Format(time.RFC3339Nano)
srcInfo.UserDefined[xhttp.AmzObjectTagging] = objTags
}
}
} else {
if dstOpts.ReplicationRequest {
srcTimestamp := dstOpts.ReplicationSourceTaggingTimestamp
if !srcTimestamp.IsZero() {
// An empty value with a timestamp is an ordered deletion. Recheck
// the captured state even if metadata REPLACE rebuilt the map.
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = srcTimestamp.UTC().Format(time.RFC3339Nano)
srcInfo.UserDefined[xhttp.AmzObjectTagging] = objTags
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = UTCNow().Format(time.RFC3339Nano)
reconcileStoredObjectTags(srcInfo.UserDefined, srcInfo.UserTags, storedTagTimestamp)
} else {
srcInfo.UserDefined[xhttp.AmzObjectTagging] = srcInfo.UserTags
if storedTagTimestamp != "" {
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = storedTagTimestamp
} else {
delete(srcInfo.UserDefined, ReservedMetadataPrefixLower+TaggingTimestamp)
}
}
} else {
srcInfo.UserDefined[xhttp.AmzObjectTagging] = objTags
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
// SSE-C rotation snapshots reserved metadata before the tag decision. Its
// later merge must not put the old timestamp back over the accepted state.
delete(encMetadata, ReservedMetadataPrefixLower+TaggingTimestamp)
srcInfo.UserDefined = filterReplicationStatusMetadata(srcInfo.UserDefined)
srcInfo.UserDefined = objectlock.FilterObjectLockMetadata(srcInfo.UserDefined, true, true)
@@ -2313,6 +2320,9 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
if opts.ReplicationRequest && !opts.ReplicationSourceTaggingTimestamp.IsZero() {
metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = opts.ReplicationSourceTaggingTimestamp.UTC().Format(time.RFC3339Nano)
}
actualSize := size
var idxCb func() []byte
@@ -3760,11 +3770,11 @@ func (api objectAPIHandlers) PutObjectTaggingHandler(w http.ResponseWriter, r *h
}
dsc := mustReplicate(ctx, bucket, object, getMustReplicateOptions(objInfo.UserDefined, tagsStr, objInfo.ReplicationStatus, replication.MetadataReplicationType, opts))
stamp := UTCNow().Format(time.RFC3339Nano)
opts.UserDefined = map[string]string{ReservedMetadataPrefixLower + TaggingTimestamp: stamp}
if dsc.ReplicateAny() {
opts.UserDefined = make(map[string]string)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = UTCNow().Format(time.RFC3339Nano)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = stamp
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationStatus] = dsc.PendingStatus()
opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
// Put object tags
@@ -3863,9 +3873,10 @@ func (api objectAPIHandlers) DeleteObjectTaggingHandler(w http.ResponseWriter, r
}
dsc := mustReplicate(ctx, bucket, object, oi.getMustReplicateOptions(replication.MetadataReplicationType, opts))
stamp := UTCNow().Format(time.RFC3339Nano)
opts.UserDefined = map[string]string{ReservedMetadataPrefixLower + TaggingTimestamp: stamp}
if dsc.ReplicateAny() {
opts.UserDefined = make(map[string]string)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = UTCNow().Format(time.RFC3339Nano)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = stamp
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationStatus] = dsc.PendingStatus()
}
+4
View File
@@ -312,6 +312,10 @@ func (api objectAPIHandlers) NewMultipartUploadHandler(w http.ResponseWriter, r
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
// Completion orders the upload's persisted tag state under the object lock.
if opts.ReplicationRequest && !opts.ReplicationSourceTaggingTimestamp.IsZero() {
metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = opts.ReplicationSourceTaggingTimestamp.UTC().Format(time.RFC3339Nano)
}
if r.Header.Get(xhttp.IfMatch) != "" {
opts.HasIfMatch = true
+11 -4
View File
@@ -204,7 +204,9 @@ func (s *peerRESTServer) DeleteServiceAccountHandler(mss *grid.MSS) (np grid.NoP
return np, grid.NewRemoteErr(errors.New("service account name is missing"))
}
if err := globalIAMSys.DeleteServiceAccount(context.Background(), accessKey, false); err != nil {
ctx, cancel := context.WithTimeout(GlobalContext, defaultContextTimeout)
defer cancel()
if err := globalIAMSys.LoadServiceAccount(ctx, accessKey); err != nil {
return np, grid.NewRemoteErr(err)
}
@@ -230,7 +232,8 @@ func (s *peerRESTServer) LoadServiceAccountHandler(mss *grid.MSS) (np grid.NoPay
return np, nerr
}
// DeleteUserHandler - deletes a user on the server.
// DeleteUserHandler reloads the state committed by another node. A delayed
// notification must not delete an identity recreated since that commit.
func (s *peerRESTServer) DeleteUserHandler(mss *grid.MSS) (np grid.NoPayload, nerr *grid.RemoteErr) {
objAPI := newObjectLayerFn()
if objAPI == nil {
@@ -242,7 +245,9 @@ func (s *peerRESTServer) DeleteUserHandler(mss *grid.MSS) (np grid.NoPayload, ne
return np, grid.NewRemoteErr(errors.New("username is missing"))
}
if err := globalIAMSys.DeleteUser(context.Background(), accessKey, false); err != nil {
ctx, cancel := context.WithTimeout(GlobalContext, defaultContextTimeout)
defer cancel()
if err := globalIAMSys.LoadUserAfterDelete(ctx, accessKey); err != nil {
return np, grid.NewRemoteErr(err)
}
@@ -271,7 +276,9 @@ func (s *peerRESTServer) LoadUserHandler(mss *grid.MSS) (np grid.NoPayload, nerr
userType = stsUser
}
if err = globalIAMSys.LoadUser(context.Background(), objAPI, accessKey, userType); err != nil {
ctx, cancel := context.WithTimeout(GlobalContext, defaultContextTimeout)
defer cancel()
if err = globalIAMSys.LoadUser(ctx, objAPI, accessKey, userType); err != nil {
return np, grid.NewRemoteErr(err)
}
+276
View File
@@ -0,0 +1,276 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"archive/tar"
"bytes"
"compress/gzip"
"encoding/xml"
"maps"
"net/http"
"net/http/httptest"
"net/url"
"reflect"
"strings"
"testing"
"github.com/minio/minio/internal/auth"
xhttp "github.com/minio/minio/internal/http"
)
// Exercise authenticated handlers and actual disk metadata, including the
// response headers consumers see after replication has completed.
func TestAPIReplicaContentEncoding(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPIReplicaContentEncoding})
}
func testAPIReplicaContentEncoding(obj ObjectLayer, instance, bucket string, router http.Handler, owner auth.Credentials, t *testing.T) {
ordinary := newObjectAttributesAuthzUser(t, instance, bucket, `"s3:PutObject","s3:GetObject"`)
replicator := newObjectAttributesAuthzUser(t, instance, bucket, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
for _, mode := range []string{"ordinary", "untrusted-marker", "replica"} {
for _, tc := range []struct{ name, wire, want string }{
{"bare", "aws-chunked", ""}, {"mixed", "aws-chunked,gzip", "gzip"}, {"gzip", "gzip", "gzip"},
} {
for _, operation := range []string{"put", "copy-replace", "multipart"} {
t.Run(instance+"/"+mode+"/"+tc.name+"/"+operation, func(t *testing.T) {
object := mode + "/" + tc.name + "/" + operation
payload := replicaEncodingPayload(t, tc.want)
creds := ordinary
headers := map[string]string{xhttp.ContentEncoding: tc.wire, xhttp.ContentType: "application/octet-stream", "X-Amz-Meta-Source": "encoding-test"}
if mode != "ordinary" {
headers[xhttp.MinIOSourceReplicationRequest] = "true"
}
if mode == "replica" {
creds = replicator
headers[xhttp.AmzBucketReplicationStatus] = "REPLICA"
}
send := func(method, target string, data []byte, hdrs map[string]string) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, target, int64(len(data)), bytes.NewReader(data), creds.AccessKey, creds.SecretKey, hdrs)
if err != nil {
t.Fatal(err)
}
return replicaEncodingServe(t, router, req, http.StatusOK)
}
switch operation {
case "put":
if strings.Contains(tc.wire, "aws-chunked") {
req := replicaEncodingStream(t, getPutObjectURL("", bucket, object), payload, creds, headers)
replicaEncodingServe(t, router, req, http.StatusOK)
} else {
send(http.MethodPut, getPutObjectURL("", bucket, object), payload, headers)
}
case "copy-replace":
source := object + "-source"
if _, err := obj.PutObject(t.Context(), bucket, source, mustGetPutObjReader(t, bytes.NewReader(payload), int64(len(payload)), "", ""), ObjectOptions{}); err != nil {
t.Fatal(err)
}
headers[xhttp.AmzCopySource] = url.QueryEscape("/" + bucket + "/" + source)
headers[xhttp.AmzMetadataDirective] = replaceDirective
send(http.MethodPut, getCopyObjectURL("", bucket, object), nil, headers)
case "multipart":
rec := send(http.MethodPost, getNewMultipartURL("", bucket, object), nil, headers)
var init InitiateMultipartUploadResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &init); err != nil {
t.Fatal(err)
}
// Part/completion metadata must not replace the encoding saved at initiation.
partHeaders := map[string]string{xhttp.ContentEncoding: "br"}
if mode == "replica" {
partHeaders[xhttp.MinIOSourceReplicationRequest] = "true"
partHeaders[xhttp.AmzBucketReplicationStatus] = "REPLICA"
}
part := send(http.MethodPut, getPutObjectPartURL("", bucket, object, init.UploadID, "1"), payload, partHeaders)
partETags := part.Header()[xhttp.ETag]
if len(partETags) != 1 {
t.Fatalf("missing part ETag: %#v", part.Header())
}
complete, err := xml.Marshal(CompleteMultipartUpload{Parts: []CompletePart{{PartNumber: 1, ETag: canonicalizeETag(partETags[0])}}})
if err != nil {
t.Fatal(err)
}
send(http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, init.UploadID), complete, partHeaders)
}
assertReplicaEncodingObject(t, obj, router, owner, bucket, object, tc.want, payload)
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if got := info.UserDefined[xhttp.AmzBucketReplicationStatus]; (got == "REPLICA") != (mode == "replica") {
t.Errorf("replica status %q for mode %s", got, mode)
}
if info.ContentType != "application/octet-stream" {
t.Errorf("content-type=%q", info.ContentType)
}
if value, ok := caseInsensitiveMap(info.UserDefined).Lookup("x-amz-meta-source"); !ok || value != "encoding-test" {
t.Errorf("user metadata lost: %#v", info.UserDefined)
}
})
}
}
}
t.Run(instance+"/unauthorized-replica", func(t *testing.T) {
object := "denied-replica"
req := replicaEncodingStream(t, getPutObjectURL("", bucket, object), []byte("denied"), ordinary, map[string]string{xhttp.ContentEncoding: "aws-chunked", xhttp.MinIOSourceReplicationRequest: "true", xhttp.AmzBucketReplicationStatus: "REPLICA"})
rec := replicaEncodingServe(t, router, req, http.StatusForbidden)
var response APIErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatal(err)
}
if response.Code != "AccessDenied" {
t.Fatalf("expected permission denial, got %s", response.Code)
}
if _, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); err == nil {
t.Error("denied replica created an object")
}
})
}
func replicaEncodingPayload(t *testing.T, encoding string) []byte {
t.Helper()
data := bytes.Repeat([]byte("replica encoding payload\n"), 128)
if encoding != "gzip" {
return data
}
var b bytes.Buffer
w := gzip.NewWriter(&b)
if _, err := w.Write(data); err != nil {
t.Fatal(err)
}
if err := w.Close(); err != nil {
t.Fatal(err)
}
return b.Bytes()
}
func replicaEncodingStream(t *testing.T, target string, data []byte, creds auth.Credentials, headers map[string]string) *http.Request {
t.Helper()
const chunkSize = 64
body := bytes.NewReader(data)
req, err := newTestStreamingRequest(http.MethodPut, target, int64(len(data)), chunkSize, body)
if err != nil {
t.Fatal(err)
}
for k, v := range headers {
req.Header.Set(k, v)
}
now := UTCNow()
signature, err := signStreamingRequest(req, creds.AccessKey, creds.SecretKey, now)
if err != nil {
t.Fatal(err)
}
req, err = assembleStreamingChunks(req, body, chunkSize, creds.SecretKey, signature, now)
if err != nil {
t.Fatal(err)
}
return req
}
func replicaEncodingServe(t *testing.T, router http.Handler, req *http.Request, want int) *httptest.ResponseRecorder {
t.Helper()
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != want {
t.Fatalf("%s %s: status=%d want=%d body=%s", req.Method, req.URL, rec.Code, want, rec.Body.String())
}
return rec
}
func assertReplicaEncodingObject(t *testing.T, obj ObjectLayer, router http.Handler, creds auth.Credentials, bucket, object, encoding string, data []byte) {
t.Helper()
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if info.ContentEncoding != encoding {
t.Errorf("persisted content-encoding=%q want=%q", info.ContentEncoding, encoding)
}
if encoding == "" {
if _, present := info.UserDefined["content-encoding"]; present {
t.Error("transport-only content-encoding key persisted")
}
}
for _, method := range []string{http.MethodGet, http.MethodHead} {
req, err := newTestSignedRequestV4(method, getPutObjectURL("", bucket, object), 0, nil, creds.AccessKey, creds.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := replicaEncodingServe(t, router, req, http.StatusOK)
if got := rec.Header().Get(xhttp.ContentEncoding); got != encoding {
t.Errorf("%s content-encoding=%q want=%q", method, got, encoding)
}
if encoding == "" {
if _, present := rec.Header()[xhttp.ContentEncoding]; present {
t.Errorf("%s sent an empty/transport encoding header", method)
}
}
if method == http.MethodGet && !bytes.Equal(rec.Body.Bytes(), data) {
t.Errorf("GET body differs: got %d bytes want %d", rec.Body.Len(), len(data))
}
}
}
func TestAPISnowballReplicaContentEncoding(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, creds auth.Credentials, t *testing.T) {
for _, tc := range []struct {
name string
pax map[string]string
want string
}{
{name: "no-pax"},
{name: "pax-without-encoding", pax: map[string]string{"minio.metadata.Content-Type": "application/octet-stream"}},
{name: "pax-bare", pax: map[string]string{"minio.metadata.Content-Encoding": "aws-chunked"}},
{name: "pax-mixed", pax: map[string]string{"minio.metadata.Content-Encoding": "aws-chunked,gzip"}, want: "gzip"},
} {
t.Run(instance+"/"+tc.name, func(t *testing.T) {
object := "snowball/" + tc.name
data := replicaEncodingPayload(t, tc.want)
var archive bytes.Buffer
tw := tar.NewWriter(&archive)
if err := tw.WriteHeader(&tar.Header{Name: object, Mode: 0o600, Size: int64(len(data)), PAXRecords: tc.pax}); err != nil {
t.Fatal(err)
}
if _, err := tw.Write(data); err != nil {
t.Fatal(err)
}
if err := tw.Close(); err != nil {
t.Fatal(err)
}
var ordinaryMetadata map[string]string
// An unauthorized entry in a REPLICA request is rejected. Compare the
// same archive across ordinary and authorized replica requests instead.
for _, replica := range []bool{false, true} {
headers := map[string]string{
xhttp.ContentEncoding: "aws-chunked", xhttp.AmzSnowballExtract: "true",
xhttp.ContentType: "application/x-tar", xhttp.CacheControl: "max-age=123",
"X-Amz-Meta-Archive": "outer-request",
}
if replica {
headers[xhttp.MinIOSourceReplicationRequest] = "true"
headers[xhttp.AmzBucketReplicationStatus] = "REPLICA"
}
req := replicaEncodingStream(t, getPutObjectURL("", bucket, "archive.tar"), archive.Bytes(), creds, headers)
replicaEncodingServe(t, router, req, http.StatusOK)
assertReplicaEncodingObject(t, obj, router, creds, bucket, object, tc.want, data)
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
metadata := maps.Clone(info.UserDefined)
for _, key := range []string{xhttp.AmzBucketReplicationStatus, ReservedMetadataPrefixLower + ReplicaStatus, ReservedMetadataPrefixLower + ReplicaTimestamp, "etag"} {
delete(metadata, key)
}
if !replica {
ordinaryMetadata = metadata
} else if !reflect.DeepEqual(metadata, ordinaryMetadata) {
t.Errorf("replica inherited ordinary archive metadata: got %#v want %#v", metadata, ordinaryMetadata)
}
}
})
}
}})
}
+5 -19
View File
@@ -106,6 +106,7 @@ func TestReplicateDeleteMarkerTargetSemantics(t *testing.T) {
func testReplicateDeleteMarkerPurge(obj ObjectLayer, instanceType, bucket string, router http.Handler, creds auth.Credentials, t *testing.T, legacy bool) {
ctx := t.Context()
defer replicationTestCapacity(obj)()
const arn = "arn:minio:replication::af470089-d354-4473-934c-9e1f52f6da89:bucket"
const name = "marker"
version := mustGetUUID()
@@ -208,27 +209,12 @@ func testReplicateDeleteMarkerPurge(obj ObjectLayer, instanceType, bucket string
t.Errorf("purge scheduled as marker creation: version=%q marker=%q", deletion.VersionID, deletion.DeleteMarkerVersionID)
}
if legacy {
// Reproduce the old producer's state and let the existing scanner/heal
// path recover it. Upgrades must also finish purges already left pending.
// Old task shapes must complete directly; scanner/MRF recovery is
// covered separately by TestReplicationMRFMarkerRecovery.
deletion.VersionID, deletion.DeleteMarkerVersionID = "", version
replicateDelete(ctx, deletion, obj)
oi, _ := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version, Versioned: true})
if oi.VersionPurgeStatus != replication.VersionPurgePending {
t.Fatalf("legacy source purge = %s, want PENDING", oi.VersionPurgeStatus)
}
targets, err := globalBucketTargetSys.ListBucketTargets(ctx, bucket)
if err != nil {
t.Fatal(err)
}
queueReplicationHeal(ctx, bucket, oi, replicationConfig{Config: &cfg, remotes: targets}, 0)
select {
case op := <-worker:
deletion = op.(DeletedObjectReplicationInfo)
case <-time.After(time.Second):
t.Fatal("legacy pending purge was not scheduled for healing")
}
}
result := replicateDelete(context.Background(), deletion, obj)
result := replicateDelete(context.Background(), deletion, markerPurgeUpdateLayer{ObjectLayer: obj, t: t})
if result.VersionPurgeStatus() != replication.VersionPurgeComplete {
t.Errorf("remote purge result = %s, want COMPLETE", result.VersionPurgeStatus())
}
+628
View File
@@ -0,0 +1,628 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/logger"
loghttp "github.com/minio/minio/internal/logger/target/http"
"github.com/minio/minio/internal/once"
xnet "github.com/pgsty/silo-pkg/v3/net"
)
type markerRecoveryCase struct {
name string
legacy, creation, lostReply, partial, exhaust, lockFirst, invalidMRF, replicaSource, offline, unrecordedPurge bool
targets int
}
func TestReplicationMRFMarkerRecovery(t *testing.T) {
for _, tc := range []markerRecoveryCase{
{name: "canonical", targets: 1},
{name: "lock-failure", lockFirst: true, targets: 1},
{name: "invalid-MRF-metadata", invalidMRF: true, targets: 1},
{name: "legacy", legacy: true, targets: 1},
{name: "unrecorded-purge", unrecordedPurge: true, targets: 2},
{name: "two-targets", legacy: true, targets: 2},
{name: "two-targets-offline", offline: true, targets: 2},
{name: "replica-source", replicaSource: true, targets: 1},
{name: "lost-reply-canonical", lostReply: true, targets: 1},
{name: "lost-reply-legacy", legacy: true, lostReply: true, targets: 1},
{name: "creation", creation: true, targets: 1},
{name: "partial-creation-block", partial: true, targets: 2},
{name: "retry-budget-and-scanner", exhaust: true, targets: 1},
} {
t.Run(tc.name, func(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, endpoints: []string{"DeleteObject"}, objAPITest: func(obj ObjectLayer, backend, bucket string, router http.Handler, creds auth.Credentials, t *testing.T) {
testReplicationMRFMarkerRecovery(t, obj, backend, bucket, router, creds, tc)
}})
})
}
}
type markerRecoveryTarget struct {
arn, bucket string
client *minio.Client
reject, loseReply atomic.Bool
deletes atomic.Int32
}
// Only adapt host capacity accounting, using the existing real-disk adapter.
// Reads, object metadata, MRF files and writes all still use the fixture disks.
func replicationTestCapacity(obj ObjectLayer) func() {
var restore []func()
for _, pool := range obj.(*erasureServerPools).serverPools {
for _, set := range pool.sets {
original := set.getDisks
disks := append([]StorageAPI(nil), original()...)
for i, disk := range disks {
if disk != nil {
disks[i] = tagTestCapacityDisk{StorageAPI: disk}
}
}
set.getDisks = func() []StorageAPI { return disks }
restore = append(restore, func() { set.getDisks = original })
}
}
return func() {
for _, fn := range restore {
fn()
}
}
}
func testReplicationMRFMarkerRecovery(t *testing.T, obj ObjectLayer, backend, bucket string, router http.Handler, creds auth.Credentials, tc markerRecoveryCase) {
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
defer replicationTestCapacity(obj)()
stats := NewReplicationStats(ctx, nil)
oldStats := globalReplicationStats.Swap(stats)
defer globalReplicationStats.Store(oldStats)
oldPool := globalReplicationPool
defer func() { globalReplicationPool = oldPool }()
const name = "marker"
version := mustGetUUID()
creationTime := UTCNow().Add(-time.Hour).Truncate(time.Second)
if _, err := globalBucketMetadataSys.Update(ctx, bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
if _, err := obj.PutObject(ctx, bucket, name, mustGetPutObjReader(t, bytes.NewReader([]byte("data")), 4, "", ""), ObjectOptions{Versioned: true}); err != nil {
t.Fatal(err)
}
var targets []*markerRecoveryTarget
cfg := replication.Config{}
creationStates := make(map[string]replication.StatusType)
var creationInternal string
for i := 0; i < tc.targets; i++ {
target := &markerRecoveryTarget{arn: "arn:minio:replication::" + mustGetUUID() + ":bucket", bucket: getRandomBucketName()}
if err := obj.MakeBucket(ctx, target.bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
if _, err := obj.PutObject(ctx, target.bucket, name, mustGetPutObjReader(t, bytes.NewReader([]byte("data")), 4, "", ""), ObjectOptions{Versioned: true}); err != nil {
t.Fatal(err)
}
if !tc.creation {
opts := ObjectOptions{VersionID: version, Versioned: true, DeleteMarker: true, ReplicationRequest: true, MTime: creationTime}
opts.SetReplicaStatus(replication.Replica)
if _, err := obj.DeleteObject(ctx, target.bucket, name, opts); err != nil {
t.Fatal(err)
}
}
remote := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
opts := ObjectOptions{VersionID: r.URL.Query().Get("versionId"), Versioned: true}
switch r.Method {
case http.MethodHead:
oi, err := obj.GetObjectInfo(r.Context(), target.bucket, name, opts)
if oi.DeleteMarker {
w.Header().Set(xhttp.AmzDeleteMarker, "true")
w.Header().Set(xhttp.AmzVersionID, oi.VersionID)
}
if err != nil {
writeErrorResponseHeadersOnly(w, toAPIError(r.Context(), err))
return
}
w.WriteHeader(http.StatusOK)
case http.MethodDelete:
target.deletes.Add(1)
if got := r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true"; got != tc.creation {
t.Errorf("purge/creation wire flag=%v creation=%v", got, tc.creation)
}
if opts.VersionID == "" {
t.Error("missing remote versionId")
}
if target.reject.Load() {
w.WriteHeader(http.StatusForbidden)
fmt.Fprint(w, `<Error><Code>AccessDenied</Code></Error>`)
return
}
opts.DeleteMarker = r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true"
opts.ReplicationRequest = true
opts.SetReplicaStatus(replication.Replica)
_, err := obj.DeleteObject(r.Context(), target.bucket, name, opts)
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
writeErrorResponse(r.Context(), w, toAPIError(r.Context(), err), r.URL)
return
}
if target.loseReply.Swap(false) {
conn, _, err := w.(http.Hijacker).Hijack()
if err != nil {
t.Error(err)
return
}
conn.Close()
return
}
w.WriteHeader(http.StatusNoContent)
default:
t.Errorf("unexpected remote method %s", r.Method)
}
}))
defer remote.Close()
client, err := minio.New(strings.TrimPrefix(remote.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
target.client = client
globalBucketTargetSys.Lock()
globalBucketTargetSys.arnRemotesMap[target.arn] = arnTarget{Client: &TargetClient{Client: client, ARN: target.arn, Bucket: target.bucket}, lastRefresh: UTCNow()}
globalBucketTargetSys.targetsMap[bucket] = append(globalBucketTargetSys.targetsMap[bucket], madmin.BucketTarget{Arn: target.arn, TargetBucket: target.bucket})
globalBucketTargetSys.Unlock()
globalBucketTargetSys.hMutex.Lock()
globalBucketTargetSys.hc[client.EndpointURL().Host] = epHealth{Online: true}
globalBucketTargetSys.hMutex.Unlock()
rule := configs[0].Rules[0]
rule.Priority = i + 1
rule.Destination = replication.Destination{ARN: target.arn, Bucket: target.bucket}
cfg.Rules = append(cfg.Rules, rule)
creationStates[target.arn] = replication.Completed
creationInternal += target.arn + "=COMPLETED;"
targets = append(targets, target)
}
if tc.targets == 1 {
cfg.RoleArn = targets[0].arn
}
if !tc.creation {
opts := ObjectOptions{VersionID: version, Versioned: true, DeleteMarker: true, MTime: creationTime, DeleteReplication: ReplicationState{Targets: creationStates, ReplicationStatusInternal: creationInternal, ReplicationTimeStamp: creationTime}}
if tc.replicaSource {
opts.DeleteReplication = ReplicationState{}
opts.SetReplicaStatus(replication.Replica)
}
if _, err := obj.DeleteObject(ctx, bucket, name, opts); err != nil {
t.Fatal(err)
}
}
meta, err := globalBucketMetadataSys.Get(bucket)
if err != nil {
t.Fatal(err)
}
meta.replicationConfig = &cfg
globalBucketMetadataSys.Set(bucket, meta)
newPool := func() *ReplicationPool {
p := &ReplicationPool{ctx: ctx, objLayer: obj, workers: []chan ReplicationWorkerOperation{make(chan ReplicationWorkerOperation, 8)}, stats: stats, mrfSaveCh: make(chan MRFReplicateEntry, 8)}
globalReplicationPool = once.NewSingleton[ReplicationPool]()
globalReplicationPool.Set(p)
return p
}
p := newPool()
receive := func() DeletedObjectReplicationInfo {
select {
case op := <-p.workers[0]:
d, ok := op.(DeletedObjectReplicationInfo)
if !ok {
t.Fatalf("wrong queued operation %T", op)
}
return d
case <-time.After(3 * time.Second):
t.Fatal("no marker task from actual MRF/handler/scanner queue")
return DeletedObjectReplicationInfo{}
}
}
uri := "/" + bucket + "/" + name
if !tc.creation {
uri += "?versionId=" + version
}
req, err := newTestSignedRequestV4(http.MethodDelete, uri, 0, nil, creds.AccessKey, creds.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
if w.Code != http.StatusNoContent {
t.Fatalf("source DELETE status %d: %s", w.Code, w.Body)
}
deletion := receive()
if tc.creation {
version = deletion.DeleteMarkerVersionID
} else if deletion.VersionID != version || deletion.DeleteMarkerVersionID != "" {
t.Fatalf("current producer emitted noncanonical purge: %+v", deletion)
}
if tc.unrecordedPurge {
// Exercise a task carrying purge state while the disk marker still has
// only creation metadata. Recreate only this isolated fixture marker.
if _, err := obj.DeleteObject(ctx, bucket, name, ObjectOptions{VersionID: version, Versioned: true}); err != nil {
t.Fatal(err)
}
opts := ObjectOptions{VersionID: version, Versioned: true, DeleteMarker: true, MTime: creationTime, DeleteReplication: ReplicationState{Targets: creationStates, ReplicationStatusInternal: creationInternal, ReplicationTimeStamp: creationTime}}
if _, err := obj.DeleteObject(ctx, bucket, name, opts); err != nil {
t.Fatal(err)
}
}
if tc.legacy {
deletion.VersionID, deletion.DeleteMarkerVersionID = "", version
}
if tc.partial {
deletion.TargetArn = targets[len(targets)-1].arn
}
if tc.invalidMRF {
testReplicationMRFInvalidLookups(ctx, t, obj, p, newPool, deletion)
return
}
var auditStatuses <-chan string
if tc.name == "canonical" {
var stopAudit func()
auditStatuses, stopAudit = replicationTestAudit(ctx, t, bucket)
defer stopAudit()
}
assertAudit := func(want string) {
if auditStatuses == nil {
return
}
select {
case got := <-auditStatuses:
if got != want {
t.Fatalf("replication audit status=%q want %q", got, want)
}
case <-time.After(3 * time.Second):
t.Fatal("missing replication audit event")
}
}
deletion.OpType = replication.HealReplicationType // Exercise operation-status statistics too.
failing := targets[len(targets)-1]
failing.reject.Store(!tc.lostReply)
failing.loseReply.Store(tc.lostReply)
if tc.offline {
globalBucketTargetSys.hMutex.Lock()
globalBucketTargetSys.hc[failing.client.EndpointURL().Host] = epHealth{Online: false}
globalBucketTargetSys.hMutex.Unlock()
}
before, _ := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version})
beforeCreation := before.ReplicationStatusInternal
beforeStamp := before.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp]
beforeReplica := before.UserDefined[ReservedMetadataPrefixLower+ReplicaStatus]
beforeReplicaStamp := before.UserDefined[ReservedMetadataPrefixLower+ReplicaTimestamp]
if !tc.creation && !tc.replicaSource && (len(replicationStatusesMap(beforeCreation)) != tc.targets || beforeStamp == "") {
t.Fatalf("missing seeded creation block: %+v", before)
}
if tc.replicaSource && (beforeReplica != "REPLICA" || beforeReplicaStamp == "") {
t.Fatalf("missing replica block: %+v", before)
}
updateObj := obj
if !tc.creation {
updateObj = markerPurgeUpdateLayer{ObjectLayer: obj, t: t}
}
rounds := 2
if tc.lostReply {
rounds = 1
}
if tc.exhaust {
rounds = mrfRetryLimit + 1
}
for round := 1; round <= rounds; round++ {
callObj := updateObj
if tc.lockFirst && round == 1 {
callObj = markerLockFailureLayer{ObjectLayer: updateObj}
}
result := replicateDelete(ctx, deletion, callObj)
switch {
case tc.lockFirst && round == 1:
if len(result.Targets) != 0 {
t.Fatal("lock failure attempted a target")
}
case tc.creation:
if result.ReplicationStatus() != replication.Failed {
t.Fatalf("creation result: %+v", result)
}
case result.VersionPurgeStatus() != replication.VersionPurgeFailed:
t.Fatalf("purge failure result: %+v", result)
}
assertAudit("FAILED")
if round == 1 && !tc.lockFirst && !tc.creation {
stats.RLock()
failed := stats.Cache[bucket].Stats[failing.arn].FailStats.SinceUptime
stats.RUnlock()
if failed.Count != 1 || failed.Bytes != 0 {
t.Fatalf("purge failure stats=%+v, want count 1 and zero bytes", failed)
}
}
oi, err := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version})
if !isErrMethodNotAllowed(err) || !oi.DeleteMarker {
t.Fatalf("source marker metadata/405 missing: %+v %v", oi, err)
}
if !tc.creation && (oi.ReplicationStatusInternal != beforeCreation || oi.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] != beforeStamp) {
t.Fatalf("purge rewrote creation block: before=%q/%q after=%q/%q", beforeCreation, beforeStamp, oi.ReplicationStatusInternal, oi.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp])
}
if !tc.creation && (oi.UserDefined[ReservedMetadataPrefixLower+ReplicaStatus] != beforeReplica || oi.UserDefined[ReservedMetadataPrefixLower+ReplicaTimestamp] != beforeReplicaStamp) {
t.Fatal("purge rewrote replica block")
}
if tc.targets == 2 && !tc.partial {
purges := versionPurgeStatusesMap(oi.VersionPurgeStatusInternal)
if purges[targets[0].arn] != replication.VersionPurgeComplete || purges[failing.arn] != replication.VersionPurgeFailed {
t.Fatalf("incorrect persisted target states: %v", purges)
}
if targets[0].deletes.Load() != 1 {
t.Fatalf("successful target resent %d times", targets[0].deletes.Load())
}
}
if tc.partial {
select {
case <-p.mrfSaveCh:
default:
t.Fatal("partial failure missing MRF")
}
dsc := deletion.ReplicationState.ReplicateDecisionStr
deletion.ReplicationState = oi.ReplicationState()
deletion.ReplicationState.ReplicateDecisionStr = dsc
continue
}
if tc.exhaust && round > mrfRetryLimit {
if len(p.mrfSaveCh) != 0 || atomic.LoadUint64(&stats.mrfStats.TotalDroppedCount) != 1 {
t.Fatalf("retry budget not applied: queued=%d drops=%d", len(p.mrfSaveCh), stats.mrfStats.TotalDroppedCount)
}
break
}
var entry MRFReplicateEntry
select {
case entry = <-p.mrfSaveCh:
default:
t.Fatal("failure did not enter MRF")
}
if entry.RetryCount != round || entry.versionID != version {
t.Fatalf("MRF entry lost identity/budget: %+v, round=%d", entry, round)
}
p.saveMRFEntries(ctx, map[string]MRFReplicateEntry{entry.versionID: entry})
record, err := p.loadMRF()
if err != nil || len(record.Entries) != 1 {
t.Fatalf("disk MRF missing: %+v %v", record, err)
}
if got := record.Entries[version]; got.RetryCount != round || got.Object != name || got.Bucket != bucket {
t.Fatalf("disk MRF mismatch: %+v", got)
}
p.saveMRFEntries(ctx, record.Entries) // loadMRF consumes the file; each replay uses a fresh disk record.
p = newPool() // no in-memory entries carried to the replacement pool.
if err := p.queueMRFHeal(); err != nil {
t.Fatal(err)
}
deletion = receive()
if deletion.RetryCount != round {
t.Fatalf("MRF retry count=%d want %d", deletion.RetryCount, round)
}
if !tc.creation && (deletion.VersionID != version || deletion.DeleteMarkerVersionID != "") {
t.Fatalf("MRF emitted wrong purge: %+v", deletion)
}
}
if tc.partial {
return
}
failing.reject.Store(false)
globalBucketTargetSys.hMutex.Lock()
globalBucketTargetSys.hc[failing.client.EndpointURL().Host] = epHealth{Online: true}
globalBucketTargetSys.hMutex.Unlock()
if tc.exhaust {
oi, _ := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version})
QueueReplicationHeal(ctx, bucket, oi, 0) // Separately prove the existing scanner fallback after MRF exhaustion.
deletion = receive()
if deletion.RetryCount != 0 {
t.Fatal("scanner did not start a fresh retry budget")
}
}
result := replicateDelete(ctx, deletion, updateObj)
assertAudit("COMPLETED")
if tc.creation {
if result.ReplicationStatus() != replication.Completed {
t.Fatalf("creation recovery failed: %+v", result)
}
} else if result.VersionPurgeStatus() != replication.VersionPurgeComplete {
t.Fatalf("purge recovery failed: %+v", result)
}
for _, b := range append([]string{bucket}, func() []string {
var b []string
for _, target := range targets {
b = append(b, target.bucket)
}
return b
}()...) {
oi, err := obj.GetObjectInfo(ctx, b, name, ObjectOptions{VersionID: version})
if tc.creation {
if !oi.DeleteMarker || !isErrMethodNotAllowed(err) {
t.Fatalf("creation missing in %s: %+v %v", b, oi, err)
}
} else if !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
t.Fatalf("purge left marker in %s: %+v %v", b, oi, err)
}
}
if !tc.creation {
// Re-deliver an ambiguous purge directly. An absent marker must stay absent.
duplicate := deletion
duplicate.VersionID, duplicate.DeleteMarkerVersionID = "", version
duplicate.ReplicationState.PurgeTargets = map[string]VersionPurgeStatusType{failing.arn: replication.VersionPurgePending}
duplicate.ReplicationState.VersionPurgeStatusInternal = ""
got := replicateDeleteToTarget(ctx, duplicate, &TargetClient{Client: failing.client, ARN: failing.arn, Bucket: failing.bucket})
if got.VersionPurgeStatus != replication.VersionPurgeComplete {
t.Fatalf("duplicate purge: %+v", got)
}
if oi, err := obj.GetObjectInfo(ctx, failing.bucket, name, ObjectOptions{VersionID: version}); !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
t.Fatalf("duplicate recreated marker: %+v %v", oi, err)
}
}
if len(p.mrfSaveCh) != 0 || len(p.workers[0]) != 0 {
t.Fatalf("completion left queued work: mrf=%d worker=%d", len(p.mrfSaveCh), len(p.workers[0]))
}
if !tc.creation && !tc.exhaust && !tc.lostReply {
stats.RLock()
got := stats.Cache[bucket].Stats[failing.arn].ReplicatedCount
size := stats.Cache[bucket].Stats[failing.arn].ReplicatedSize
stats.RUnlock()
if got != 1 || size != 0 {
t.Fatalf("purge completion statistic=%d want 1", got)
}
}
t.Logf("%s: %s recovered; %d target(s), persisted MRF, source/target metadata checked", backend, tc.name, tc.targets)
}
func TestReplicationDeleteQueueFullRetryBudget(t *testing.T) {
stats := NewReplicationStats(t.Context(), nil)
p := &ReplicationPool{ctx: t.Context(), objLayer: &replicationMRFTestObjectLayer{}, workers: []chan ReplicationWorkerOperation{make(chan ReplicationWorkerOperation)}, stats: stats, mrfSaveCh: make(chan MRFReplicateEntry, 1), priority: "slow"}
d := DeletedObjectReplicationInfo{Bucket: "bucket", DeletedObject: DeletedObject{ObjectName: "marker", VersionID: mustGetUUID()}, RetryCount: 2}
p.queueReplicaDeleteTask(d)
if got := <-p.mrfSaveCh; got.RetryCount != 3 {
t.Fatalf("queue-full retry count=%d", got.RetryCount)
}
d.RetryCount = mrfRetryLimit
p.queueReplicaDeleteTask(d)
if len(p.mrfSaveCh) != 0 || atomic.LoadUint64(&stats.mrfStats.TotalDroppedCount) != 1 {
t.Fatal("queue-full retry exceeded budget without a visible drop")
}
}
type markerLockFailureLayer struct{ ObjectLayer }
func (markerLockFailureLayer) NewNSLock(string, ...string) RWLocker { return markerFailedLock{} }
type markerFailedLock struct{ RWLocker }
func (markerFailedLock) GetLock(context.Context, *dynamicTimeout) (LockContext, error) {
return LockContext{}, context.DeadlineExceeded
}
type markerLookupLayer struct {
ObjectLayer
mutate func(ObjectInfo, error) (ObjectInfo, error)
lookedUp chan struct{}
}
func (l markerLookupLayer) GetObjectInfo(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
oi, err := l.ObjectLayer.GetObjectInfo(ctx, bucket, object, opts)
defer close(l.lookedUp)
return l.mutate(oi, err)
}
func testReplicationMRFInvalidLookups(ctx context.Context, t *testing.T, obj ObjectLayer, p *ReplicationPool, newPool func() *ReplicationPool, d DeletedObjectReplicationInfo) {
for _, tc := range []struct {
name string
mutate func(ObjectInfo, error) (ObjectInfo, error)
}{
{"empty-info", func(_ ObjectInfo, e error) (ObjectInfo, error) { return ObjectInfo{}, e }},
{"not-a-marker", func(o ObjectInfo, e error) (ObjectInfo, error) { o.DeleteMarker = false; return o, e }},
{"wrong-version", func(o ObjectInfo, e error) (ObjectInfo, error) { o.VersionID = mustGetUUID(); return o, e }},
{"wrong-bucket", func(o ObjectInfo, e error) (ObjectInfo, error) { o.Bucket = "different-bucket"; return o, e }},
{"wrong-object", func(o ObjectInfo, e error) (ObjectInfo, error) { o.Name = "different-object"; return o, e }},
{"zero-modtime", func(o ObjectInfo, e error) (ObjectInfo, error) { o.ModTime = time.Time{}; return o, e }},
{"missing", func(o ObjectInfo, _ error) (ObjectInfo, error) {
return o, ObjectNotFound{Bucket: o.Bucket, Object: o.Name}
}},
{"read-error", func(o ObjectInfo, _ error) (ObjectInfo, error) { return o, InsufficientReadQuorum{} }},
} {
t.Run(tc.name, func(t *testing.T) {
entry := d.ToMRFEntry()
p.saveMRFEntries(ctx, map[string]MRFReplicateEntry{entry.versionID: entry})
fresh := newPool()
looked := make(chan struct{})
fresh.objLayer = markerLookupLayer{ObjectLayer: obj, mutate: tc.mutate, lookedUp: looked}
if err := fresh.queueMRFHeal(); err != nil {
t.Fatal(err)
}
select {
case <-looked:
case <-time.After(3 * time.Second):
t.Fatal("disk MRF entry was not read")
}
select {
case op := <-fresh.workers[0]:
t.Fatalf("invalid lookup scheduled %T", op)
case <-time.After(100 * time.Millisecond):
}
})
}
}
// Capture the actual internal audit event through the supported webhook sink.
func replicationTestAudit(ctx context.Context, t *testing.T, bucket string) (<-chan string, func()) {
t.Helper()
if len(logger.AuditTargets()) != 0 {
t.Fatal("unexpected pre-existing test audit targets")
}
statuses := make(chan string, 16)
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var entry struct {
API struct{ Name, Bucket, Status string }
}
if err := json.NewDecoder(r.Body).Decode(&entry); err != nil {
t.Error(err)
} else if entry.API.Name == ReplicateDeleteAPI && entry.API.Bucket == bucket {
statuses <- entry.API.Status
}
w.WriteHeader(http.StatusOK)
}))
endpoint, err := xnet.ParseHTTPURL(server.URL)
if err != nil {
t.Fatal(err)
}
if errs := logger.UpdateAuditWebhooks(ctx, map[string]loghttp.Config{"r6": {Enabled: true, Name: "r6", Endpoint: endpoint, BatchSize: 1, QueueSize: 128, MaxRetry: 1, RetryIntvl: time.Millisecond, HTTPTimeout: time.Second}}); len(errs) > 0 {
t.Fatal(errs)
}
targets := logger.AuditTargets()
return statuses, func() {
logger.UpdateAuditWebhooks(ctx, nil)
for _, target := range targets {
target.Cancel()
}
server.Close()
}
}
type markerPurgeUpdateLayer struct {
ObjectLayer
t *testing.T
}
func (l markerPurgeUpdateLayer) DeleteObject(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
l.t.Helper()
rs := opts.DeleteReplication
if rs.ReplicationStatusInternal != "" || rs.Targets != nil || rs.ReplicaStatus != "" || !rs.CompositeReplicationStatus().Empty() {
l.t.Fatalf("purge has a nonempty creation update: %+v", rs)
}
if rs.CompositeVersionPurgeStatus().Empty() {
l.t.Fatal("purge update lost its purge status")
}
return l.ObjectLayer.DeleteObject(ctx, bucket, object, opts)
}
+215
View File
@@ -0,0 +1,215 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
)
// Purge uses the version-delete wire form regardless of the in-memory shape.
// Creation status, purge status and successful resync markers are independent.
func TestReplicateDeleteOperationExits(t *testing.T) {
for _, shape := range []string{"marker-creation", "legacy-marker-purge", "marker-purge", "object-purge"} {
purge := shape != "marker-creation"
for _, tc := range []struct {
name string
creation replication.StatusType
priorPurge VersionPurgeStatusType
headCode, deleteCode int
headError string
offline, resync bool
}{
{name: "pending-success", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 204},
{name: "completed-creation", creation: replication.Completed, priorPurge: replication.VersionPurgePending, headCode: 405, deleteCode: 204},
{name: "existing-marker", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 405, deleteCode: 204},
// Use the quorum S3 code without 503, which is classified as backend-down first.
{name: "head-read-quorum", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 400, headError: "SlowDownRead", deleteCode: 204},
{name: "replica-creation-status", creation: replication.Replica, priorPurge: replication.VersionPurgeFailed, headCode: 405, deleteCode: 204},
{name: "head-unavailable", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 503, deleteCode: 204},
{name: "head-forbidden", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 403, deleteCode: 204},
{name: "delete-forbidden", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 403},
{name: "delete-method-rejected", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 405},
{name: "delete-unavailable", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 503},
{name: "offline", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 204, offline: true},
{name: "retry-failed", creation: replication.Completed, priorPurge: replication.VersionPurgeFailed, headCode: 404, deleteCode: 204},
{name: "purge-complete", creation: replication.Pending, priorPurge: replication.VersionPurgeComplete, headCode: 404, deleteCode: 403},
{name: "resync-success", creation: replication.Completed, priorPurge: replication.VersionPurgePending, headCode: 405, deleteCode: 204, resync: true},
{name: "resync-already-purged", creation: replication.Pending, priorPurge: replication.VersionPurgeComplete, headCode: 404, deleteCode: 403, resync: true},
{name: "resync-failure", creation: replication.Completed, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 403, resync: true},
} {
t.Run(shape+"/"+tc.name, func(t *testing.T) {
version := mustGetUUID()
var heads, deletes atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Query().Get("versionId") != version {
t.Errorf("wrong request version: %s", r.URL)
}
switch r.Method {
case http.MethodHead:
heads.Add(1)
w.Header().Set(xhttp.AmzVersionID, version)
w.Header().Set(xhttp.LastModified, time.Now().UTC().Format(http.TimeFormat))
if tc.headCode == 405 {
w.Header().Set(xhttp.AmzDeleteMarker, "true")
}
if tc.headError != "" {
w.Header().Set("x-minio-error-code", tc.headError)
}
w.WriteHeader(tc.headCode)
case http.MethodDelete:
deletes.Add(1)
if got := r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true"; got == purge {
t.Errorf("source delete-marker header=%v, purge=%v", got, purge)
}
w.WriteHeader(tc.deleteCode)
if tc.deleteCode != 204 {
fmt.Fprintf(w, `<Error><Code>%s</Code><Message>injected rejection</Message></Error>`, map[int]string{403: "AccessDenied", 405: "MethodNotAllowed", 503: "ServiceUnavailable"}[tc.deleteCode])
}
default:
t.Errorf("unexpected method %s", r.Method)
}
}))
defer server.Close()
client, err := minio.New(strings.TrimPrefix(server.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
old := globalBucketTargetSys
globalBucketTargetSys = &BucketTargetSys{hc: map[string]epHealth{client.EndpointURL().Host: {Online: !tc.offline}}}
defer func() { globalBucketTargetSys = old }()
d := DeletedObjectReplicationInfo{Bucket: "source", DeletedObject: DeletedObject{ObjectName: "marker", DeleteMarker: shape != "object-purge"}}
if shape == "marker-creation" || shape == "legacy-marker-purge" {
d.DeleteMarkerVersionID = version
} else {
d.VersionID = version
}
d.ReplicationState.Targets = map[string]replication.StatusType{"arn1": tc.creation}
d.ReplicationState.ResetStatusesMap = map[string]string{"arn1": "previous-reset"}
if purge {
d.ReplicationState.PurgeTargets = map[string]VersionPurgeStatusType{"arn1": tc.priorPurge}
}
if tc.resync {
d.OpType = replication.ExistingObjectReplicationType
}
got := replicateDeleteToTarget(t.Context(), d, &TargetClient{Client: client, ARN: "arn1", Bucket: "target", ResetID: "current-reset"})
var wantCreation replication.StatusType
var wantPurge VersionPurgeStatusType
var wantHeads, wantDeletes int32
success := false
if purge {
wantCreation = ""
wantPurge = replication.VersionPurgeComplete
switch {
case tc.priorPurge == replication.VersionPurgeComplete:
success = true
case tc.offline:
wantPurge = replication.VersionPurgeFailed
default:
wantDeletes = 1
success = tc.deleteCode == 204
if !success {
wantPurge = replication.VersionPurgeFailed
}
}
} else {
switch {
case tc.creation == replication.Completed && !tc.resync:
success = true
case tc.offline:
default:
wantHeads = 1
switch tc.headCode {
case 405:
success = true
case 403, 503:
default:
wantDeletes = 1
success = tc.deleteCode == 204
}
}
if success {
wantCreation = replication.Completed
} else {
wantCreation = replication.Failed
}
}
if got.PrevReplicationStatus != tc.creation {
t.Error("previous creation state changed")
}
if got.ReplicationStatus != wantCreation || got.VersionPurgeStatus != wantPurge {
t.Errorf("status=%+v, want creation=%s purge=%s", got, wantCreation, wantPurge)
}
if heads.Load() != wantHeads || deletes.Load() != wantDeletes {
t.Errorf("HEAD/DELETE=%d/%d, want %d/%d", heads.Load(), deletes.Load(), wantHeads, wantDeletes)
}
if (got.Err != nil) == success {
t.Errorf("success=%v, error=%v", success, got.Err)
}
if tc.resync && success {
if !strings.HasSuffix(got.ResyncTimestamp, ";current-reset") {
t.Errorf("successful resync missing reset: %q", got.ResyncTimestamp)
}
} else if got.ResyncTimestamp != "previous-reset" {
t.Errorf("unexpected reset: %q", got.ResyncTimestamp)
}
})
}
}
}
func TestReplicateDeletePurgeMissingTargetState(t *testing.T) {
// Classification belongs to the operation, even if this target has no
// previous purge entry (another target supplies the tracked purge state).
var deletes atomic.Int32
version := mustGetUUID()
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodDelete || r.URL.Query().Get("versionId") != version || r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true" {
t.Errorf("wrong purge request: %s %s %v", r.Method, r.URL, r.Header)
}
deletes.Add(1)
w.WriteHeader(http.StatusNoContent)
}))
defer server.Close()
client, err := minio.New(strings.TrimPrefix(server.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
old := globalBucketTargetSys
globalBucketTargetSys = &BucketTargetSys{hc: map[string]epHealth{client.EndpointURL().Host: {Online: true}}}
defer func() { globalBucketTargetSys = old }()
d := DeletedObjectReplicationInfo{Bucket: "source", DeletedObject: DeletedObject{ObjectName: "marker", DeleteMarker: true, DeleteMarkerVersionID: version, ReplicationState: ReplicationState{Targets: map[string]replication.StatusType{"arn1": replication.Completed}, PurgeTargets: map[string]VersionPurgeStatusType{"arn2": replication.VersionPurgePending}}}}
got := replicateDeleteToTarget(t.Context(), d, &TargetClient{Client: client, ARN: "arn1", Bucket: "target", ResetID: "reset"})
if deletes.Load() != 1 || got.VersionPurgeStatus != replication.VersionPurgeComplete || got.ReplicationStatus != "" {
t.Fatalf("operation misclassified: %+v, deletes=%d", got, deletes.Load())
}
got.ResyncTimestamp = "resync;reset"
state := getReplicationState(replicatedInfos{Targets: []replicatedTargetInfo{got}}, ReplicationState{}, "")
if state.ResetStatusesMap[targetResetHeader("arn1")] != got.ResyncTimestamp {
t.Fatal("resync timestamp not preserved with nil reset map")
}
}
+121
View File
@@ -0,0 +1,121 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"maps"
"net/http"
"net/textproto"
"reflect"
"strings"
"testing"
xhttp "github.com/minio/minio/internal/http"
)
func TestExtractReplicationMetadataPreservesNormalizedMetadata(t *testing.T) {
for _, tc := range []struct {
name string
wire []string
want string
}{
{name: "absent"},
{name: "transport-only", wire: []string{"aws-chunked"}},
{name: "mixed", wire: []string{"aws-chunked,gzip"}, want: "gzip"},
{name: "gzip", wire: []string{"gzip"}, want: "gzip"},
{name: "transport-last", wire: []string{"gzip,aws-chunked"}, want: "gzip"},
{name: "multiple-values", wire: []string{"aws-chunked", "gzip"}, want: "gzip"},
// Preserve the existing exact-token grammar; whitespace is not normalized here.
{name: "space-before-gzip", wire: []string{"aws-chunked, gzip"}, want: " gzip"},
{name: "space-before-transport", wire: []string{"gzip, aws-chunked"}, want: "gzip, aws-chunked"},
} {
for _, lowercase := range []bool{false, true} {
name := tc.name + "/canonical"
if lowercase {
name = tc.name + "/lowercase"
}
t.Run(name, func(t *testing.T) {
header := http.Header{
"Content-Type": []string{"application/octet-stream"},
"X-Amz-Meta-Source": []string{"raw"},
"X-Minio-Replication-Server-Side-Encryption-Sealed-Key": []string{"sealed-key"},
"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm": []string{"DAREv2-HMAC-SHA256"},
"X-Minio-Replication-Server-Side-Encryption-Iv": []string{"iv"},
"X-Minio-Replication-Encrypted-Multipart": []string{""},
"X-Minio-Replication-Actual-Object-Size": []string{"1"},
ReplicationSsecChecksumHeader: []string{"checksum"},
xhttp.AmzMetaUnencryptedContentLength: []string{"injected-length"},
xhttp.AmzMetaUnencryptedContentMD5: []string{"injected-md5"},
}
if tc.wire != nil {
header[xhttp.ContentEncoding] = tc.wire
}
if lowercase {
h := make(http.Header, len(header))
for k, v := range header {
h[strings.ToLower(k)] = v
}
header = h
}
metadata, err := extractMetadata(t.Context(), textproto.MIMEHeader(header))
if err != nil {
t.Fatal(err)
}
if metadata["content-encoding"] != tc.want {
t.Fatalf("ordinary encoding=%q want=%q", metadata["content-encoding"], tc.want)
}
for _, internal := range replicationToInternalHeaders {
if _, ok := metadata[internal]; ok {
t.Fatalf("ordinary request accepted internal field %s", internal)
}
}
// Callers own ordinary metadata and may transform it after extraction.
metadata["content-type"] = "application/wasm"
for k := range metadata {
if strings.EqualFold(k, "x-amz-meta-source") {
metadata[k] = "caller"
}
}
want := maps.Clone(metadata)
maps.Copy(want, map[string]string{
"X-Minio-Internal-Server-Side-Encryption-Sealed-Key": "sealed-key",
"X-Minio-Internal-Server-Side-Encryption-Seal-Algorithm": "DAREv2-HMAC-SHA256",
"X-Minio-Internal-Server-Side-Encryption-Iv": "iv",
"X-Minio-Internal-Encrypted-Multipart": "",
"X-Minio-Internal-Actual-Object-Size": "1",
ReplicationSsecChecksumHeader: "checksum",
})
for range 2 {
if err := extractReplicationMetadataFromMime(t.Context(), textproto.MIMEHeader(header), metadata); err != nil {
t.Fatal(err)
}
if !reflect.DeepEqual(metadata, want) {
t.Errorf("restoration changed normalized metadata: got %#v want %#v", metadata, want)
}
}
if tc.want == "" {
if _, present := metadata["content-encoding"]; present {
t.Error("transport-only content-encoding key restored")
}
}
for _, key := range []string{xhttp.AmzMetaUnencryptedContentLength, xhttp.AmzMetaUnencryptedContentMD5} {
if _, present := caseInsensitiveMap(metadata).Lookup(key); present {
t.Errorf("redacted metadata restored: %s", key)
}
}
})
}
}
}
func TestExtractReplicationMetadataNilHeader(t *testing.T) {
metadata := map[string]string{"content-type": "application/wasm"}
want := maps.Clone(metadata)
if err := extractReplicationMetadataFromMime(t.Context(), nil, metadata); err != errInvalidArgument {
t.Fatalf("nil header: got %v want %v", err, errInvalidArgument)
}
if !reflect.DeepEqual(metadata, want) {
t.Fatalf("nil input changed metadata: %#v", metadata)
}
}
+599
View File
@@ -0,0 +1,599 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"context"
"encoding/xml"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
const r5TagStamp = ReservedMetadataPrefixLower + TaggingTimestamp
func r5Capacity(z *erasureServerPools) func() {
var restores []func()
for _, pool := range z.serverPools {
for _, set := range pool.sets {
old := set.getDisks
disks := append([]StorageAPI(nil), old()...)
for i := range disks {
disks[i] = tagTestCapacityDisk{StorageAPI: disks[i]}
}
set.getDisks = func() []StorageAPI { return disks }
restores = append(restores, func() { set.getDisks = old })
}
}
return func() {
for _, restore := range restores {
restore()
}
}
}
func r5Request(t *testing.T, router http.Handler, cred auth.Credentials, method, path, body string, headers map[string]string) *httptest.ResponseRecorder {
t.Helper()
r, err := newTestSignedRequestV4(method, path, int64(len(body)), strings.NewReader(body), cred.AccessKey, cred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, r)
return w
}
func r5Stored(t *testing.T, obj interface {
GetObjectInfo(context.Context, string, string, ObjectOptions) (ObjectInfo, error)
}, bucket, name, vid, wantTags, wantStamp string,
) ObjectInfo {
t.Helper()
oi, err := obj.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: vid})
if err != nil {
t.Fatal(err)
}
if oi.UserTags != wantTags || oi.UserDefined[r5TagStamp] != wantStamp {
t.Fatalf("%s(%s): tags=%q stamp=%q, want %q %q", name, vid, oi.UserTags, oi.UserDefined[r5TagStamp], wantTags, wantStamp)
}
return oi
}
// The request bytes are signed and enter the real API and storage implementation.
// Equal/stale retransmits may retain the existing 412 duplicate response.
func r5Receive(t *testing.T, obj ObjectLayer, router http.Handler, cred auth.Credentials, bucket, operation string, source ObjectInfo, stamp string, afterInit func()) {
t.Helper()
opts, _, err := putReplicationOpts(t.Context(), "", source)
if err != nil {
t.Fatal(err)
}
if operation == "multipart" {
opts.Internal.SourceMTime = time.Time{}
}
headers := make(map[string]string)
for k, vs := range opts.Header() {
if len(vs) > 0 {
headers[k] = vs[0]
}
}
if stamp == "" {
delete(headers, http.CanonicalHeaderKey(xhttp.MinIOSourceTaggingTimestamp))
delete(headers, xhttp.MinIOSourceTaggingTimestamp)
} else {
headers[http.CanonicalHeaderKey(xhttp.MinIOSourceTaggingTimestamp)] = stamp
}
path := "/" + bucket + "/" + source.Name + "?versionId=" + source.VersionID
var w *httptest.ResponseRecorder
switch operation {
case "copy", "copy-default":
maps.Copy(headers, getCopyObjMetadata(source, ""))
headers[xhttp.AmzCopySource] = "/" + bucket + "/" + source.Name + "?versionId=" + source.VersionID
if operation == "copy" {
headers[xhttp.AmzMetadataDirective] = "REPLACE"
}
w = r5Request(t, router, cred, http.MethodPut, path, "", headers)
case "put":
w = r5Request(t, router, cred, http.MethodPut, path, "data", headers)
case "multipart":
w = r5Request(t, router, cred, http.MethodPost, path+"&uploads", "", headers)
if w.Code == http.StatusPreconditionFailed {
return
}
if w.Code != http.StatusOK {
t.Fatalf("init: %d %s", w.Code, w.Body.String())
}
var init struct {
UploadID string `xml:"UploadId"`
}
if err := xml.Unmarshal(w.Body.Bytes(), &init); err != nil || init.UploadID == "" {
t.Fatalf("init XML: %v %s", err, w.Body.String())
}
mi, err := obj.GetMultipartInfo(t.Context(), bucket, source.Name, init.UploadID, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if stamp != "" && mi.UserDefined[r5TagStamp] != stamp {
t.Fatalf("upload persisted stamp=%q, want %q", mi.UserDefined[r5TagStamp], stamp)
}
partPath := "/" + bucket + "/" + source.Name + "?uploadId=" + url.QueryEscape(init.UploadID)
ph := map[string]string{xhttp.MinIOSourceReplicationRequest: "true"}
part := r5Request(t, router, cred, http.MethodPut, partPath+"&partNumber=1", "data", ph)
if part.Code != http.StatusOK {
t.Fatalf("part: %d %s", part.Code, part.Body.String())
}
if afterInit != nil {
afterInit()
}
body := "<CompleteMultipartUpload><Part><PartNumber>1</PartNumber><ETag>" + canonicalizeETag(part.Header()[xhttp.ETag][0]) + "</ETag></Part></CompleteMultipartUpload>"
ph[xhttp.MinIOSourceMTime] = source.ModTime.Format(time.RFC3339Nano)
ph[xhttp.MinIOSourceETag] = source.ETag
w = r5Request(t, router, cred, http.MethodPost, partPath, body, ph)
}
if w.Code != http.StatusOK && w.Code != http.StatusPreconditionFailed {
t.Fatalf("%s: %d %s", operation, w.Code, w.Body.String())
}
}
func TestAPITaggingReplicationOrdering(t *testing.T) { r5APIOrdering(t, false) }
func TestAPITaggingReplicationOrderingKMS(t *testing.T) { r5APIOrdering(t, true) }
func r5APIOrdering(t *testing.T, encrypted bool) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
if encrypted {
prev := GlobalKMS
GlobalKMS = kms.NewStub("r5-tag-order")
defer func() { GlobalKMS = prev }()
sse := []byte(`<ServerSideEncryptionConfiguration xmlns="http://s3.amazonaws.com/doc/2006-03-01/"><Rule><ApplyServerSideEncryptionByDefault><SSEAlgorithm>aws:kms</SSEAlgorithm><KMSMasterKeyID>r5-tag-order</KMSMasterKeyID></ApplyServerSideEncryptionByDefault></Rule></ServerSideEncryptionConfiguration>`)
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketSSEConfig, sse); err != nil {
t.Fatal(err)
}
}
base := time.Now().UTC().Add(-5 * time.Hour)
for _, op := range []string{"copy", "copy-default", "put", "multipart"} {
for _, version := range []string{"uuid", "null"} {
t.Run(instance+"/"+op+"/"+version, func(t *testing.T) {
name := op + "-" + version
vid := mustGetUUID()
if version == "null" {
vid = nullVersionID
}
original, err := obj.PutObject(t.Context(), bucket, name, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true, VersionID: vid, UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=original", r5TagStamp: base.Format(time.RFC3339Nano)}})
if err != nil {
t.Fatal(err)
}
if op == "multipart" {
// The production sender uses multipart only for multipart
// sources. Preserve a real multipart ETag and part layout.
metadata := maps.Clone(original.UserDefined)
metadata[xhttp.AmzObjectTagging] = original.UserTags
mp, err := obj.NewMultipartUpload(t.Context(), bucket, name, ObjectOptions{Versioned: true, VersionID: vid, UserDefined: metadata})
if err != nil {
t.Fatal(err)
}
part, err := obj.PutObjectPart(t.Context(), bucket, name, mp.UploadID, 1, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
original, err = obj.CompleteMultipartUpload(t.Context(), bucket, name, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
}
// A later unrelated version must not contribute tags to an explicitly addressed UUID/null version.
latest, err := obj.PutObject(t.Context(), bucket, name, mustGetPutObjReader(t, strings.NewReader("other"), 5, "", ""), ObjectOptions{Versioned: true, UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=latest", r5TagStamp: base.Add(10 * time.Hour).Format(time.RFC3339Nano)}})
if err != nil {
t.Fatal(err)
}
for _, event := range []struct {
name, tags string
hours int
wantTags string
wantHours int
}{
{"delete", "", 3, "", 3}, {"stale", "key=stale", 2, "", 3}, {"equal-conflict", "key=conflict", 3, "", 3}, {"newer", "key=new", 4, "key=new", 4},
} {
t.Run(event.name, func(t *testing.T) {
source := original
source.VersionID = vid
source.UserDefined = maps.Clone(original.UserDefined)
source.UserTags = event.tags
stamp := base.Add(time.Duration(event.hours) * time.Hour).Format(time.RFC3339Nano)
source.UserDefined[r5TagStamp] = stamp
r5Receive(t, obj, router, cred, bucket, op, source, stamp, nil)
r5Stored(t, obj, bucket, name, vid, event.wantTags, base.Add(time.Duration(event.wantHours)*time.Hour).Format(time.RFC3339Nano))
r5Stored(t, obj, bucket, name, latest.VersionID, "key=latest", base.Add(10*time.Hour).Format(time.RFC3339Nano))
})
}
source := original
source.VersionID = vid
source.UserTags = "key=unversioned-event"
r5Receive(t, obj, router, cred, bucket, op, source, "", nil)
r5Stored(t, obj, bucket, name, vid, "key=new", base.Add(4*time.Hour).Format(time.RFC3339Nano))
get := r5Request(t, router, cred, http.MethodGet, "/"+bucket+"/"+name+"?versionId="+vid, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "data" {
t.Fatalf("plaintext GET: %d %q", get.Code, get.Body.String())
}
})
}
}
}})
}
func TestAPITaggingMultipartCommitRechecksRevision(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
base := time.Now().UTC().Add(-time.Hour)
source := ObjectInfo{Name: "commit-recheck", VersionID: mustGetUUID(), ModTime: base, UserTags: "key=incoming", UserDefined: map[string]string{r5TagStamp: base.Format(time.RFC3339Nano)}}
later := base.Add(time.Minute).Format(time.RFC3339Nano)
r5Receive(t, obj, router, cred, bucket, "multipart", source, base.Format(time.RFC3339Nano), func() {
_, err := obj.PutObject(t.Context(), bucket, source.Name, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true, VersionID: source.VersionID, MTime: base, UserDefined: map[string]string{r5TagStamp: later}})
if err != nil {
t.Fatal(err)
}
})
r5Stored(t, obj, bucket, source.Name, source.VersionID, "", later)
t.Logf("%s: deletion committed between initiation and completion survived", instance)
}})
}
func TestAPILocalTaggingAlwaysAdvancesRevision(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
oi, err := obj.PutObject(t.Context(), bucket, "local-tags", mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
path := "/" + bucket + "/local-tags?tagging&versionId=" + oi.VersionID
for n, method := range []string{http.MethodPut, http.MethodDelete, http.MethodDelete, http.MethodPut} {
body := ""
want := ""
status := http.StatusNoContent
if method == http.MethodPut {
status = http.StatusOK
body = "<Tagging><TagSet/></Tagging>"
if n == 0 {
want = "key=local"
body = "<Tagging><TagSet><Tag><Key>key</Key><Value>local</Value></Tag></TagSet></Tagging>"
}
}
before := time.Now().UTC()
w := r5Request(t, router, cred, method, path, body, nil)
if w.Code != status {
t.Fatalf("%s: %d %s", method, w.Code, w.Body.String())
}
now, err := obj.GetObjectInfo(t.Context(), bucket, "local-tags", ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
stamp, err := time.Parse(time.RFC3339Nano, now.UserDefined[r5TagStamp])
if err != nil || stamp.Before(before) || now.UserTags != want {
t.Fatalf("local %s: tags=%q timestamp=%q err=%v", method, now.UserTags, now.UserDefined[r5TagStamp], err)
}
if !now.ModTime.Equal(oi.ModTime) {
t.Fatal("tagging changed the object's data modification time")
}
t.Logf("%s mutation %d persisted tags=%q timestamp=%s", instance, n, now.UserTags, stamp)
}
// Ordinary COPY with empty REPLACE has the same local-deletion semantics.
before := time.Now().UTC()
headers := map[string]string{xhttp.AmzCopySource: "/" + bucket + "/local-tags?versionId=" + oi.VersionID, xhttp.AmzMetadataDirective: "REPLACE", xhttp.AmzTagDirective: "REPLACE"}
w := r5Request(t, router, cred, http.MethodPut, "/"+bucket+"/copied-empty", "", headers)
if w.Code != http.StatusOK {
t.Fatalf("copy: %d %s", w.Code, w.Body.String())
}
copied, err := obj.GetObjectInfo(t.Context(), bucket, "copied-empty", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
stamp, err := time.Parse(time.RFC3339Nano, copied.UserDefined[r5TagStamp])
if err != nil || stamp.Before(before) || copied.UserTags != "" {
t.Fatalf("local COPY: %v %+v", err, copied)
}
}})
}
func TestTaggingTimestampWire(t *testing.T) {
now := time.Now().UTC()
for _, tag := range []string{"", "key=value"} {
for _, stamp := range []string{"", now.Format(time.RFC3339Nano), "invalid"} {
t.Run(fmt.Sprintf("%s/%s", tag, stamp), func(t *testing.T) {
source := ObjectInfo{ModTime: now.Add(-time.Hour), UserTags: tag, UserDefined: map[string]string{}}
if stamp != "" {
source.UserDefined[r5TagStamp] = stamp
}
opts, _, err := putReplicationOpts(t.Context(), "", source)
if stamp == "invalid" {
if err == nil {
t.Fatal("invalid timestamp accepted")
}
return
}
if err != nil {
t.Fatal(err)
}
want := stamp
if want == "" && tag != "" {
want = source.ModTime.Format(time.RFC3339Nano)
}
if got := opts.Header().Get(xhttp.MinIOSourceTaggingTimestamp); got != want {
t.Fatalf("wire timestamp=%q want %q", got, want)
}
})
}
}
}
func TestAPIPoolsTaggingReplicaDeletion(t *testing.T) {
z, bucket := consistencyPools(t)
defer r5Capacity(z)()
if err := newTestConfig(globalMinioDefaultRegion, z); err != nil {
t.Fatal(err)
}
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
router := initTestAPIEndPoints(z, nil)
base := time.Now().UTC().Add(-5 * time.Hour)
for _, op := range []string{"copy", "copy-default", "put", "multipart"} {
for _, kind := range []string{"uuid", "null"} {
t.Run(op+"/"+kind, func(t *testing.T) {
vid := mustGetUUID()
if kind == "null" {
vid = nullVersionID
}
name := "pool-tags-" + op + "-" + kind
var source ObjectInfo
for pool := range 2 {
tag := "key=stale-pool"
ts := base
if pool == 1 {
tag = ""
ts = base.Add(3 * time.Hour)
}
source = putConsistencyObject(t, z, bucket, name, pool, "data", ObjectOptions{Versioned: true, VersionID: vid, MTime: base, UserDefined: map[string]string{xhttp.AmzObjectTagging: tag, r5TagStamp: ts.Format(time.RFC3339Nano)}})
}
// Clear all copies through the signed local handler, then replay a stale incoming state.
req := r5Request(t, router, globalActiveCred, http.MethodDelete, "/"+bucket+"/"+name+"?tagging&versionId="+vid, "", nil)
if req.Code != http.StatusNoContent {
t.Fatalf("DELETE: %d %s", req.Code, req.Body.String())
}
current, err := z.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: vid})
if err != nil {
t.Fatal(err)
}
deletedAt := current.UserDefined[r5TagStamp]
for pool := range 2 {
r5Stored(t, z.serverPools[pool], bucket, name, vid, "", deletedAt)
}
source.VersionID = vid
source.UserTags = "key=delayed"
source.UserDefined = map[string]string{r5TagStamp: base.Add(time.Hour).Format(time.RFC3339Nano)}
r5Receive(t, z, router, globalActiveCred, bucket, op, source, source.UserDefined[r5TagStamp], nil)
// The addressed version must remain readable through normal routing.
r5Stored(t, z, bucket, name, vid, "", deletedAt)
// Existing duplicate suppression may leave both identical tombstones; any retained copy must be correct.
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: vid})
if isErrVersionNotFound(err) {
continue
}
if err != nil {
t.Fatal(err)
}
if got.UserTags != "" || got.UserDefined[r5TagStamp] != deletedAt {
t.Fatalf("pool %d: tags=%q stamp=%q want deletion %q", pool, got.UserTags, got.UserDefined[r5TagStamp], deletedAt)
}
}
})
}
}
}
func TestAPITaggingSSECRotationPreservesDeletionRevision(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
oldTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = oldTLS }()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
keys := [][]byte{[]byte(strings.Repeat("a", 32)), []byte(strings.Repeat("b", 32)), []byte(strings.Repeat("c", 32)), []byte(strings.Repeat("d", 32))}
const name = "tag-rotation"
putCopyChecksumSource(t, router, cred, bucket, name, []byte("data"), ssecKeyHeaders(keys[0], false))
oi, err := obj.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
base := time.Now().UTC().Add(-time.Hour)
for i, event := range []struct {
tags string
delta int
wantTags string
wantDelta int
}{
{"key=live", 1, "key=live", 1}, {"", 3, "", 3}, {"key=stale", 2, "", 3},
} {
h := ssecKeyHeaders(keys[i], true)
maps.Copy(h, ssecKeyHeaders(keys[i+1], false))
h[xhttp.AmzObjectTagging] = event.tags
h[xhttp.AmzTagDirective] = "REPLACE"
h[xhttp.MinIOSourceTaggingTimestamp] = base.Add(time.Duration(event.delta) * time.Minute).Format(time.RFC3339Nano)
sendReplicaLockCopy(t, router, cred, bucket, name, oi.VersionID, h)
r5Stored(t, obj, bucket, name, oi.VersionID, event.wantTags, base.Add(time.Duration(event.wantDelta)*time.Minute).Format(time.RFC3339Nano))
}
get := r5Request(t, router, cred, http.MethodGet, "/"+bucket+"/"+name+"?versionId="+oi.VersionID, "", ssecKeyHeaders(keys[3], false))
if get.Code != http.StatusOK || get.Body.String() != "data" {
t.Fatalf("%s GET after rotations: %d %q", instance, get.Code, get.Body.String())
}
}})
}
func TestTaggingRepeatedValueNeedsRevisionDelivery(t *testing.T) {
for _, tag := range []string{"", "key=same"} {
now := time.Now().UTC()
source := ObjectInfo{ModTime: now, UserTags: tag, UserDefined: map[string]string{r5TagStamp: now.Add(time.Hour).Format(time.RFC3339Nano)}}
target := minio.ObjectInfo{LastModified: now}
if tag != "" {
target.UserTags = map[string]string{"key": "same"}
target.UserTagCount = 1
}
if got := getReplicationAction(source, target, replication.MetadataReplicationType); got != replicateMetadata {
t.Errorf("equal tags %q suppress a newer revision: got %s; an intervening delayed deletion can win", tag, got)
}
}
}
func TestTaggingProductionCopyWireShape(t *testing.T) {
got := make(chan http.Header, 1)
peer := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
got <- r.Header.Clone()
w.Header().Set(xhttp.ContentType, "application/xml")
w.Write([]byte(`<CopyObjectResult><LastModified>2026-09-15T01:00:00Z</LastModified><ETag>"abc"</ETag></CopyObjectResult>`))
}))
defer peer.Close()
c, err := minio.New(strings.TrimPrefix(peer.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
now := time.Now().UTC()
oi := ObjectInfo{Name: "object", ModTime: now, ETag: "abc", VersionID: mustGetUUID()}
core := minio.Core{Client: c}
_, err = core.CopyObject(t.Context(), "bucket", "object", "bucket", "object", getCopyObjMetadata(oi, ""), minio.CopySrcOptions{VersionID: oi.VersionID}, minio.PutObjectOptions{Internal: minio.AdvancedPutOptions{SourceVersionID: oi.VersionID, ReplicationRequest: true, TaggingTimestamp: now}})
if err != nil {
t.Fatal(err)
}
h := <-got
t.Logf("SDK wire: metadata-directive=%q tagging-directive=%q tagging=%q time=%q", h.Get(xhttp.AmzMetadataDirective), h.Get(xhttp.AmzTagDirective), h.Get(xhttp.AmzObjectTagging), h.Get(xhttp.MinIOSourceTaggingTimestamp))
if h.Get(xhttp.AmzMetadataDirective) != "" || h.Get(xhttp.AmzTagDirective) != "REPLACE" || h.Get(xhttp.MinIOSourceTaggingTimestamp) != now.Format(time.RFC3339Nano) {
t.Fatalf("unexpected SDK wire: %v", h)
}
}
func TestLocalTaggingCommitCannotRegressRevision(t *testing.T) {
z, bucket := consistencyPools(t)
incoming := "2026-09-15T01:00:00Z"
newer := "2026-09-15T02:00:00Z"
newest := "2026-09-15T03:00:00Z"
vid := mustGetUUID()
name := "inverted-local-tags"
for pool := range 2 {
putConsistencyObject(t, z, bucket, name, pool, "data", ObjectOptions{Versioned: true, VersionID: vid, UserDefined: map[string]string{r5TagStamp: []string{newer, newest}[pool]}})
}
oi, err := z.PutObjectTags(t.Context(), bucket, name, "key=after-delete", ObjectOptions{VersionID: vid, UserDefined: map[string]string{r5TagStamp: incoming}})
if err != nil {
t.Fatal(err)
}
want := "2026-09-15T03:00:00.000000001Z"
for pool := range 2 {
r5Stored(t, z.serverPools[pool], bucket, name, vid, "key=after-delete", want)
}
if oi.UserDefined[r5TagStamp] != want {
t.Fatalf("response stamp=%q want %q", oi.UserDefined[r5TagStamp], want)
}
// The single-set guard also applies when the pools dispatcher is bypassed.
_, err = z.serverPools[0].PutObjectTags(t.Context(), bucket, name, "", ObjectOptions{VersionID: vid, UserDefined: map[string]string{r5TagStamp: incoming}})
if err != nil {
t.Fatal(err)
}
r5Stored(t, z.serverPools[0], bucket, name, vid, "", "2026-09-15T03:00:00.000000002Z")
}
func TestTaggingReplicaContentDuplicateGuard(t *testing.T) {
stamp := time.Date(2026, 9, 15, 1, 0, 0, 0, time.UTC)
for _, tc := range []struct {
name string
delta int
stored string
trusted, replica bool
ifMatch, ifNone string
wantSkip bool
}{
{name: "newer", delta: 1, trusted: true, replica: true},
{name: "equal", trusted: true, replica: true, wantSkip: true},
{name: "older", delta: -1, trusted: true, replica: true, wantSkip: true},
{name: "invalid-stored", stored: "invalid", delta: 1, trusted: true, replica: true},
{name: "untrusted", delta: 1, wantSkip: true},
{name: "marker-only", delta: 1, trusted: true, wantSkip: true},
{name: "if-match-fails", delta: 1, trusted: true, replica: true, ifMatch: "other", wantSkip: true},
{name: "if-none-match-fails", delta: 1, trusted: true, replica: true, ifNone: "etag", wantSkip: true},
{name: "if-match-passes", delta: 1, trusted: true, replica: true, ifMatch: "etag"},
} {
t.Run(tc.name, func(t *testing.T) {
ctx := withReplicationTrust(t.Context(), tc.trusted, tc.replica)
r := httptest.NewRequest(http.MethodPut, "/bucket/object", nil).WithContext(ctx)
if tc.ifMatch != "" {
r.Header.Set(xhttp.IfMatch, tc.ifMatch)
}
if tc.ifNone != "" {
r.Header.Set(xhttp.IfNoneMatch, tc.ifNone)
}
stored := tc.stored
if stored == "" {
stored = stamp.Format(time.RFC3339Nano)
}
oi := ObjectInfo{ModTime: stamp, VersionID: mustGetUUID(), ETag: "etag", UserDefined: map[string]string{r5TagStamp: stored}}
opts := ObjectOptions{VersionID: oi.VersionID, PreserveETag: oi.ETag, ReplicationRequest: tc.trusted, ReplicationSourceTaggingTimestamp: stamp.Add(time.Duration(tc.delta) * time.Second)}
if skip := checkPreconditionsPUT(ctx, httptest.NewRecorder(), r, oi, opts); skip != tc.wantSkip {
t.Fatalf("skip=%v, want %v", skip, tc.wantSkip)
}
})
}
}
func TestAPITaggingUnqualifiedCopyOrdering(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
stamp := time.Now().UTC().Add(-time.Hour)
oi, err := obj.PutObject(t.Context(), bucket, "unqualified-tags", mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=stored", r5TagStamp: stamp.Format(time.RFC3339Nano)}})
if err != nil {
t.Fatal(err)
}
source := oi
source.UserTags = ""
source.UserDefined = maps.Clone(oi.UserDefined)
r5Receive(t, obj, router, cred, bucket, "copy", source, stamp.Format(time.RFC3339Nano), nil)
r5Stored(t, obj, bucket, oi.Name, "", "key=stored", stamp.Format(time.RFC3339Nano))
later := stamp.Add(time.Minute).Format(time.RFC3339Nano)
r5Receive(t, obj, router, cred, bucket, "copy-default", source, later, nil)
r5Stored(t, obj, bucket, oi.Name, "", "", later)
source.UserTags = "key=delayed"
r5Receive(t, obj, router, cred, bucket, "copy-default", source, stamp.Format(time.RFC3339Nano), nil)
r5Stored(t, obj, bucket, oi.Name, "", "", later)
t.Logf("%s: unqualified COPY keeps stored ties and ordered deletion", instance)
}})
}
+154
View File
@@ -0,0 +1,154 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/once"
)
func r5ReplicationFixture(t *testing.T, obj ObjectLayer, bucket string, client *minio.Client) (chan ReplicationWorkerOperation, func()) {
t.Helper()
const arn = "arn:minio:replication::af470089-d354-4473-934c-9e1f52f6da89:bucket"
target := &TargetClient{Client: client, ARN: arn, Bucket: bucket}
globalBucketTargetSys.arnRemotesMap[arn] = arnTarget{Client: target, lastRefresh: UTCNow()}
globalBucketTargetSys.targetsMap[bucket] = []madmin.BucketTarget{{Arn: arn, TargetBucket: bucket}}
meta, err := globalBucketMetadataSys.Get(bucket)
if err != nil {
t.Fatal(err)
}
cfg := configs[0]
cfg.RoleArn = arn
meta.replicationConfig = &cfg
globalBucketMetadataSys.Set(bucket, meta)
worker := make(chan ReplicationWorkerOperation, 10)
previous := globalReplicationPool
globalReplicationPool = once.NewSingleton[ReplicationPool]()
globalReplicationPool.Set(&ReplicationPool{ctx: t.Context(), objLayer: obj, workers: []chan ReplicationWorkerOperation{worker}, stats: globalReplicationStats.Load(), mrfSaveCh: make(chan MRFReplicateEntry, 10)})
return worker, func() { globalReplicationPool = previous }
}
func TestTaggingReplicationSenderRetryAndAcknowledgment(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
const name = "tagging-old-queue"
oi, err := obj.PutObject(t.Context(), bucket, name, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
requests := make(chan http.Header, 4)
var attempts atomic.Int32
peer := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set(xhttp.AmzVersionID, oi.VersionID)
w.Header().Set(xhttp.ETag, "\""+oi.ETag+"\"")
w.Header().Set(xhttp.LastModified, oi.ModTime.Format(http.TimeFormat))
w.Header().Set(xhttp.ContentType, oi.ContentType)
if r.Method == http.MethodHead {
w.Header().Set(xhttp.ContentLength, "4")
w.WriteHeader(http.StatusOK)
return
}
requests <- r.Header.Clone()
w.Header().Set(xhttp.ContentType, "application/xml")
if attempts.Add(1) == 1 {
w.WriteHeader(http.StatusServiceUnavailable)
w.Write([]byte(`<Error><Code>SlowDown</Code><Message>retry fixture</Message></Error>`))
return
}
w.Write([]byte("<CopyObjectResult><LastModified>" + oi.ModTime.Format(time.RFC3339Nano) + "</LastModified><ETag>\"" + oi.ETag + "\"</ETag></CopyObjectResult>"))
}))
defer peer.Close()
client, err := minio.New(strings.TrimPrefix(peer.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
worker, cleanup := r5ReplicationFixture(t, obj, bucket, client)
defer cleanup()
body := `<Tagging><TagSet><Tag><Key>key</Key><Value>queued</Value></Tag></TagSet></Tagging>`
w := r5Request(t, router, cred, http.MethodPut, "/"+bucket+"/"+name+"?tagging&versionId="+oi.VersionID, body, nil)
if w.Code != http.StatusOK || len(worker) != 1 {
t.Fatalf("tagging PUT: %d %s queued=%d", w.Code, w.Body.String(), len(worker))
}
old := (<-worker).(ReplicateObjectInfo)
// Delete through the actual handler before processing the old task.
w = r5Request(t, router, cred, http.MethodDelete, "/"+bucket+"/"+name+"?tagging&versionId="+oi.VersionID, "", nil)
if w.Code != http.StatusNoContent || len(worker) != 1 {
t.Fatalf("tagging DELETE: %d %s queued=%d", w.Code, w.Body.String(), len(worker))
}
deleted, err := obj.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
stamp := deleted.UserDefined[r5TagStamp]
for attempt := 0; attempt < 2; attempt++ {
result := replicateObject(t.Context(), old, obj)
want := replication.Failed
if attempt == 1 {
want = replication.Completed
}
if result.ReplicationStatus() != want {
t.Fatalf("attempt %d result=%+v want %s", attempt, result, want)
}
if len(result.Targets) != 1 || result.Targets[0].ReplicationAction != replicateMetadata || (result.Targets[0].Err != nil) != (attempt == 0) {
t.Fatalf("attempt %d reported wrong action/error: %+v", attempt, result)
}
r5Stored(t, obj, bucket, name, oi.VersionID, "", stamp)
select {
case h := <-requests:
if h.Get(xhttp.MinIOSourceTaggingTimestamp) != stamp || h.Get(xhttp.AmzObjectTagging) != "" || h.Get(xhttp.AmzTagDirective) != "REPLACE" || h.Get(xhttp.AmzMetadataDirective) != "" {
t.Fatalf("sender did not carry current deletion: %v", h)
}
t.Logf("%s attempt %d sent tags=%q timestamp=%s status=%s", instance, attempt, h.Get(xhttp.AmzObjectTagging), stamp, want)
default:
t.Fatal("no metadata COPY sent for same-empty target")
}
}
// With a real outgoing rule enabled, a signed incoming replica COPY
// must not queue another outgoing event and create a feedback loop.
before := len(worker)
r5Receive(t, obj, router, cred, bucket, "copy", deleted, stamp, nil)
if len(worker) != before {
t.Fatal("incoming replica COPY scheduled another outgoing event")
}
// The metadata sender must fail malformed stored revisions before COPY,
// just as the full retransmission option builder does.
_, err = obj.PutObjectTags(t.Context(), bucket, name, "", ObjectOptions{VersionID: oi.VersionID, UserDefined: map[string]string{r5TagStamp: "invalid"}})
if err != nil {
t.Fatal(err)
}
target := globalBucketTargetSys.GetRemoteTargetClient(bucket, globalBucketTargetSys.targetsMap[bucket][0].Arn)
invalid := old.replicateAll(t.Context(), obj, target)
if invalid.ReplicationStatus != replication.Failed || invalid.Err == nil || len(requests) != 0 {
t.Fatalf("invalid timestamp was not rejected before send: %+v", invalid)
}
}})
}
+2
View File
@@ -901,6 +901,8 @@ func serverMain(ctx *cli.Context) {
close(globalGridStart)
close(globalLockGridStart)
// The HTTP/1 listener preserves absolute header deadlines and renews the
// body read/write idle limits, so transfers may outlast IdleTimeout.
httpServer := xhttp.NewServer(getServerListenAddrs()).
UseHandler(setCriticalErrorHandler(corsHandler(handler))).
UseTLSConfig(newTLSConfig(getCert)).
+97
View File
@@ -0,0 +1,97 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"os"
"testing"
"time"
"github.com/minio/cli"
xhttp "github.com/minio/minio/internal/http"
)
func TestServerReadHeaderTimeoutConfig(t *testing.T) {
for _, tc := range []struct {
name, env, idle string
args []string
want time.Duration
fmtgen bool
}{
{name: "default", want: xhttp.DefaultReadHeaderTimeout},
{name: "flag", args: []string{"--read-header-timeout=100ms"}, want: 100 * time.Millisecond},
{name: "environment", env: "170ms", want: 170 * time.Millisecond},
{name: "flag-over-environment", env: "170ms", args: []string{"--read-header-timeout=80ms"}, want: 80 * time.Millisecond},
{name: "yaml-retains-flag", args: []string{"--config=testdata/config/1.yaml", "--read-header-timeout=100ms"}, want: 100 * time.Millisecond},
{name: "zero-fallback", args: []string{"--read-header-timeout=0s"}},
{name: "zero-idle-default-header", idle: "0s", want: xhttp.DefaultReadHeaderTimeout},
{name: "negative-idle-default-header", idle: "-1s", want: xhttp.DefaultReadHeaderTimeout},
{name: "fmt-gen-unregistered-duration", env: "100ms", fmtgen: true},
{name: "negative-disabled", args: []string{"--read-header-timeout=-1s"}, want: -time.Second},
} {
t.Run(tc.name, func(t *testing.T) {
for _, key := range []string{"MINIO_ARGS", "MINIO_VOLUMES", "MINIO_ENDPOINTS", "MINIO_CONFIG", "MINIO_ERASURE_SET_DRIVE_COUNT"} {
t.Setenv(key, "")
}
t.Setenv("MINIO_READ_HEADER_TIMEOUT", tc.env)
if tc.env == "" {
if err := os.Unsetenv("MINIO_READ_HEADER_TIMEOUT"); err != nil {
t.Fatal(err)
}
}
idle := tc.idle
if idle == "" {
idle = "2s"
}
t.Setenv("MINIO_IDLE_TIMEOUT", idle)
idleWant, err := time.ParseDuration(idle)
if err != nil {
t.Fatal(err)
}
commandName, flags := "server", serverCmd.Flags
if tc.fmtgen {
commandName, flags = "fmt-gen", fmtGenFlags
idleWant = 0
}
var got serverCtxt
called := false
app := cli.NewApp()
app.Commands = []cli.Command{{Name: commandName, Flags: flags, Action: func(ctx *cli.Context) error {
called = true
if parsed := ctx.Duration("read-header-timeout"); parsed != tc.want {
t.Errorf("CLI read-header-timeout=%s, expected=%s", parsed, tc.want)
}
return buildServerCtxt(ctx, &got)
}}}
args := append([]string{"silo", commandName}, tc.args...)
args = append(args, t.TempDir())
if err := app.Run(args); err != nil {
t.Fatal(err)
}
if !called {
t.Fatal("server action did not run")
}
if got.ReadHeaderTimeout != tc.want {
t.Errorf("parsed ReadHeaderTimeout = %s, want %s", got.ReadHeaderTimeout, tc.want)
}
if got.IdleTimeout != idleWant {
t.Errorf("parsed IdleTimeout = %s, want %s", got.IdleTimeout, idleWant)
}
})
}
}
@@ -274,6 +274,10 @@ func TestBucketMetadataInitialSyncPhysicalCreated(t *testing.T) {
}
events = append(events, event)
}
if r.URL.Path == "/minio/admin/v3/site-replication/peer/iam-revisions" {
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "initial-peer", Instance: "initial-boot", Digest: "ack"}})
return
}
w.WriteHeader(http.StatusOK)
}))
defer peer.Close()
+87 -34
View File
@@ -28,6 +28,7 @@ import (
"fmt"
"maps"
"math/rand"
"net/http"
"net/url"
"reflect"
"runtime"
@@ -209,7 +210,11 @@ type SiteReplicationSys struct {
// In-memory and persisted multi-site replication state.
state srState
iamMetaCache srIAMCache
iamMetaCache srIAMCache
healOnce sync.Once // Configuration reloads must not spawn more healing loops.
iamHealMu sync.Mutex
iamRevisionProgress map[string]iamRevisionProgress
iamRevisionMetrics iamRevisionMetrics
}
type srState srStateV1
@@ -234,7 +239,7 @@ type srStateData struct {
// Init - initialize the site replication manager.
func (c *SiteReplicationSys) Init(ctx context.Context, objAPI ObjectLayer) error {
go c.startHealRoutine(ctx, objAPI)
c.healOnce.Do(func() { go c.startHealRoutine(ctx, objAPI) })
r := rand.New(rand.NewSource(time.Now().UnixNano()))
for {
err := c.loadFromDisk(ctx, objAPI)
@@ -1257,13 +1262,36 @@ func (c *SiteReplicationSys) IAMChangeHook(ctx context.Context, item madmin.SRIA
return nil
}
if path := iamDeletionPath(item); path != "" {
r, err := loadIAMRevision(ctx, globalIAMSys.store, path)
if err != nil {
return err
}
if !r.Deleted && ((item.Type != madmin.SRIAMItemIAMUser && item.Type != madmin.SRIAMItemGroupInfo) || r.RevokedBefore.IsZero()) {
// A concurrent recreation has already superseded this delete.
return nil
}
item.UpdatedAt = r.timestamp()
if !r.Deleted {
item.UpdatedAt = r.RevokedBefore
}
}
versioned, err := c.replicationItem(ctx, item)
if errors.Is(err, errIAMStaleUpdate) {
return nil
}
if err != nil {
return err
}
cerr := c.concDo(nil, func(d string, p madmin.PeerInfo) error {
admClient, err := c.getAdminClient(ctx, d)
if err != nil {
return wrapSRErr(err)
}
return c.annotatePeerErr(p.Name, replicateIAMItem, admClient.SRPeerReplicateIAMItem(ctx, item))
_, err = executeIAMRevisionRequest(ctx, admClient, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: []iamReplicationItem{versioned}})
return c.annotatePeerErr(p.Name, replicateIAMItem, err)
},
replicateIAMItem,
)
@@ -1273,6 +1301,7 @@ func (c *SiteReplicationSys) IAMChangeHook(ctx context.Context, item madmin.SRIA
// PeerAddPolicyHandler - copies IAM policy to local. A nil policy argument,
// causes the named policy to be deleted.
func (c *SiteReplicationSys) PeerAddPolicyHandler(ctx context.Context, policyName string, p *policy.Policy, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
var err error
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
@@ -1286,18 +1315,19 @@ func (c *SiteReplicationSys) PeerAddPolicyHandler(ctx context.Context, policyNam
_, err = globalIAMSys.SetPolicy(ctx, policyName, *p)
}
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerIAMUserChangeHandler - copies IAM user to local.
func (c *SiteReplicationSys) PeerIAMUserChangeHandler(ctx context.Context, change *madmin.SRIAMUser, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if change == nil {
return errSRInvalidRequest(errInvalidArgument)
}
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
if !change.IsDeleteReq && !updatedAt.IsZero() {
if ui, err := globalIAMSys.GetUserInfo(ctx, change.AccessKey); err == nil && ui.UpdatedAt.After(updatedAt) {
return nil
}
@@ -1326,13 +1356,14 @@ func (c *SiteReplicationSys) PeerIAMUserChangeHandler(ctx context.Context, chang
}
}
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerGroupInfoChangeHandler - copies group changes to local.
func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, change *madmin.SRGroupInfo, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if change == nil {
return errSRInvalidRequest(errInvalidArgument)
}
@@ -1340,7 +1371,7 @@ func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, cha
var err error
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
if !updatedAt.IsZero() && (!updReq.IsRemove || len(updReq.Members) != 0) {
if gd, err := globalIAMSys.GetGroupDescription(updReq.Group); err == nil && gd.UpdatedAt.After(updatedAt) {
return nil
}
@@ -1349,7 +1380,8 @@ func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, cha
if updReq.IsRemove {
_, err = globalIAMSys.RemoveUsersFromGroup(ctx, updReq.Group, updReq.Members)
} else {
if updReq.Status != "" && len(updReq.Members) == 0 {
snapshot, _ := ctx.Value(iamGroupSnapshotKey{}).(bool)
if !snapshot && updReq.Status != "" && len(updReq.Members) == 0 {
_, err = globalIAMSys.SetGroupStatus(ctx, updReq.Group, updReq.Status == madmin.GroupEnabled)
} else {
if globalIAMSys.LDAPConfig.Enabled() {
@@ -1365,13 +1397,14 @@ func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, cha
}
}
if err != nil && !errors.Is(err, errNoSuchGroup) {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerSvcAccChangeHandler - copies service-account change to local.
func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change *madmin.SRSvcAccChange, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if change == nil {
return errSRInvalidRequest(errInvalidArgument)
}
@@ -1382,7 +1415,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
if len(change.Create.SessionPolicy) > 0 {
sp, err = policy.ParseConfig(bytes.NewReader(change.Create.SessionPolicy))
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
// skip overwrite of local update if peer sent stale info
@@ -1394,6 +1427,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
opts := newServiceAccountOpts{
accessKey: change.Create.AccessKey,
secretKey: change.Create.SecretKey,
status: change.Create.Status,
sessionPolicy: sp,
claims: change.Create.Claims,
name: change.Create.Name,
@@ -1402,7 +1436,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
}
_, _, err = globalIAMSys.NewServiceAccount(ctx, change.Create.Parent, change.Create.Groups, opts)
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
case change.Update != nil:
@@ -1411,7 +1445,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
if len(change.Update.SessionPolicy) > 0 {
sp, err = policy.ParseConfig(bytes.NewReader(change.Update.SessionPolicy))
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
// skip overwrite of local update if peer sent stale info
@@ -1431,7 +1465,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
_, err = globalIAMSys.UpdateServiceAccount(ctx, change.Update.AccessKey, opts)
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
case change.Delete != nil:
@@ -1442,7 +1476,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
}
}
if err := globalIAMSys.DeleteServiceAccount(ctx, change.Delete.AccessKey, true); err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
@@ -1451,24 +1485,17 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
// PeerPolicyMappingHandler - copies policy mapping to local.
func (c *SiteReplicationSys) PeerPolicyMappingHandler(ctx context.Context, mapping *madmin.SRPolicyMapping, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if mapping == nil {
return errSRInvalidRequest(errInvalidArgument)
}
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
mp, ok := globalIAMSys.store.GetMappedPolicy(mapping.Policy, mapping.IsGroup)
if ok && mp.UpdatedAt.After(updatedAt) {
return nil
}
}
// When LDAP is enabled, we verify that the user or group exists in LDAP and
// use the normalized form of the entityName (which will be an LDAP DN).
userType := IAMUserType(mapping.UserType)
isGroup := mapping.IsGroup
entityName := mapping.UserOrGroup
if globalIAMSys.GetUsersSysType() == LDAPUsersSysType && userType == stsUser {
if mapping.Policy != "" && globalIAMSys.GetUsersSysType() == LDAPUsersSysType && userType == stsUser {
// Validate that the user or group exists in LDAP and use the normalized
// form of the entityName (which will be an LDAP DN).
var err error
@@ -1491,19 +1518,20 @@ func (c *SiteReplicationSys) PeerPolicyMappingHandler(ctx context.Context, mappi
entityName = foundUserDN.NormDN
}
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
_, err := globalIAMSys.PolicyDBSet(ctx, entityName, mapping.Policy, userType, isGroup)
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerSTSAccHandler - replicates STS credential locally.
func (c *SiteReplicationSys) PeerSTSAccHandler(ctx context.Context, stsCred *madmin.SRSTSCredential, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if stsCred == nil {
return errSRInvalidRequest(errInvalidArgument)
}
@@ -1555,7 +1583,7 @@ func (c *SiteReplicationSys) PeerSTSAccHandler(ctx context.Context, stsCred *mad
// Set these credentials to IAM.
if _, err := globalIAMSys.SetTempUser(ctx, cred.AccessKey, cred, stsCred.ParentPolicyMapping); err != nil {
return fmt.Errorf("unable to save STS credential and/or parent policy mapping: %w", err)
return iamReplicationError(fmt.Errorf("unable to save STS credential and/or parent policy mapping: %w", err))
}
return nil
@@ -4193,7 +4221,8 @@ func (c *SiteReplicationSys) SiteReplicationMetaInfo(ctx context.Context, objAPI
}
info.UserInfoMap[k] = madmin.UserInfo{
Status: madmin.AccountStatus(v.Credentials.Status),
Status: madmin.AccountStatus(v.Credentials.Status),
UpdatedAt: v.UpdatedAt,
}
}
}
@@ -4549,9 +4578,25 @@ func (c *SiteReplicationSys) PeerStateEditReq(ctx context.Context, arg madmin.SR
const siteHealTimeInterval = 30 * time.Second
func (c *SiteReplicationSys) startHealRoutine(ctx context.Context, objAPI ObjectLayer) {
ctx, cancel := globalLeaderLock.GetLock(ctx)
defer cancel()
for ctx.Err() == nil {
var leadership LockContext
select {
case <-ctx.Done():
return
case leadership = <-globalLeaderLock.lockContext:
}
if leadership.Context().Err() != nil {
continue
}
leaderCtx, cancel := mergeContext(leadership.Context(), ctx)
c.healWithLeadership(leaderCtx, objAPI)
cancel()
// Quorum loss cancels a leadership lease, not the subsystem. Wait
// for a new lease so revocations can still reach offline sites.
}
}
func (c *SiteReplicationSys) healWithLeadership(ctx context.Context, objAPI ObjectLayer) {
healTimer := time.NewTimer(siteHealTimeInterval)
defer healTimer.Stop()
@@ -5170,13 +5215,16 @@ func (c *SiteReplicationSys) healBucketReplicationConfig(ctx context.Context, ob
}
func (c *SiteReplicationSys) healIAMSystem(ctx context.Context, objAPI ObjectLayer) error {
// A peer rejecting a deletion must not stop unrelated live updates from
// reaching healthy peers. Retain and report the error for the next retry.
deletionErr := c.healIAMDeletions(ctx)
info, err := c.siteReplicationStatus(ctx, objAPI, madmin.SRStatusOptions{
Users: true,
Policies: true,
Groups: true,
})
if err != nil {
return err
return errors.Join(deletionErr, err)
}
for policy := range info.PolicyStats {
c.healPolicies(ctx, objAPI, policy, info)
@@ -5194,7 +5242,7 @@ func (c *SiteReplicationSys) healIAMSystem(ctx context.Context, objAPI ObjectLay
c.healGroupPolicies(ctx, objAPI, group, info)
}
return nil
return deletionErr
}
// heal iam policies present on this site to peers, provided current cluster has the most recent update.
@@ -5429,8 +5477,10 @@ func (c *SiteReplicationSys) healUsers(ctx context.Context, objAPI ObjectLayer,
peerName := info.Sites[dID].Name
u, ok := globalIAMSys.GetUser(ctx, user)
if !ok {
// Disabled identities are valid replication sources. CheckKey returns
// their stored record even though authentication is denied.
u, _, err := globalIAMSys.CheckKey(ctx, user)
if err != nil || u.Credentials.AccessKey == "" || u.Credentials.IsExpired() {
continue
}
creds := u.Credentials
@@ -5629,7 +5679,10 @@ func isGroupDescEqual(g1, g2 madmin.GroupDesc) bool {
}
func isUserInfoEqual(u1, u2 madmin.UserInfo) bool {
if u1.PolicyName != u2.PolicyName ||
// Full-site summaries omit secrets and claims. Equal status alone cannot
// distinguish a recreated identity or an edited service-account policy.
if !u1.UpdatedAt.Equal(u2.UpdatedAt) ||
u1.PolicyName != u2.PolicyName ||
u1.Status != u2.Status ||
u1.SecretKey != u2.SecretKey {
return false
+4
View File
@@ -624,6 +624,10 @@ func (sts *stsAPIHandlers) AssumeRole(w http.ResponseWriter, r *http.Request) {
claims[expClaim] = UTCNow().Add(duration).Unix()
claims[parentClaim] = user.AccessKey
if err := setIAMParentRevocationClaim(ctx, globalIAMSys.store, user.AccessKey, claims); err != nil {
writeSTSErrorResponse(ctx, w, ErrSTSInternalError, err)
return
}
tokenRevokeType := r.Form.Get(stsRevokeTokenType)
if tokenRevokeType != "" {
@@ -0,0 +1,57 @@
# Conditional multipart completion across pools
## Defect and scope
An unfinished multipart upload can remain in one pool while another pool holds
the logical current object. Evaluating `If-Match` against the upload pool's
local copy can accept a stale ETag or reject the current ETag. The pools object
lock serializes writes, but a local read still does not identify the logical
current object.
This layout does not require rebalance. Commit
`83b2ad418b15ff0fa78175e2014d78d02046edd8` (upstream #21115) made `getPoolIdx`
choose an available pool even when `pinfo.Err == nil`: normal overwrites and
upload initiation can select different pools. This change predates the SILO
multi-pool consistency work. The deterministic regression fixtures place copies
and uploads directly in real erasure pools; they do not claim to run rebalance.
## Minimal correction
For multi-pool conditional completion, retain the existing object lock and use
`objectPoolInfos` to read the logical current object before completing the
upload. An unreadable pool is an error, not proof of absence. The first sorted
copy supplies the ETag and encryption metadata used by the existing callback.
A current delete marker is treated as an absent key. `If-Match` then fails for
an absent object; `If-None-Match: *` may proceed.
Use explicit read options with an empty `VersionID` and `NoAuditLog: true`.
The precondition concerns the logical current object, independently of an
internal completion's destination version. After a successful check, clear
the callback before entering the set layer, so it is evaluated only once.
No new lock, storage format, replica cleanup algorithm or distributed protocol
is introduced. Single-pool and unconditional completion retain their existing
paths.
## Availability and validation
If any pool cannot supply the required metadata, conditional completion fails,
even when GET/HEAD can still read a copy from another pool. The unreadable pool
might hold a newer object, a delete marker, or no copy at all; none of these
possibilities can be assumed. Retry after recovery. This behavior is recorded
in the unreleased changelog.
The regression suite covers both upload-pool directions, stale/current ETags,
`If-None-Match: *`, absent objects and delete markers, read-quorum errors,
explicit destination versions, tied modification times, callback counts,
upload preservation, signed HTTP error bodies and concurrent completions.
The ordinary HTTP routing control accepts every valid placement; deterministic
fixtures provide the cross-pool regression gate.
## Separate follow-up scope
The pool-placement change and conditional checks in `PutObject` and
`NewMultipartUpload` require separate assessment. This completion fix does not
repair those paths. In particular, PUT has live destination-version and
preserved-ETag semantics, so its repair must not copy this completion-specific
empty-VersionID rule without examining that contract. Parallelizing the shared
pool metadata reader is also outside this correctness fix.
@@ -0,0 +1,95 @@
# R4–R8 集成核验
## 结论与范围
2026-09-16,R5、R6、R8 的修复在已包含 R4、R7 的 main 基线上完成集成。
本地完整 `cmd`、`internal` 测试、相关 race 检查、仓库 verifiers、构建和
HTTP 超时进程探针均通过。环境中的真实 `claude-opus-5`(effort `max`)独立
阅读合并差异与调用链,结论为 **GO_WITH_NONBLOCKING_NOTES,零阻断项**。
本记录对应 main 合并前的代码核验。最终 PR 的 Linux CI、DCO 和合并结果以该
PR 的实际提交及检查为准;这里的本地结果不代表发布、部署或多站点生产验收。
## 提交对应关系
基线:`9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe`。
| 问题 | 原修复提交 | 集成提交 | 行为 |
| --- | --- | --- | --- |
| R4 | PR [#193](https://github.com/pgsty/silo/pull/193),已在基线 | `af2b1794d38d9e70e1d2c3ee692426e4b6cab4bd`(merge) | SSE-KMS 复制保留标签修订时间 |
| R7 | PR [#194](https://github.com/pgsty/silo/pull/194),已在基线 | `9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe`(merge) | 复制元数据恢复不再重新写入传输用 aws-chunked |
| R5 | `115fe8b12329d147adbaf817faa1737392ecbf9b` | `680eac66e40b0980bc20e70d7ad34185096e63f5` | 删除标签推进修订,接收端抵御乱序事件,重试及 ACK 保留新状态 |
| R6 | `cf381a7151ef25fc95ace5fedcd767fa19410de2` | `0c61128d23f05ce6b37e7ace713c3ffbfb68f4cb` | 旧形态 marker purge 正确分类,MRF 恢复 marker 并保留重试次数 |
| R6 核验记录 | `d38edb2c46182d3a8fa96e040604493d20a4b478` | `aea3882c95d16ec5598a07b40d593e04054137a9` | 保存 v3 共识及验证边界 |
| R8 | `0d48d32d7e038ae1ea5966f3d7e0cb86780a6311` | `055030ea53ca92ee22ce1e601ef4757c247edde8` | 配置绑定到读头绝对超时,正文继续采用滚动空闲超时 |
各原任务先取得 Opus 方案共识,再实施修复。原始方案、实现复核及验证记录保留在
[R5](../r5/verification.md)、[R6](../r6/README.md)、[R8](../r8/README.md)。
R5、R6、R8 原任务又分别只读核验了集成后的交叉影响,未发现新增生产阻断项。
集成使用 `git cherry-pick -x -s`,保留原作者、来源及 DCO。后续
`80684fed59f556d579e268c5a855d936c1347b68` 仅处理两类贡献规范问题:
- 六个新建测试文件统一使用实际贡献者姓名及 AGPL-3.0-or-later 头部;从
`package` 开始的内容逐字节不变,Linux build tag 保留。
- 按 CONTRIBUTING 的规则更新兼容标识清单。唯一新增条目是 R6 测试拼接既有
replication ARN 所用的 `arn:minio:replication::`,没有新增协议名称或生产行为。
22 个源码/测试文件的最终哈希见 [manifest.json](manifest.json)。共享文件中的
R5、R6 补丁与原修复具有相同稳定 patch ID;其余源文件直接比较,六个测试仅允许
上述头部差异。[等价检查](evidence/integration-equivalence.json)全部通过。
## Opus 集成复核与处置
实际 CLI 为 2.1.270,显式指定 `claude-opus-5 --effort max`;只允许 Read、Grep、
Glob,未执行测试或修改代码。实际返回模型为 `claude-opus-5`,进程和结果均成功。
复核基于 `055030ea53ca92ee22ce1e601ef4757c247edde8` 的 22 个文件及完整差异;
此后的代码变化仅为上文已证明等价的头部与兼容清单调整。
原文、调用元数据及提示词分别见 [复核结果](evidence/opus-review.md)、
[metadata](evidence/opus-integration.metadata.json)、[prompt](evidence/opus-integration.prompt.md)。
保留原文中的判断,再用直接证据逐项处置,避免把模型意见当作测试结果:
| 非阻断意见 | 核验与决定 |
| --- | --- |
| 非法或空的历史标签时间戳可能使复制失败并重试 | 保留 R5 共识中的失败关闭行为;历史异常数据修复另行处理 |
| 带标签修订的版本在 resync 时可能多一次 metadata COPY | R5 已接受的可靠性成本;正常 COMPLETED 路径保持原有门控 |
| purge 审计状态由 COMPLETE 规范为 COMPLETED,统计开始记录实际目标结果 | R6 的预期行为;后续发布说明应告知审计/指标使用者 |
| 配置的较短 ReadHeaderTimeout 同时缩短 TLS 握手窗口 | Go net/http 的预期语义,已在 R8 共识中说明 |
| 新增多池标签测试单独运行可能缺少全局初始化 | **未成立**:精确单独运行通过;`consistencyPools` 经 `prepareErasurePoolsWithContext` → `initObjectLayer` → `newTestObjectLayer` 调用 `initAllSubsystems`。保留测试原样 |
| 审计 fixture 重复取消可能输出栈信息 | 本地完整及 race 测试通过;不扩大本次生产修复范围 |
| 新测试文件头部应按实际贡献者整理 | 已在 `80684fed` 修正,测试代码及 build tag 不变 |
## 本地直接验证
下表全部针对 `80684fed59f556d579e268c5a855d936c1347b68`,未使用额外的源码或
容量 overlay;测试代码自身的容量 fixture 保留。限制并行度仅为
`GOMAXPROCS=4`、`GOFLAGS=-p=2`,并使用
仓库 CI 的 `MINIO_API_REQUESTS_MAX=10000`。详细命令、时间和日志哈希在
[validation-results.json](evidence/validation-results.json)。
| 检查 | 结果 |
| --- | --- |
| `make verifiers`:lint、生成文件、rebrand guard | 通过,76.9 秒;可选 typos 工具按现有 Makefile 规则跳过 |
| `make build`、`./silo --version` | 通过,产物为 silo |
| `CGO_ENABLED=0 go test -p 2 ./cmd ./internal/... -count=1 -timeout=30m` | 全部通过,340.2 秒,50 个有测试的包 |
| `CGO_ENABLED=1 go test -race`,cmd/deadlineconn/http 中变更测试的函数集合 | 通过,46.3 秒;Linux build tag 用例由最终 Linux CI 覆盖 |
| `TestAPIPoolsTaggingReplicaDeletion` 精确单独执行 | 通过,无需其他测试预先运行 |
| 实际 silo 进程的 CLI/环境变量读头超时探针 | 两种配置均在 100ms 读头限制下拒绝 400ms 才完成的请求头;空闲超时为 2s,随后健康请求成功 |
进程探针使用二进制 SHA-256
`1cc536f1a3c8d8372ff2d5b140b1fd2bc98a299324fea0f73f67288d48104ce4`。
原始输出见 [runtime-probe.json](evidence/runtime-probe.json)。
各子任务较早遇到的磁盘容量不足或筛选测试初始化问题,不作为这次通过的证据。
本地完整测试已重新执行并成功;原失败记录仍保留在各自调查档案。
## 验收边界
- 最终 PR 必须通过实际提交的全部仓库检查,尤其 Linux internal 测试、完整 cmd
测试、构建/vet、lint/生成文件、交叉编译、S3 Select race、DCO 和漏洞检查。
- 既有无时间戳对象、异常时间戳、标签筛选的目标选择、任意站点时钟偏差等不由本次
修复追溯重建。共享状态解析器的历史限制按 R6 v3 共识在写入点规避。
- R8 原先未完成的 S3 长传输脚本不计为通过;本次进程探针验证配置生效,不替代
S3 长传输、独立多进程、多节点或跨区域生产验收。
- 本次不引入依赖变更或上游 MinIO 兼容硬门槛;R9 不属于这五项修复。
@@ -0,0 +1,79 @@
ok github.com/minio/minio/cmd 320.458s
ok github.com/minio/minio/internal/amztime 1.605s
ok github.com/minio/minio/internal/arn 0.392s
ok github.com/minio/minio/internal/auth 0.440s
ok github.com/minio/minio/internal/bpool 0.422s
ok github.com/minio/minio/internal/bucket/bandwidth 0.450s
ok github.com/minio/minio/internal/bucket/cors 0.484s
ok github.com/minio/minio/internal/bucket/encryption 0.644s
ok github.com/minio/minio/internal/bucket/lifecycle 0.647s
ok github.com/minio/minio/internal/bucket/object/lock 0.653s
ok github.com/minio/minio/internal/bucket/replication 0.496s
ok github.com/minio/minio/internal/bucket/versioning 0.449s
ok github.com/minio/minio/internal/cachevalue 5.457s
? github.com/minio/minio/internal/color [no test files]
ok github.com/minio/minio/internal/config 0.656s
? github.com/minio/minio/internal/config/api [no test files]
? github.com/minio/minio/internal/config/batch [no test files]
? github.com/minio/minio/internal/config/browser [no test files]
? github.com/minio/minio/internal/config/callhome [no test files]
ok github.com/minio/minio/internal/config/compress 0.624s
ok github.com/minio/minio/internal/config/dns 0.731s
? github.com/minio/minio/internal/config/drive [no test files]
ok github.com/minio/minio/internal/config/etcd 1.015s
? github.com/minio/minio/internal/config/heal [no test files]
ok github.com/minio/minio/internal/config/identity/ldap 0.593s
ok github.com/minio/minio/internal/config/identity/openid 0.675s
? github.com/minio/minio/internal/config/identity/openid/provider [no test files]
? github.com/minio/minio/internal/config/identity/plugin [no test files]
? github.com/minio/minio/internal/config/identity/tls [no test files]
ok github.com/minio/minio/internal/config/ilm 1.057s
? github.com/minio/minio/internal/config/lambda [no test files]
ok github.com/minio/minio/internal/config/lambda/event 0.485s
? github.com/minio/minio/internal/config/lambda/target [no test files]
ok github.com/minio/minio/internal/config/notify 0.752s
? github.com/minio/minio/internal/config/policy/opa [no test files]
? github.com/minio/minio/internal/config/policy/plugin [no test files]
? github.com/minio/minio/internal/config/scanner [no test files]
ok github.com/minio/minio/internal/config/storageclass 0.585s
ok github.com/minio/minio/internal/config/subnet 0.582s
ok github.com/minio/minio/internal/crypto 0.843s
ok github.com/minio/minio/internal/deadlineconn 4.934s
ok github.com/minio/minio/internal/disk 0.418s
ok github.com/minio/minio/internal/dsync 131.144s
ok github.com/minio/minio/internal/etag 0.504s
ok github.com/minio/minio/internal/event 0.595s
ok github.com/minio/minio/internal/event/target 0.957s
ok github.com/minio/minio/internal/grid 7.616s
ok github.com/minio/minio/internal/handlers 0.717s
ok github.com/minio/minio/internal/hash 0.700s
? github.com/minio/minio/internal/hash/sha256 [no test files]
ok github.com/minio/minio/internal/http 13.643s
? github.com/minio/minio/internal/init [no test files]
ok github.com/minio/minio/internal/ioutil 1.921s
ok github.com/minio/minio/internal/jwt 0.423s
ok github.com/minio/minio/internal/kms 0.538s
ok github.com/minio/minio/internal/lock 1.077s
ok github.com/minio/minio/internal/logger 0.546s
? github.com/minio/minio/internal/logger/message/audit [no test files]
? github.com/minio/minio/internal/logger/target/console [no test files]
? github.com/minio/minio/internal/logger/target/http [no test files]
? github.com/minio/minio/internal/logger/target/kafka [no test files]
? github.com/minio/minio/internal/logger/target/loggertypes [no test files]
? github.com/minio/minio/internal/logger/target/testlogger [no test files]
ok github.com/minio/minio/internal/lsync 10.501s
? github.com/minio/minio/internal/mcontext [no test files]
? github.com/minio/minio/internal/mountinfo [no test files]
? github.com/minio/minio/internal/net [no test files]
? github.com/minio/minio/internal/once [no test files]
ok github.com/minio/minio/internal/pubsub 0.511s
ok github.com/minio/minio/internal/rest 0.525s
ok github.com/minio/minio/internal/ringbuffer 1.316s
ok github.com/minio/minio/internal/s3select 0.603s
ok github.com/minio/minio/internal/s3select/csv 0.446s
ok github.com/minio/minio/internal/s3select/json 0.439s
ok github.com/minio/minio/internal/s3select/jstream 0.417s
? github.com/minio/minio/internal/s3select/parquet [no test files]
? github.com/minio/minio/internal/s3select/simdj [no test files]
ok github.com/minio/minio/internal/s3select/sql 0.425s
ok github.com/minio/minio/internal/store 1.451s
@@ -0,0 +1,50 @@
[
{
"file": "cmd/replication-delete-mrf_test.go",
"before_sha256": "5e160f7e19cbbb8cd5fa4e7ffd9cff9e09361b3fc4c5ee5ae61458f99c777781",
"after_sha256": "6fcfaf505d28f98e42dff9c0d965895be2c2d112e9e996220d424aea1f76d691",
"body_sha256": "5b3cecec74273540f4a5f82ca55e8a28545d0252822af3568f7e36b19011bda8",
"body_identical": true,
"build_prefix_preserved": true
},
{
"file": "cmd/replication-delete-operation_test.go",
"before_sha256": "2888a04a543324776041316de2821f388d28c3c1a6f5d0031e5ac9d34f58d504",
"after_sha256": "2e674cab5ca4dbc38cb2c1ddcca117269276e6419e1869c154c3bd6a05676715",
"body_sha256": "7921af42a9f123450e3e567d0bf65cd030678404709c77fde4a9cbb366230c35",
"body_identical": true,
"build_prefix_preserved": true
},
{
"file": "cmd/server_deadline_config_test.go",
"before_sha256": "1013157f83baa5f7882ec2d41c7b1fccb9e05fb418d0fa61263953037c9698c4",
"after_sha256": "a8259d273922a8973b44d9a468791761d373fe265fff8b2b53871e4626b11405",
"body_sha256": "5714a9275aee7230d54ba8ed03e3705a3cf6e150bda122929a9131b4b2b8dd5c",
"body_identical": true,
"build_prefix_preserved": true
},
{
"file": "internal/deadlineconn/deadlineconn_strict_test.go",
"before_sha256": "f405690c9ff044595f48323d68f4a9b33ce695b3ad6820f54f17151067bfae5e",
"after_sha256": "b4aec28c5daddb36e6ebb98dd8af2a44b7a2c48f0660068534f69669609f67a7",
"body_sha256": "ab2b54e29ed34de6e59ba2fa4984d61ed5f93d74ae4f6cd22dd44ba77a37e569",
"body_identical": true,
"build_prefix_preserved": true
},
{
"file": "internal/http/dial_deadline_linux_test.go",
"before_sha256": "0939d05b72a09760d53fcdf249775989e3f89bca824b9961d0b2a657ebfdf41e",
"after_sha256": "c4c83bda92ba9bac53166453920456134b3d8265cd59101b4203067c67eca59e",
"body_sha256": "fac091ef08f29fe32c2668eccd8cd505106c9fe3c4715e4c3ed9063d4b71c86e",
"body_identical": true,
"build_prefix_preserved": true
},
{
"file": "internal/http/server_deadline_test.go",
"before_sha256": "a6e687b3904a876fa92a4c5b86453159f3e5a38a4b9412dc213c7772f47fbf0b",
"after_sha256": "87303cc389bf1cf759898489f06b001073de2dd9c4b3687c4eca14aa186b62ab",
"body_sha256": "20b238f34090c356b2254388e863d77d0a316f33b849291d15d94fad260f23dd",
"body_identical": true,
"build_prefix_preserved": true
}
]
@@ -0,0 +1,35 @@
{
"head": "80684fed59f556d579e268c5a855d936c1347b68",
"checks": {
"r6_shared_file_patch": true,
"r5_shared_file_patch": true,
"cmd/erasure-object.go": true,
"cmd/erasure-server-pool-consistency.go": true,
"cmd/erasure-server-pool.go": true,
"cmd/object-handlers-common.go": true,
"cmd/object-handlers.go": true,
"cmd/object-multipart-handlers.go": true,
"cmd/replication-tagging-order_test.go": true,
"cmd/replication-tagging-sender_test.go": true,
"cmd/bucket-replication-utils.go": true,
"cmd/replication-delete-marker_test.go": true,
"cmd/replication-delete-operation_test.go": true,
"cmd/replication-delete-mrf_test.go": true,
"cmd/common-main.go": true,
"cmd/server-main.go": true,
"internal/deadlineconn/deadlineconn.go": true,
"internal/http/listener.go": true,
"internal/http/server.go": true,
"cmd/server_deadline_config_test.go": true,
"internal/deadlineconn/deadlineconn_strict_test.go": true,
"internal/http/dial_deadline_linux_test.go": true,
"internal/http/server_deadline_test.go": true,
"only_reviewed_hygiene_changes": true,
"dco:80684fed59f556d579e268c5a855d936c1347b68": true,
"dco:055030ea53ca92ee22ce1e601ef4757c247edde8": true,
"dco:aea3882c95d16ec5598a07b40d593e04054137a9": true,
"dco:0c61128d23f05ce6b37e7ace713c3ffbfb68f4cb": true,
"dco:680eac66e40b0980bc20e70d7ad34185096e63f5": true
},
"all_pass": true
}
@@ -0,0 +1 @@
ok github.com/minio/minio/cmd 1.583s
@@ -0,0 +1,17 @@
{
"command": [
"go",
"test",
"-p",
"2",
"./cmd",
"-run",
"^TestAPIPoolsTaggingReplicaDeletion$",
"-count=1",
"-timeout=2m"
],
"exit_code": 0,
"seconds": 5.096,
"log": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/isolated-pools-before.log",
"log_sha256": "f014a1a72dda7b973cd1e0b7e77c70b470eeebca0d04805d98f07c2622c87b1e"
}
@@ -0,0 +1,2 @@
Checking dependencies
Building Silo binary to './silo'
@@ -0,0 +1,7 @@
Running lint check
0 issues.
typos binary is not found.. skipping..
compatibility manifest: imports=119 env=428 metrics=19 headers=87 routes=224 roots=1 grid=3 storage=16 policy=59 brand=181 sha256=ad05829578cf879b462a12fa65f3c10c8c7c6aa8d6c329e04779459eec2645be
Silo rebrand compatibility baseline is unchanged
Silo delivery and runtime rebrand checks passed
docker entrypoint argv compatibility tests passed
@@ -0,0 +1,66 @@
{
"candidate": "055030ea53ca92ee22ce1e601ef4757c247edde8",
"base": "9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe",
"requested_model": "claude-opus-5",
"requested_effort": "max",
"cli_version": "2.1.270",
"command": [
"/opt/homebrew/bin/claude",
"--print",
"--model",
"claude-opus-5",
"--effort",
"max",
"--safe-mode",
"--permission-mode",
"plan",
"--tools",
"Read,Grep,Glob",
"--strict-mcp-config",
"--no-session-persistence",
"--add-dir",
"/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab",
"--add-dir",
"/Users/vonng/pgsty/silo",
"--output-format",
"stream-json",
"--verbose"
],
"source_sha256": {
"cmd/bucket-replication-utils.go": "365641c760901641e8320cc5123697ab92a46f1add612048b991ed2d1cfad43b",
"cmd/bucket-replication.go": "2e766c5946dcabaea79455b50a8e426f404e55c92dcbccbe28d55906e2e43843",
"cmd/common-main.go": "f8777fe8a07d175aceee07b4dd13792b2384449c404c004c38a06893be997843",
"cmd/erasure-object.go": "4bc848685ea714d88cabbd5d1b8585fbcc06f7b19c775e1a811030e783d0e1a4",
"cmd/erasure-server-pool-consistency.go": "d2736ef6bffbb5c5758eba8df38f8d4ecb888a838ab0de8ad3cf015c051f8ad7",
"cmd/erasure-server-pool.go": "87ad0b25dfa3081d0e63d0073b788614a9c88e2498a2ce0956b93f8a0a03ef53",
"cmd/object-handlers-common.go": "101bd7d7447072d13fed50983b69b562e4725632645e623d7fdd490f388ecdec",
"cmd/object-handlers.go": "61897a260f3f5f660f41edcb50956c60e914ef98f9a987da824f16d78171fde2",
"cmd/object-multipart-handlers.go": "d9622c69c540ab32dd23916e3f534b6886473a98370c9dd17673e69a423b2a7e",
"cmd/replication-delete-marker_test.go": "d967787804d558ac6266b113228fdf4a4f9fcb7cab39138a4fb07558814ccca4",
"cmd/replication-delete-mrf_test.go": "5e160f7e19cbbb8cd5fa4e7ffd9cff9e09361b3fc4c5ee5ae61458f99c777781",
"cmd/replication-delete-operation_test.go": "2888a04a543324776041316de2821f388d28c3c1a6f5d0031e5ac9d34f58d504",
"cmd/replication-tagging-order_test.go": "c8260b4ccf82fa615e1e24b35a07f2d1aacbcf776e5c6f9dadffea4a09ad6ea8",
"cmd/replication-tagging-sender_test.go": "3770a1a48a6efe58fe8127e1e4fdf6bd7cf171e17db20f15222ea2f7b85db1af",
"cmd/server-main.go": "04c265de211412ba0297396928096d7f2d971244a957a3126154846035263514",
"cmd/server_deadline_config_test.go": "1013157f83baa5f7882ec2d41c7b1fccb9e05fb418d0fa61263953037c9698c4",
"internal/deadlineconn/deadlineconn.go": "b9272ef640f1d4403b3d0af6cdbaba9186c51ad9a0226dfe449e8ef738e1ec4b",
"internal/deadlineconn/deadlineconn_strict_test.go": "f405690c9ff044595f48323d68f4a9b33ce695b3ad6820f54f17151067bfae5e",
"internal/http/dial_deadline_linux_test.go": "0939d05b72a09760d53fcdf249775989e3f89bca824b9961d0b2a657ebfdf41e",
"internal/http/listener.go": "49628575367f6ab9b6986caf594726d74d370f7d2ac4eed582903600b6eb3fa2",
"internal/http/server.go": "b7b0355f2781f8c5f7c77bc910cd4180cd3e5f22a87de41bd35ef119d36b4cdf",
"internal/http/server_deadline_test.go": "a6e687b3904a876fa92a4c5b86453159f3e5a38a4b9412dc213c7772f47fbf0b"
},
"diff_sha256": "d8c4e60f9e4a5b1338f3e6e07b758b1bb279db80eec26847e3b35fde0d049485",
"prompt_sha256": "f95edcf71afedceab190ff57bfbec812b06c8c18b5370d887131a760228c2b2c",
"started_at": "2026-09-15T16:41:37.757693+00:00",
"status": "completed",
"exit_code": 0,
"actual_models": [
"claude-opus-5"
],
"finished_at": "2026-09-15T16:53:09.761932+00:00",
"raw_sha256": "3e1a45ebdf3f4420fb05647bb383c00d0a86452a6c357a92679485c99d8f4eb3",
"stderr_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"result_subtype": "success",
"is_error": false
}
@@ -0,0 +1,18 @@
Independently review the complete SILO R4-R8 integration candidate for a user-authorized merge to main. You are the real Claude Opus reviewer; provide your own conclusion from source inspection. Read-only: no edits, no GitHub actions, no test execution claims.
Exact candidate: 055030ea53ca92ee22ce1e601ef4757c247edde8; integration branch codex/merge-r4-r8. Base: 9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe, already contains separately reviewed and CI-accepted R4 (SSE-KMS tag timestamp) and R7 (replication metadata/aws-chunked). This candidate adds the final R5, R6, and R8 local repairs, cherry-picked without conflicts and with provenance/DCO preserved.
Read /Users/vonng/pgsty/silo/AGENTS.md and CONTRIBUTING.md. The maintained PGSTY stack is the release target; upstream MinIO compatibility is best effort. Read /Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/integration-code.diff and /Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/reviewed-source.json, then inspect complete relevant functions and tests in this worktree. Each repair already has real same-version Opus plan consensus: docs/investigations/r5/plan-v2.md and consensus.md; r6/plan-v3.md and consensus.md; r8/plan-v2.md and consensus.md. Prior implementation reviews and validation reports are supporting evidence, not substitutes for this integration review. The R6 v2 multi-target parser proof was disproved by a real storage counterexample and fixed only after v3 consensus; verify the final empty creation-update invariant.
Focus on concrete integration correctness:
1. R5 tag revision persistence and empty/nonempty ordering through COPY, PUT, multipart, retries and source ACK, with R6 purge/MRF state writes and shared bucket-replication.go functions.
2. R7 restores only six replication-specific fields. Verify tag values/timestamps remain handled correctly and aws-chunked is not reintroduced, including R4 KMS options.
3. R6 marker creation versus canonical/legacy purge, all exits, per-target statistics, disk creation/replica metadata preservation, identity-checked marker 405 recovery, retry counts and bounded scanner fallback.
4. R8 absolute request-header deadlines and CLI/env propagation, HTTP/1 streaming bodies, keep-alive/TLS/h2 boundaries, default DeadlineConn callers and effects on replication I/O.
5. Full-package test global state, cleanup/initialization ordering and any compile/dependency conflicts that separate scoped tests would miss.
Current local make verifiers is running; full candidate runtime tests and actual PR CI are still pending. Do not label them as passed. The host previously had ENOSPC; it now has more than 700 GiB free, so standard full-package tests can run without capacity overlays. Historical reports accurately retain earlier environment limits and incomplete multi-process deployment acceptance.
Two new R6 and four new R8 test files have boilerplate copyright headers queued for repository-policy normalization, after confirming original authorship. They will be header-only changes with package-and-body byte equivalence checked; do not conflate them with production semantic edits. Any actual implementation blocker will be fixed and re-reviewed as needed.
Return GO_WITH_NONBLOCKING_NOTES or REQUEST_CHANGES, number of blocking findings, exact candidate identity and inspected scope. For each blocker give path/line, real triggering conditions, impact and smallest compatible correction. Clearly distinguish inherited out-of-scope limitations from newly introduced defects and required CI from optional further deployment tests. Do not invent a blocker merely from absence of production deployment. This request is main merge, not release/deploy.
@@ -0,0 +1,49 @@
The plan file can't be written (Write is disabled this session), so here is the complete review inline.
# SILO R4–R8 integration review — independent, source-only
**Verdict: GO_WITH_NONBLOCKING_NOTES — 0 blocking findings.**
**Candidate identity:** head `055030ea53ca92ee22ce1e601ef4757c247edde8`, branch `codex/merge-r4-r8`, base `9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe`, worktree `/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/silo`. Scope of this review is a **main merge**, not a release or deployment. I executed nothing — no tests, no builds; every statement below is from reading source.
**Inspected scope:** `integration-code.diff` in full (3367 lines) and all 22 files in `reviewed-source.json`, then the complete surrounding functions in the worktree: `cmd/bucket-replication.go` (`replicateDelete`, `replicateDeleteToTarget`, `replicateObject`, `replicateAll`, `putReplicationOpts`, `getReplicationAction`, `queueMRFHeal`, `queueReplicaDeleteTask`, `queueMRFSave`, `queueReplicationHeal`), `bucket-replication-utils.go`, `bucket-replication-stats.go`, `xl-storage-format-v2.go DeleteVersion`, `erasure-object.go`, `erasure-multipart.go`, `erasure-server-pool.go`, `erasure-server-pool-consistency.go`, `object-handlers.go`, `object-handlers-common.go`, `object-multipart-handlers.go`, `handler-utils.go`, `object-api-options.go`, `bucket-handlers.go`, `common-main.go`, `server-main.go`, `internal/http/{server,listener}.go`, `internal/deadlineconn/deadlineconn.go`, all six new test files plus the base tests they interact with, `Makefile`, `.golangci.yml`, `AGENTS.md`, `CONTRIBUTING.md`, and the r5/r6/r8 plan + consensus records. Cross-checked against Go 1.27.1 `net/http/server.go` and pinned `minio-go/v7 v7.3.1-0.20260910142817-60bd07042d49`.
## Independent confirmation of the key claims
**R6 final empty creation-update invariant (the disproved-v2 point).** The v2 proof fails exactly as recorded: `replStatusRegex` (`bucket-replication-utils.go:168`) matches `arn1=;arn2=;` yielding `{arn1: ";arn2="}` — non-empty — while single-target `arn1=;` does not match at all. The v3 fix at `bucket-replication.go:575-582` sets `ReplicationStatusInternal=""`, `Targets=nil`, `ReplicaStatus=""`; `CompositeReplicationStatus` (`bucket-replication-utils.go:356-379`) then returns empty via both the internal string and the replica fallback, so `xlMetaV2.DeleteVersion` (`xl-storage-format-v2.go:1396-1405`, `1438-1447`) skips the creation/replica write while still writing `VersionPurgeStatusKey` (`1406-1408`, `1448-1450`). `ReplicationTimeStamp` is therefore inert (consensus N5 holds). I also confirmed the counterexample's precondition independently: `erasure-object.go:2099-2105` leaves `deleteMarker=true` when the stored marker carries no purge status, which is what sets `fi.Deleted=true` and reaches the rewriting branch. COMPLETE purges still remove the version (`1379-1393`, `1457-1459`).
**R6 classification / all exits.** `isVersionPurge()` (`1954-1956`) parses as `VersionID != "" || (DeleteMarkerVersionID != "" && !VersionPurgeStatus().Empty())`. Every live producer emits only one shape (`object-handlers.go:3232-3248`, `bucket-replication.go:3362-3378`, `3814-3836`), and the dir-object `nullVersionID` re-add (`bucket-handlers.go:552-555`, `679-681`) classifies as a purge under both old and new code, so no wire-form change there. All exits of `replicateDeleteToTarget` select the purge field consistently, HEAD probing is creation-only, `ReplicationDeleteMarker` is false for purges. `purgeReplicationStatus` maps only the legacy `COMPLETE` spelling; `ReplicationStats.Update` (`bucket-replication-stats.go:179-237`) consumes the passed status and never `rinfo.ReplicationStatus`, with `ri.Size == 0` giving count-only/zero-byte deltas. The `ResetStatusesMap` nil guard fixes a genuine nil-map assignment panic, since `ObjectToDelete.ReplicationState()` never initialises that map.
**R6 MRF recovery / budget.** `queueMRFHeal:4118-4125` parses as `(err != nil && !validMarker) || oi.Name == ""`; `decodeDirObject` is identity for both `obj` and `dir/`, matching `GetObjectInfo`'s decoded name, and `erasure-server-pool-consistency.go:143-145` is what returns a populated marker `ObjectInfo` with `MethodNotAllowed`. `RetryCount int` matches the persisted `MRFReplicateEntry.RetryCount` and `QueueReplicationHeal`'s parameter — no on-disk format change. All three increment sites feed the existing `> mrfRetryLimit` drop accounting (`3926-3931`), and the scanner fallback restarts with a fresh budget.
**R5 tag revision flow.** Sender: `replicationTaggingTimestamp` (`804-812`) serves both the full retransmit and the metadata-COPY branch, the latter now failing closed symmetrically (`1720-1725`). `getCopyObjMetadata:765-766` always emits `X-Amz-Tagging` (possibly empty) + `REPLACE`; minio-go `copyObjectDo` forwards empty header values; the receiver's `getRequestHeaderOrQueryValue` (`handler-utils.go:160-171`) treats presence-with-empty-value as authoritative — so an ordered deletion is genuinely representable on the wire. Receiver: `CopyObjectHandler:1799-1840` captures the stored stamp before REPLACE rebuilds the map, every branch writes or deletes the key explicitly, and the unconditional `delete(encMetadata, …)` is safe *because* of that, blocking the SSE-C rotation snapshot (`1655-1659`) from re-merging at `1910`. `PutObjectHandler:2323-2325` and `NewMultipartUploadHandler:315-318` mutate the same map that becomes `opts.UserDefined` (`object-api-options.go:451`), and the header is parsed only under trusted replication (`388-396`). Ordering is re-applied under the write lock by the existing `reconcileStoredObjectTags` callers (`erasure-object.go:136-139`, `1312-1315`; `erasure-multipart.go:1161-1189`; `erasure-server-pool.go:1443-1456`) — which is also what keeps the base R4 KMS test's `missing-timestamp` expectation intact. Dropping the `ri.UserTags` re-injection in the source ACK (`1294-1304`) is right: `cleanMetadata` strips the tagging key from `UserDefined`, so stored tags are now left alone, and the pools path re-derives them from merged `UserTags`.
**R7 boundary.** `replicationToInternalHeaders` has exactly six entries (`handler-utils.go:106-114`); `extractReplicationMetadataFromMime` restores only those and re-extracts no ordinary metadata, so the `aws-chunked` normalisation in `extractMetadata:225-241` is not undone. R5 touches neither, and R4's KMS options (`object-api-options.go:449-460`) still carry the three replication timestamps unmodified.
**R8 deadlines.** Against Go 1.27.1: header window set at `server.go:2038`/`2177`, whole-request deadline unconditionally at `1103`, `StateActive` at `2056-2058` firing after *every* successful `readRequest` because `readRequest` calls `setInfiniteReadLimit()` at `1067`. That is the one place where the naive reading of the `c.r.remain` comment is wrong — the pipelined/fully-buffered request does get the strict→rolling flip, so consensus N2 is correct. `startBackgroundRead` (`741`) and `hijackLocked` zero the deadline, which `infReads` honours — that is why background reads and hijacked grid/websocket conns still work. The strict cap only shortens, never extends; zero/past semantics unchanged; `readExplicit`/`readDeadlineStrict` only touched under `mu`. The h2 skip is defensive rather than load-bearing (net/http uses `skipHooks` for ALPN h2; the h2 server zeroes the conn deadline), and a nil `raw` fails the type assertion safely. Strict mode is opt-in, so every other `DeadlineConn` caller — the optional Linux internode dialer (`dial_linux.go:126-131`, currently disabled at `server-main.go:422`) and all outbound replication transports — keeps legacy rolling reads. The real fix is propagation: flag, field and `UseReadHeaderTimeout` already existed; `ctxt.ReadHeaderTimeout` was simply never populated before `common-main.go:448`.
**Full-package test state.** No duplicate symbols (`tagTestCapacityDisk` defined once in base `erasure-server-pool-tags_test.go:258`); no helper collisions in `internal/http`; `testdata/config/1.yaml`, `fmtGenFlags`, `serverCmd.Flags` all exist; `buildServerCtxt` mutates no globals. Globals are swapped/restored, and `prepareFS`/`prepareErasure`/`initAPIHandlerTest` re-run `initAllSubsystems` between backends, so leaked target-sys entries can't cross a fixture boundary. `logger.UpdateAuditWebhooks(ctx, nil)` really clears the list (`targets.go:227-273`), so the audit fixture is re-enterable across the SD and Erasure passes. `make verifiers` = lint + check-gen + rebrand-guard, and `.golangci.yml` enables no header linter.
## Blocking findings
**None.**
## Non-blocking notes (newly introduced here)
1. `bucket-replication.go:804-812` — an unparseable stored tagging timestamp now fails replication for *untagged* objects too. Trigger: a corrupt/legacy value, or the empty value `erasure-server-pool-consistency.go:211-214` can persist when a non-primary pool holds a malformed stamp. Fail-closed (FAILED + MRF retry), never silent divergence. Optional hardening: treat missing-or-empty as absent, keeping the hard error for non-empty malformed values.
2. `bucket-replication.go:1031-1036` — any version with a tag revision now resolves to `replicateMetadata` instead of `replicateNone`; since `object-handlers.go:1835-1836` gives every non-replica COPY a revision, large resyncs issue one extra metadata COPY per object. Steady state is unaffected (COMPLETED gate at `3810`). Accepted in r5 consensus; flagged as operational cost.
3. `bucket-replication.go:435-445`, `551-566`, `1958-1964` — purge audit status changes from legacy `COMPLETE` to `COMPLETED`, and per-target purge outcomes now reach `globalReplicationStats` where the empty creation status previously recorded nothing. Intended (r6 dispositions 1–2); worth one release-note line for audit/metric consumers.
4. `common-main.go:448` — now that `ReadHeaderTimeout` is populated it also participates in `net/http`'s `tlsHandshakeTimeout()` minimum. No change at defaults (30s vs 30s); a small configured header timeout also shortens the TLS handshake window. Documented in r8 consensus N1.
5. `replication-tagging-order_test.go:1805-1814` — `TestAPIPoolsTaggingReplicaDeletion` calls `globalBucketMetadataSys.Update` after `consistencyPools`, which (unlike `initAPIHandlerTest`) does not run `initAllSubsystems`; run in isolation with `-run` it would nil-dereference. Harmless in full-package CI. Smallest fix: use `initAPIHandlerTest`, as the neighbouring consistency tests do.
6. `replication-delete-mrf_test.go:1208-1215` — the audit cleanup cancels targets twice; non-panicking because `SafeClose(nil)` only prints a stack (`internal/ioutil/ioutil.go:463-471`), but it can add stderr noise.
7. Headers (already queued, header-only): the two new R6 test files use `Copyright (c) 2026 PGSTY` + SPDX-only, and the four new R8 test files carry `Copyright (c) 2015-2021 MinIO, Inc.` despite being new original work; CONTRIBUTING.md asks new original files to name their actual holders with AGPL boilerplate (the two new R5 files already comply). No verifier enforces headers, so this does not gate CI and is distinct from production semantics.
## Inherited / out-of-scope limitations (not introduced by this candidate)
Purge-target subset merging under narrowed fan-out; nil/missing target clients cannot restore already-lost tracking; the `ResetStatusesMap` key asymmetry between `targetResetHeader(arn)` writes and bare-`arn` reads (`bucket-replication-utils.go:392-399` vs `419-426`) — the candidate only adds the panic-preventing nil guard; the shared `replStatusRegex` still mis-parses serialized empty statuses (R6 deliberately works around it at the write site); unqualified (no `versionId`) replica PUT has no stored-tag reconcile (`object-handlers.go:2442`); TLS handshake *write* deadlines keep rolling behaviour (r8 N9); historical objects without a revision, tag-filter target selection and real multi-site clock skew remain unsolved (r5 consensus).
## Verification still owed (I ran nothing; nothing below is "passed")
**Required for this merge — repository gates only:** `make verifiers` (lint, check-gen, rebrand-guard), `make build` producing `silo`, and full-package tests, at minimum `./cmd`, `./internal/http`, `./internal/deadlineconn`; a Linux CI leg to compile and run the build-tagged `TestInternodeDialReadDeadline`; DCO sign-off/authorship checks on the cherry-picked commits. With >700 GiB free, no capacity overlay is needed.
**Optional, explicitly not a merge gate:** `SILO_TEST_LONG_UPLOAD=1` (>30s default-idle transfer regressions), multi-process/multi-node deployment acceptance, cross-region replication mesh. Their absence is a coverage boundary, not a defect; I do not treat missing production deployment as a blocker.
@@ -0,0 +1,27 @@
{
"binary": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/silo/silo",
"binary_sha256": "1cc536f1a3c8d8372ff2d5b140b1fd2bc98a299324fea0f73f67288d48104ce4",
"expected_rejection": true,
"cases": [
{
"source": "flag",
"header_timeout_ms": 100,
"idle_timeout_ms": 2000,
"header_completion_delay_ms": 400,
"rejected": true,
"status": "",
"elapsed_seconds": 0.403,
"still_alive": true
},
{
"source": "environment",
"header_timeout_ms": 100,
"idle_timeout_ms": 2000,
"header_completion_delay_ms": 400,
"rejected": true,
"status": "",
"elapsed_seconds": 0.402,
"still_alive": true
}
]
}
@@ -0,0 +1,12 @@
{
"command": [
"python3",
"docs/investigations/r8/evidence/runtime_probe.py",
"/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/silo/silo",
"fixed",
"/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/runtime-probe"
],
"exit_code": 0,
"seconds": 1.638,
"stdout_sha256": "e493d9757c130c2531072dc0eaee42b8fc1fa0f45fc43d0d0e098aac55f0d387"
}
@@ -0,0 +1,6 @@
silo version DEVELOPMENT.2026-09-15T16-44-34Z (commit-id=80684fed59f556d579e268c5a855d936c1347b68)
Runtime: go1.27.1 darwin/arm64
License: GNU AGPLv3 - https://www.gnu.org/licenses/agpl-3.0.html
Copyright: 2015-2025 MinIO, Inc.
Modifications: Copyright 2025-2026 PGSTY
Source compatibility: based on MinIO technology
@@ -0,0 +1,3 @@
ok github.com/minio/minio/cmd 13.654s
ok github.com/minio/minio/internal/deadlineconn 2.528s
ok github.com/minio/minio/internal/http 14.301s
@@ -0,0 +1,175 @@
{
"head": "80684fed59f556d579e268c5a855d936c1347b68",
"baseline": "9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe",
"source_sha256": {
"cmd/bucket-replication-utils.go": "365641c760901641e8320cc5123697ab92a46f1add612048b991ed2d1cfad43b",
"cmd/bucket-replication.go": "2e766c5946dcabaea79455b50a8e426f404e55c92dcbccbe28d55906e2e43843",
"cmd/common-main.go": "f8777fe8a07d175aceee07b4dd13792b2384449c404c004c38a06893be997843",
"cmd/erasure-object.go": "4bc848685ea714d88cabbd5d1b8585fbcc06f7b19c775e1a811030e783d0e1a4",
"cmd/erasure-server-pool-consistency.go": "d2736ef6bffbb5c5758eba8df38f8d4ecb888a838ab0de8ad3cf015c051f8ad7",
"cmd/erasure-server-pool.go": "87ad0b25dfa3081d0e63d0073b788614a9c88e2498a2ce0956b93f8a0a03ef53",
"cmd/object-handlers-common.go": "101bd7d7447072d13fed50983b69b562e4725632645e623d7fdd490f388ecdec",
"cmd/object-handlers.go": "61897a260f3f5f660f41edcb50956c60e914ef98f9a987da824f16d78171fde2",
"cmd/object-multipart-handlers.go": "d9622c69c540ab32dd23916e3f534b6886473a98370c9dd17673e69a423b2a7e",
"cmd/replication-delete-marker_test.go": "d967787804d558ac6266b113228fdf4a4f9fcb7cab39138a4fb07558814ccca4",
"cmd/replication-delete-mrf_test.go": "6fcfaf505d28f98e42dff9c0d965895be2c2d112e9e996220d424aea1f76d691",
"cmd/replication-delete-operation_test.go": "2e674cab5ca4dbc38cb2c1ddcca117269276e6419e1869c154c3bd6a05676715",
"cmd/replication-tagging-order_test.go": "c8260b4ccf82fa615e1e24b35a07f2d1aacbcf776e5c6f9dadffea4a09ad6ea8",
"cmd/replication-tagging-sender_test.go": "3770a1a48a6efe58fe8127e1e4fdf6bd7cf171e17db20f15222ea2f7b85db1af",
"cmd/server-main.go": "04c265de211412ba0297396928096d7f2d971244a957a3126154846035263514",
"cmd/server_deadline_config_test.go": "a8259d273922a8973b44d9a468791761d373fe265fff8b2b53871e4626b11405",
"internal/deadlineconn/deadlineconn.go": "b9272ef640f1d4403b3d0af6cdbaba9186c51ad9a0226dfe449e8ef738e1ec4b",
"internal/deadlineconn/deadlineconn_strict_test.go": "b4aec28c5daddb36e6ebb98dd8af2a44b7a2c48f0660068534f69669609f67a7",
"internal/http/dial_deadline_linux_test.go": "c4c83bda92ba9bac53166453920456134b3d8265cd59101b4203067c67eca59e",
"internal/http/listener.go": "49628575367f6ab9b6986caf594726d74d370f7d2ac4eed582903600b6eb3fa2",
"internal/http/server.go": "b7b0355f2781f8c5f7c77bc910cd4180cd3e5f22a87de41bd35ef119d36b4cdf",
"internal/http/server_deadline_test.go": "87303cc389bf1cf759898489f06b001073de2dd9c4b3687c4eca14aa186b62ab"
},
"no_capacity_overlay": true,
"race_test_selection": [
"TestAPILocalTaggingAlwaysAdvancesRevision",
"TestAPIPoolsTaggingReplicaDeletion",
"TestAPITaggingMultipartCommitRechecksRevision",
"TestAPITaggingReplicationOrdering",
"TestAPITaggingReplicationOrderingKMS",
"TestAPITaggingSSECRotationPreservesDeletionRevision",
"TestAPITaggingUnqualifiedCopyOrdering",
"TestConcurrentStrictReadDeadline",
"TestDefaultReadDeadlineStillRenews",
"TestInternodeDialReadDeadline",
"TestLocalTaggingCommitCannotRegressRevision",
"TestReplicateDeleteMarkerPurge",
"TestReplicateDeleteMarkerTargetSemantics",
"TestReplicateDeleteOperationExits",
"TestReplicateDeletePurgeMissingTargetState",
"TestReplicationDeleteQueueFullRetryBudget",
"TestReplicationMRFMarkerRecovery",
"TestServerBackgroundReadNoDeadline",
"TestServerConnStateHook",
"TestServerContinuousDownload",
"TestServerContinuousUpload",
"TestServerDefaultIdleLongDownload",
"TestServerDefaultIdleLongUpload",
"TestServerEarlyBodyClose",
"TestServerHTTP2Deadlines",
"TestServerHijackedDeadline",
"TestServerIdleBodyDeadline",
"TestServerKeepAliveDeadline",
"TestServerPipelinedDeadline",
"TestServerReadHeaderDeadline",
"TestServerReadHeaderTimeoutConfig",
"TestServerTLSHandshakeReadDeadline",
"TestStrictExpiredFutureReadDeadline",
"TestStrictReadDeadline",
"TestStrictReadDeadlineRepeatedRenewal",
"TestTaggingProductionCopyWireShape",
"TestTaggingRepeatedValueNeedsRevisionDelivery",
"TestTaggingReplicaContentDuplicateGuard",
"TestTaggingReplicationSenderRetryAndAcknowledgment",
"TestTaggingTimestampWire"
],
"started_at": "2026-09-15T16:45:09.640646+00:00",
"checks": [
{
"name": "make-verifiers-final",
"command": [
"make",
"verifiers",
"GOLANGCI=/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/golangci-serial"
],
"env": {
"GOMAXPROCS": "4",
"GOFLAGS": "-p=2",
"MINIO_API_REQUESTS_MAX": "10000"
},
"exit_code": 0,
"seconds": 76.933,
"log": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/make-verifiers-final.log",
"log_sha256": "54c906cff1d33d0148fbc4cac918c0785a3bcc8f7a4eb7e04a2f774c2f010bb4"
},
{
"name": "make-build",
"command": [
"make",
"build"
],
"env": {
"GOMAXPROCS": "4",
"GOFLAGS": "-p=2",
"MINIO_API_REQUESTS_MAX": "10000"
},
"exit_code": 0,
"seconds": 19.366,
"log": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/make-build.log",
"log_sha256": "6ba9b545236be964861749c72e7609edf12b8f470df30d1ede8fd62f497e629b"
},
{
"name": "silo-version",
"command": [
"./silo",
"--version"
],
"env": {
"GOMAXPROCS": "4",
"GOFLAGS": "-p=2",
"MINIO_API_REQUESTS_MAX": "10000"
},
"exit_code": 0,
"seconds": 1.259,
"log": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/silo-version.log",
"log_sha256": "ada27f2be570c33df5712e86782a7be2ce3acfef54c8bb22a9230db3606513af"
},
{
"name": "full-tests",
"command": [
"go",
"test",
"-p",
"2",
"./cmd",
"./internal/...",
"-count=1",
"-timeout=30m"
],
"env": {
"GOMAXPROCS": "4",
"GOFLAGS": "-p=2",
"MINIO_API_REQUESTS_MAX": "10000",
"CGO_ENABLED": "0"
},
"exit_code": 0,
"seconds": 340.193,
"log": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/full-tests.log",
"log_sha256": "24b986e2f5aef0e668116abaab5024bf19ac664aec010ca7aa5ef56886f47ad1"
},
{
"name": "targeted-race",
"command": [
"go",
"test",
"-race",
"-p",
"2",
"./cmd",
"./internal/deadlineconn",
"./internal/http",
"-run",
"^(TestAPILocalTaggingAlwaysAdvancesRevision|TestAPIPoolsTaggingReplicaDeletion|TestAPITaggingMultipartCommitRechecksRevision|TestAPITaggingReplicationOrdering|TestAPITaggingReplicationOrderingKMS|TestAPITaggingSSECRotationPreservesDeletionRevision|TestAPITaggingUnqualifiedCopyOrdering|TestConcurrentStrictReadDeadline|TestDefaultReadDeadlineStillRenews|TestInternodeDialReadDeadline|TestLocalTaggingCommitCannotRegressRevision|TestReplicateDeleteMarkerPurge|TestReplicateDeleteMarkerTargetSemantics|TestReplicateDeleteOperationExits|TestReplicateDeletePurgeMissingTargetState|TestReplicationDeleteQueueFullRetryBudget|TestReplicationMRFMarkerRecovery|TestServerBackgroundReadNoDeadline|TestServerConnStateHook|TestServerContinuousDownload|TestServerContinuousUpload|TestServerDefaultIdleLongDownload|TestServerDefaultIdleLongUpload|TestServerEarlyBodyClose|TestServerHTTP2Deadlines|TestServerHijackedDeadline|TestServerIdleBodyDeadline|TestServerKeepAliveDeadline|TestServerPipelinedDeadline|TestServerReadHeaderDeadline|TestServerReadHeaderTimeoutConfig|TestServerTLSHandshakeReadDeadline|TestStrictExpiredFutureReadDeadline|TestStrictReadDeadline|TestStrictReadDeadlineRepeatedRenewal|TestTaggingProductionCopyWireShape|TestTaggingRepeatedValueNeedsRevisionDelivery|TestTaggingReplicaContentDuplicateGuard|TestTaggingReplicationSenderRetryAndAcknowledgment|TestTaggingTimestampWire)$",
"-count=1",
"-timeout=15m"
],
"env": {
"GOMAXPROCS": "4",
"GOFLAGS": "-p=2",
"MINIO_API_REQUESTS_MAX": "10000",
"CGO_ENABLED": "1"
},
"exit_code": 0,
"seconds": 46.286,
"log": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/targeted-race.log",
"log_sha256": "7d021e57e513ea81617df44a6253703622145bf5e72e6180de84c2bd3c7186d3"
}
],
"source_unchanged": true,
"finished_at": "2026-09-15T16:53:13.682876+00:00"
}
@@ -0,0 +1,59 @@
{
"baseline": "9f3037e941a49ab4cd8a0eed7c0f01083fbe4bbe",
"reviewed_code_commit": "055030ea53ca92ee22ce1e601ef4757c247edde8",
"tested_commit": "80684fed59f556d579e268c5a855d936c1347b68",
"integration_branch": "codex/merge-r4-r8",
"current_source_matches_tested_commit": true,
"production_or_test_body_changes_after_review": false,
"source_sha256": {
"cmd/bucket-replication-utils.go": "365641c760901641e8320cc5123697ab92a46f1add612048b991ed2d1cfad43b",
"cmd/bucket-replication.go": "2e766c5946dcabaea79455b50a8e426f404e55c92dcbccbe28d55906e2e43843",
"cmd/common-main.go": "f8777fe8a07d175aceee07b4dd13792b2384449c404c004c38a06893be997843",
"cmd/erasure-object.go": "4bc848685ea714d88cabbd5d1b8585fbcc06f7b19c775e1a811030e783d0e1a4",
"cmd/erasure-server-pool-consistency.go": "d2736ef6bffbb5c5758eba8df38f8d4ecb888a838ab0de8ad3cf015c051f8ad7",
"cmd/erasure-server-pool.go": "87ad0b25dfa3081d0e63d0073b788614a9c88e2498a2ce0956b93f8a0a03ef53",
"cmd/object-handlers-common.go": "101bd7d7447072d13fed50983b69b562e4725632645e623d7fdd490f388ecdec",
"cmd/object-handlers.go": "61897a260f3f5f660f41edcb50956c60e914ef98f9a987da824f16d78171fde2",
"cmd/object-multipart-handlers.go": "d9622c69c540ab32dd23916e3f534b6886473a98370c9dd17673e69a423b2a7e",
"cmd/replication-delete-marker_test.go": "d967787804d558ac6266b113228fdf4a4f9fcb7cab39138a4fb07558814ccca4",
"cmd/replication-delete-mrf_test.go": "6fcfaf505d28f98e42dff9c0d965895be2c2d112e9e996220d424aea1f76d691",
"cmd/replication-delete-operation_test.go": "2e674cab5ca4dbc38cb2c1ddcca117269276e6419e1869c154c3bd6a05676715",
"cmd/replication-tagging-order_test.go": "c8260b4ccf82fa615e1e24b35a07f2d1aacbcf776e5c6f9dadffea4a09ad6ea8",
"cmd/replication-tagging-sender_test.go": "3770a1a48a6efe58fe8127e1e4fdf6bd7cf171e17db20f15222ea2f7b85db1af",
"cmd/server-main.go": "04c265de211412ba0297396928096d7f2d971244a957a3126154846035263514",
"cmd/server_deadline_config_test.go": "a8259d273922a8973b44d9a468791761d373fe265fff8b2b53871e4626b11405",
"internal/deadlineconn/deadlineconn.go": "b9272ef640f1d4403b3d0af6cdbaba9186c51ad9a0226dfe449e8ef738e1ec4b",
"internal/deadlineconn/deadlineconn_strict_test.go": "b4aec28c5daddb36e6ebb98dd8af2a44b7a2c48f0660068534f69669609f67a7",
"internal/http/dial_deadline_linux_test.go": "c4c83bda92ba9bac53166453920456134b3d8265cd59101b4203067c67eca59e",
"internal/http/listener.go": "49628575367f6ab9b6986caf594726d74d370f7d2ac4eed582903600b6eb3fa2",
"internal/http/server.go": "b7b0355f2781f8c5f7c77bc910cd4180cd3e5f22a87de41bd35ef119d36b4cdf",
"internal/http/server_deadline_test.go": "87303cc389bf1cf759898489f06b001073de2dd9c4b3687c4eca14aa186b62ab"
},
"compatibility_inventory_sha256": "208c78a9e9fc98d6de7f1e0f03cfec97d8491df65ad4de3d88f04c36c555f819",
"evidence_sha256": {
"evidence/full-tests.log": "24b986e2f5aef0e668116abaab5024bf19ac664aec010ca7aa5ef56886f47ad1",
"evidence/header-equivalence.json": "0aee56aa978515579aa59215152b685f614cd5bda323fe26690a4c21f837f109",
"evidence/integration-equivalence.json": "473711fa8660504275242b70e80622277a32c73e0c7df1de6f5ba2d6bbfcc54e",
"evidence/isolated-pools-before.log": "f014a1a72dda7b973cd1e0b7e77c70b470eeebca0d04805d98f07c2622c87b1e",
"evidence/isolated-pools-before.result.json": "3a761d8a6fd1869a9a9d2b3d506ddec4fdaaad098da2b6a2f85252acf9561774",
"evidence/make-build.log": "6ba9b545236be964861749c72e7609edf12b8f470df30d1ede8fd62f497e629b",
"evidence/make-verifiers-final.log": "54c906cff1d33d0148fbc4cac918c0785a3bcc8f7a4eb7e04a2f774c2f010bb4",
"evidence/opus-integration.metadata.json": "c468e952e2364625ffe04a7221489185705857ad073f6fdb8dee3fe6aad9345c",
"evidence/opus-integration.prompt.md": "f95edcf71afedceab190ff57bfbec812b06c8c18b5370d887131a760228c2b2c",
"evidence/opus-review.md": "463f6b97e8d19929187242724e6bcfc2406217cd82702739716b65e19a8f2360",
"evidence/runtime-probe.json": "e493d9757c130c2531072dc0eaee42b8fc1fa0f45fc43d0d0e098aac55f0d387",
"evidence/runtime-probe.result.json": "d8703c201493cf865d758d53cbc6d92b71535345dafe3c2f27ca2a29ad53e353",
"evidence/silo-version.log": "ada27f2be570c33df5712e86782a7be2ce3acfef54c8bb22a9230db3606513af",
"evidence/targeted-race.log": "7d021e57e513ea81617df44a6253703622145bf5e72e6180de84c2bd3c7186d3",
"evidence/validation-results.json": "29182a752a8ffb75367e8cce51e1f2ef080e453a059248ae2c868aa66627aede"
},
"original_raw_review": {
"path": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/opus-integration.jsonl",
"sha256": "3e1a45ebdf3f4420fb05647bb383c00d0a86452a6c357a92679485c99d8f4eb3"
},
"original_review_diff": {
"path": "/Users/vonng/tmp/silo-r4-r8-main-20260916-01a0a5ab/integration-code.diff",
"sha256": "d8c4e60f9e4a5b1338f3e6e07b758b1bb279db80eec26847e3b35fde0d049485"
},
"scope": "Local integration validation and independent source review. Final PR checks and remote merge are recorded separately."
}
+33
View File
@@ -0,0 +1,33 @@
# R4 plan consensus and review disposition
## Agreed version
- Plan: [plan v1](plan-v1.md), SHA-256 `ad539f2071155de6955b583991684ed33c4bfe2e29660005840cdc97d7e1a754`. The frozen file remains unchanged.
- Baseline: `9ebe81c1b3611f9cc73e676b5b741c2be62c467a`.
- Actual reviewer: Claude Code 2.1.270, every assistant model in the review stream is `claude-opus-5`; explicit `--effort max`.
- Opus: **GO_WITH_NONBLOCKING_NOTES**, zero blockers; explicitly agrees that this exact plan can enter local implementation. [Unedited returned review](opus-v1-review.md), [machine-readable provenance](opus-v1.metadata.json).
- Codex: agrees that adding the already-parsed timestamp to the KMS literal fixes R4, and accepts the nonblocking dispositions below. **No blocking disagreement remains on plan v1.** No production source edits were made before this record was saved.
- The agreement permits the planned local implementation and tests; it is not implementation acceptance, a merge decision or production release approval.
## Item-by-item disposition
| Opus ID | Disposition |
|---|---|
| R4-01 | Accepted citation correction here, leaving the agreed hash frozen: `ReplicaLockReconcile` is at baseline `object-handlers.go:1847`; encryption merge is at `:1903`. |
| R4-02 | Accepted scope clarification: ErasureSD and Erasure16 are both single-pool local backends. KMS rewrites use PutObject under-lock reconciliation. Multi-pool and multi-site validation are optional and deferred to the wider integration gate. Test comments and the final report will identify this boundary. |
| R4-03 | Accepted wording clarification: source encryption alone does not request destination encryption. Source-only SSE-C copy headers do not prevent destination bucket/default auto-KMS from selecting KMS. The three destination trigger categories stay unchanged. |
| R4-04 | Accepted intent. The regression matrix uses identical expected mtime, ETag, trust and all three source timestamps across all encryption modes, giving field-by-field equivalence without constructing expected values through the production function. The temporary expanded baseline matrix fails only trusted valid KMS tag timestamps. |
| R4-05 | Accepted optional test within the existing scope: a signed KMS COPY with nonempty tags and no source tag timestamp must preserve the stored value/time. This adds evidence, not production behavior. |
| R4-06 | Registered as a separate unverified-impact finding: KMS construction also omits `ProxyHeaderSet`, `ProxyRequest`, `Speedtest` relative to `getDefaultOpts`. No R4 fix or correctness claim for those flags. Send the observation to the parent for separate triage; do not assign it to R5. |
| R4-07 | Accepted. Assertions target final disk state; the REPLACE handler rebuilds metadata, while final stored-tag rejection occurs under the storage write lock. HTTP 200 alone is not acceptance. |
| R4-08 | Resolved provenance uncertainty by Codex: SHA-256 recomputed before/after review, baseline identity and current GitHub main/PR query captured in `baseline-identity.txt`. History was inspected locally with `git blame` / `git show`. Opus's read-only tools did not independently recompute the hash or check GitHub; those facts remain attributed to the local commands. |
## Raw evidence
Directory: `/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/`.
- `review-prompt-v1.md`, `opus-review-v1.jsonl`, `opus-review-v1.stderr.log`, `opus-review-v1.exit`.
- `baseline-identity.txt`, `r4_repro_test.go`, `overlay.json`, `baseline-repro.log`.
- `options_repro_test.go`, `options-overlay.json`, `baseline-options.log`.
The stream includes an attempted Write to Claude's own plan file. Its tool was disabled; the reviewer returned the full result in text and did not edit production source. The successful result and actual assistant models are checked separately from rate-limit status and auxiliary-model usage.
@@ -0,0 +1,39 @@
已完成独立复核:读了三份改动文件在 HEAD 的实际内容、`putOptsFromHeaders` 全函数与两条返回路径、`CopyObjectHandler` 的默认加密与标签排序段、`reconcileStoredObjectTags` 的三处生产调用点,以及全部指定日志与证据文件。
# R4 实现复核结论
**Verdict: GO_WITH_NONBLOCKING_NOTES(0 阻断项)**
- 复核 HEAD:`dbcf8dec589deb5d91e17d295cb70997635f5b55`
- 代码/测试 diff SHA-256(按提供值记录):`c8cd6648f8ecea835ec74a038cdeaa82acaa3f36250395f97ead3260dc2fc0a5`
- 本会话无 shell,未重算该哈希;改为逐行比对 diff 与工作树三份文件,内容完全一致(`cmd/object-api-options.go`、`cmd/object-api-options-replication_test.go`、`cmd/object-copy-replication-tagging_test.go`)。
## 核验到的事实
- 生产改动确为一个字段 + 相邻注释:`cmd/object-api-options.go:459` 的 `ReplicationSourceTaggingTimestamp: taggingtimestmp`,变量来自 `:419-425` 已解析值,与非 KMS 路径 `:473` 对齐。未动解析、信任判定、KMS key/context、返回结构。
- 影响面封闭:全仓该字段唯一消费点是 `cmd/object-handlers.go:1820`(COPY 标签排序)。PUT/POST/multipart 虽同经 `putOptsFromReq`,但无消费者,故不可能回归——与 R4/R5 切分一致。
- 三条 KMS 触发路径真实可达:`object-handlers.go:1428-1433` 在 `copyDstOpts`(`:1454`)之前套用目的端默认;`bucket-sse-config.go:139-151` 在 `nil 配置 + AutoEncrypt` 与桶默认 KMS 两种情况下都写入 `aws:kms`,因此 explicit / auto / bucket 三种模式均进入 KMS 分支。
- 回归证明成立:`baseline-final.log` 用 `-overlay` 换回未修复 constructor,失败面精确为「trusted × 有效标签时间戳 × SSE-KMS / SSE-KMS-context」和 6 个 KMS COPY 子测试(`tags="key=old"`、`kms=true`、HTTP 200),`none`/`SSE-S3`/`SSE-C`/非 trusted 全通过。修复后 `focused.log:194-204` 全 PASS。
- 测试确实覆盖被要求的维度:信任边界(trusted=false 时 mtime/ETag/三时间戳全归零)、错误路径(trusted + 畸形值必须报错且错误串含头名)、SSE 序列化回环(KMS keyID/context 原样还原)、磁盘终态(每事件 `obj.GetObjectInfo` 读真实盘)、版本一致性、签名 GET 明文可读。全局 `GlobalKMS`/`globalAutoEncryption`/`set.getDisks` 均 defer 还原。
- `race` exit 0、`vet` 空输出、`golangci-lint` 0 issues,均记录了与 HEAD 一致的三文件哈希。
- 未发现 `verification.md` / `verification.json` / `consensus.md` 中与日志矛盾的陈述。(评审者版本/模型/effort 这类 provenance 声明不在我可验证范围,未作背书。)
## 发现清单
| ID | 内容 | 阻断 |
|---|---|---|
| IMPL-01 | 单字段修复正确且充分,位置、变量、注释与 `:473` 语义一致 | 否(确认项) |
| IMPL-02 | 基线失败/修复通过的判别力成立,对照组不误报 | 否(确认项) |
| IMPL-03 | KMS 字面量相对 `getDefaultOpts` 仍缺 `ProxyHeaderSet`/`ProxyRequest`/`Speedtest`(`object-api-options.go:40-44` vs `:449-460`)。R4 范围外,已登记为 R4-06 | 否,不设为新合并门槛 |
| IMPL-04 | `metadata-directive: REPLACE` 下 `getCpObjMetadataFromHeader`(`:1143-1156`)返回全新 map,故 `:1818` 的 `lastTaggingTimestamp` 为空、`:1822` 解析失败使 handler 侧比较恒「incoming 胜」;真正的 stale 拒绝发生在写锁内的 `reconcileStoredObjectTags`(`erasure-object.go:1312-1315`)。测试终态断言仍正确,文档 R4-07 已明示此分工 | 否(R5 上下文) |
| IMPL-05 | 测试卫生:`bucket-kms` 模式写入的 `bucketSSEConfig` 未还原,仅因它是最后一个 mode、且 `ExecObjectLayerAPITest` 每后端重建对象层并 `resetTestGlobals()` 才安全;后续若在其后追加 mode 会继承默认 KMS | 否 |
| IMPL-06 | `object-api-options-replication_test.go:35` 局部变量名 `context` 遮蔽标准包名(本文件未导入该包),纯观感 | 否 |
| IMPL-07 | `focused.log` exit 1 的唯一失败是既有 `TestAPICopyObjectReplicaRetentionRemovalUnderBucketKMS`(`replication-trust_test.go:1284`,"Storage reached its minimum free drive threshold"),属本机磁盘余量环境问题,非本次引入;容量 overlay 是测试专用、未提交。新增 COPY 测试自带 `tagTestCapacityDisk` 包装,不受该阈值影响 | 否 |
**没有发现阻断性正确性问题。** 生产语义、存储格式、API 与既有排序规则均未改变,无任何既有测试断言旧(缺陷)行为。
## 合并适配性
`dbcf8dec5` 直接位于实时 main `9ebe81c1b` 之上,可快进合并。按仓库 CI 通过为前提,本实现适合合入 main。
*(我未运行任何测试,也未查询 GitHub;以上仅基于源码阅读与所提供日志。)*
@@ -0,0 +1,33 @@
{
"baseline": "9ebe81c1b3611f9cc73e676b5b741c2be62c467a",
"reviewed_head": "dbcf8dec589deb5d91e17d295cb70997635f5b55",
"requested_model": "claude-opus-5",
"requested_effort": "max",
"cli_version": "2.1.270",
"diff_sha256": "c8cd6648f8ecea835ec74a038cdeaa82acaa3f36250395f97ead3260dc2fc0a5",
"started_at": "2026-09-15T15:59:41.004963+00:00",
"status": "completed",
"command": "/opt/homebrew/bin/claude --print --model claude-opus-5 --effort max --safe-mode --permission-mode plan --tools Read,Grep,Glob --strict-mcp-config --no-session-persistence --add-dir /Users/vonng/tmp/silo-r4-evidence-20260915-a9cb --output-format stream-json --verbose",
"completed_at": "2026-09-15T16:03:52.331640+00:00",
"assistant_models": [
"claude-opus-5"
],
"observed_model": "claude-opus-5",
"verdict": "GO_WITH_NONBLOCKING_NOTES",
"blocking_findings": 0,
"session_id": "c09bc4fb-85f1-4af8-a1de-c453398e4a20",
"duration_ms": 161469,
"subtype": "success",
"is_error": false,
"used_tools": {
"Read": 20,
"Glob": 3,
"Grep": 14,
"ExitPlanMode": 1
},
"raw_stream": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/merge-review-1/opus.jsonl",
"stream_sha256": "875f617454b27e15ab44d9b777de89643ee352e0ccb948faad570fc546261acc",
"review_sha256": "962ff88d2ffcb75cd692ec17017411d624de80dc120f8dedc675e0fe25009335",
"prompt_sha256": "e1341c721161229431b943ba18d89b740e94470803c099b9ae3d597fd50544a4",
"review_extraction": "The substantive review is an earlier assistant text block; result.result only repeats CLI plan-mode merge limitations. Full raw stream and all assistant text are retained."
}
@@ -0,0 +1,84 @@
{
"original_reviewed_head": "dbcf8dec589deb5d91e17d295cb70997635f5b55",
"dco_signed_equivalent_head": "03027727d1d1b97d8beb83ac55569ea9a83dab23",
"notice_equivalence": {
"cmd/object-api-options-replication_test.go": {
"before_sha256": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"after_sha256": "c21fc8889a079085d9a882499a1cbe868278a3517580651f3bed1102e2a6aef8",
"package_body_sha256": "096f143c0b0a068581f9bb892f35ded0d65b6b60ab711f043236d27fbf51ca33",
"body_unchanged": true
},
"cmd/object-copy-replication-tagging_test.go": {
"before_sha256": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3",
"after_sha256": "73f066ed7258d430f078ecc90e551ece878bd3d4672bc762094d434ff8fec23d",
"package_body_sha256": "6b5173db2ded2d54055073c3259be208a4d7c8eac0367687082877f1fd3bef15",
"body_unchanged": true
}
},
"checks": {
"verifiers": {
"command": [
"make",
"verifiers",
"GOLANGCI=/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/merge-review-1/golangci-serial"
],
"exit_code": 0,
"started_at": "2026-09-15T16:04:55.536605+00:00",
"finished_at": "2026-09-15T16:07:03.252820+00:00",
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2",
"GOFLAGS": "-p=2"
},
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "c21fc8889a079085d9a882499a1cbe868278a3517580651f3bed1102e2a6aef8",
"cmd/object-copy-replication-tagging_test.go": "73f066ed7258d430f078ecc90e551ece878bd3d4672bc762094d434ff8fec23d"
},
"log_sha256": "e42a5bb55f5c1ebfcf02cebebf6d82cf1ec5a2d74590cdf838deba16dd80bfdf"
},
"build": {
"command": [
"make",
"build"
],
"exit_code": 0,
"started_at": "2026-09-15T16:07:03.253715+00:00",
"finished_at": "2026-09-15T16:07:35.594062+00:00",
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2",
"GOFLAGS": "-p=2"
},
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "c21fc8889a079085d9a882499a1cbe868278a3517580651f3bed1102e2a6aef8",
"cmd/object-copy-replication-tagging_test.go": "73f066ed7258d430f078ecc90e551ece878bd3d4672bc762094d434ff8fec23d"
},
"log_sha256": "6ba9b545236be964861749c72e7609edf12b8f470df30d1ede8fd62f497e629b"
},
"binary-version": {
"command": [
"./silo",
"--version"
],
"exit_code": 0,
"started_at": "2026-09-15T16:07:35.594918+00:00",
"finished_at": "2026-09-15T16:07:37.616706+00:00",
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2",
"GOFLAGS": "-p=2"
},
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "c21fc8889a079085d9a882499a1cbe868278a3517580651f3bed1102e2a6aef8",
"cmd/object-copy-replication-tagging_test.go": "73f066ed7258d430f078ecc90e551ece878bd3d4672bc762094d434ff8fec23d"
},
"log_sha256": "36317d06b691593fe0d74f88d053a24485500c15fc2001e857f2fc6fa5ba752a"
}
},
"binary_version": "silo version DEVELOPMENT.2026-09-15T16-03-52Z (commit-id=03027727d1d1b97d8beb83ac55569ea9a83dab23)\nRuntime: go1.27.1 darwin/arm64\nLicense: GNU AGPLv3 - https://www.gnu.org/licenses/agpl-3.0.html\nCopyright: 2015-2025 MinIO, Inc.\nModifications: Copyright 2025-2026 PGSTY\nSource compatibility: based on MinIO technology\n",
"raw_evidence_directory": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/merge-review-1",
"all_function_and_test_bodies_identical_to_opus_reviewed_version": true
}
@@ -0,0 +1,36 @@
# R4 合并前复核
用户已明确追加授权:使用 Opus 5 max 核实最终实现,确认无误后合并 main。本轮授权取代此前只交付本地补丁的范围限制。
## 真实实现评审
- 独立新调用:Claude Code 2.1.270,`--model claude-opus-5 --effort max`。
- 复核代码提交:`dbcf8dec589deb5d91e17d295cb70997635f5b55`;当时实时 main 与 fetch 结果均为 `9ebe81c1b3611f9cc73e676b5b741c2be62c467a`。
- 实际 assistant 模型只有 `claude-opus-5`。结论 **GO_WITH_NONBLOCKING_NOTES,0 阻断项**,明确表示仓库 CI 通过后适合合入 main。
- [原始实现评审正文](implementation-review.md)、[实际模型与输出哈希](implementation-review.metadata.json) 已保存。
- 原始流、全部 assistant 正文与最终 result 位于 `/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/merge-review-1/`。实质评审出现在较早的 assistant 消息;最终 result 只重复 Claude 只读会话不能自行合并的工具限制,不是对修复结论的撤回。本任务由 Codex 按用户明确授权完成合并。
## 意见处置
| 条目 | 处置 |
|---|---|
| IMPL-01 / IMPL-02 | 确认单字段修复和基线失败/修复通过的测试判别力,无需追加修改。 |
| IMPL-03 | Proxy/Speedtest 选项遗漏已交父任务单独核验,维持范围外,不纳入 R4 合并。 |
| IMPL-04 | REPLACE 请求的旧标签拒绝由写锁内对账完成,测试断言真实落盘状态,已有文档准确说明。 |
| IMPL-05 | 当前测试固定以 bucket-kms 为最后一种模式,且每后端重新初始化;现有执行顺序安全。后续增添模式需同步隔离桶默认配置,本次保持已评审测试逻辑。 |
| IMPL-06 | 局部变量 context 命名建议为可选观感项,不改动已评审逻辑。 |
| IMPL-07 | 既有锁测试的磁盘余量限制及仅测试容量 overlay 已如实记录;新测试和 race 不使用生产代码 overlay。 |
## 提交规范调整
按 `CONTRIBUTING.md` 补齐提交作者对应的 DCO sign-off,并将两个新原创测试文件的文件头改为 `Copyright (c) 2026 Feng Ruohang`,保留 AGPL-3.0-or-later。原有生产文件的继承声明保持原样。
生产函数和测试的 `package cmd` 之后内容与 Opus 审查版本逐字节相同。`merge-review-1/notice-equivalence.json` 记录了旧/新文件哈希及不变的代码正文哈希。原 `verification.json` 保留当时原始验证记录,不覆盖历史哈希;本轮 PR 的 CI 对最终提交重新验证。
## 合并门槛
`make verifiers` 已通过:全仓 lint 为 0 issues,生成文件检查通过,rebrand 兼容性清单未变化,交付/运行时标识检查和 entrypoint 参数兼容性测试通过。首次执行曾遇到其他任务持有 golangci-lint 进程锁;使用工具自带 `--allow-serial-runners` 串行等待后完成全部检查。可选 typos 工具未安装,由仓库 Makefile 按既有规则跳过。
`make build` 通过,已生成本地 `silo` 并成功执行 `./silo --version`。最终三个源文件哈希与本轮校验记录一致,详情见 [本轮验证清单](merge-verification.json)。
接下来由 PR CI 验证最终候选,并在合并前再次核对 main 和精确 PR head。CI 与合并事实以 GitHub PR 状态和本机原始合并证据为准,评审意见不等同于合并或发布。
+46
View File
@@ -0,0 +1,46 @@
## R4 独立评审(Opus 5.0,只读复核)
**计划**:plan v1 — `docs/investigations/r4/plan-v1.md`
**哈希(按任务给定)**:`ad539f2071155de6955b583991684ed33c4bfe2e29660005840cdc97d7e1a754`
**Baseline**:`9ebe81c1b3611f9cc73e676b5b741c2be62c467a`
### 裁定:GO_WITH_NONBLOCKING_NOTES
无阻断项。**我同意这份 exact plan(上述哈希)可以进入本地实现。** 下列 R4-01…R4-08 均为非阻断;其中 R4-02/04/05 的测试建议是**可选**的,不自动并入必做范围。
(说明:本会话 Write 工具被禁用,评审仅以正文返回,未写入任何文件,也未改动任何源码。)
### 我实际核验到的关键事实(支撑"单字段补丁正确且充分")
1. **缺陷确认**:`cmd/object-api-options.go:449-460` 的 KMS 字面量带了 MTime/PreserveETag/ReplicationRequest + 两个 Object Lock 时间戳,独缺 tagging;默认路径 `:473` 有。补丁片段中的变量名 `taggingtimestmp` 与 `:419` 完全一致,可直接编译;gofmt 对齐由更长的两个 Lock 键决定,不会扰动他行。
2. **影响面封闭**:全仓 `ReplicationSourceTaggingTimestamp` 只在 `cmd/object-handlers.go:1820` 被读取(定义于 `object-api-interface.go:99`)。因此该字段对 PUT/分段路径天然无效果——既印证 R4/R5 的切分合理,也说明补丁不可能回归其他路径。
3. **充分性的关键点(我重点查证的风险)**:`encMetadata` 只有在 SSE-C 轮换分支 `object-handlers.go:1648-1659` 才批量快照全部保留键,而该分支与 KMS options 分支互斥(目的端是 SSE-C 时 `crypto.S3KMS.IsRequested` 为假)。故 `:1903` 的 `maps.Copy(srcInfo.UserDefined, encMetadata)` **不会**覆盖 KMS COPY 新写入的 tags/时间戳 —— 单字段补丁在 R4 边界内充分。
4. **三个触发点准确**:`bucket-sse-config.go:135-153`(显式请求优先 → nil 配置 + AutoEncrypt → KMS → bucket 默认 KMS 写 header+keyID;默认 AES 走 AES 分支),配合 `object-handlers.go:1428-1433` 仅在非联邦时套用目的端默认。
5. **REPLACE 副本路径准确**:`reconcileStoredObjectTags`(`erasure-server-pool-consistency.go:232-243`)语义即"存量有效时间戳胜过缺失/更旧/相等的 incoming,并连同 tag 值一起还原"。KMS 目的端因 `isTargetEncrypted` 使 `metadataOnly=false`,实际落到 `erasure-server-pool.go:1499-1513`(`ReplicaLockReconcile` 经 `:1509` 透传)→ `erasure-object.go:1276-1316`,在 `cloneMSS`(:1324) 之前于写锁内完成对账;纯元数据路径走 `erasure-object.go:136-139`。计划同时引用 `:136` 与 `:1509`,判断正确。
6. **证据可信**:`baseline-repro.log` 中 options 用例非 KMS 保留 `...123456789Z`、KMS 返回零值;COPY 用例 6/6(ErasureSD + Erasure16 × explicit/auto/bucket KMS)失败,且均为 200、`kms=true`、明文 GET 通过、tags 停在 `key=old`。即"请求成功、加密正常,但复制标签被静默丢弃",与计划表述一致,未夸大。
7. **修复后推演**:newer/stale/duplicate/newer-again 在 handler(:1817-1833) 与写锁对账的双重排序下分别得到 new/new/new/latest,与测试期望吻合;旧发送端不带 `X-Minio-Source-Tagging-Timestamp` 时仍为零值 → 行为不变,兼容性主张成立。
### 问题清单
| ID | 阻断 | 内容与建议 |
|---|---|---|
| **R4-01** | 否 | 行号漂移:计划写的 `1851/1910`,实际是 `object-handlers.go:1847`(`ReplicaLockReconcile`)与 `:1903`(encMetadata merge)。建议更正引用。 |
| **R4-02** | 否(建议可选) | `ExecObjectLayerAPITest` 两种后端均为**单 pool**(`test-utils_test.go:216` `mustGetPoolEndpoints(0, ...)`),故 `erasure-server-pool.go:1443` 多池分支未被覆盖;且 KMS 目的端命中的是 PutObject 重写对账而非 `CopyObject:136`。建议在计划或测试注释中点明"单盘/16 盘均为单池";补多池覆盖**可选**,不必进必做范围。 |
| **R4-03** | 否 | 措辞:`crypto.Requested`(`internal/crypto/sse.go:74`)只检查**目的端** SSE 头,因此仅带 SSE-C *copy-source* 头的请求在 KMS 默认桶/自动加密下仍会进入 KMS 分支(归入触发点 2/3,枚举仍完整)。建议澄清 "source encryption alone…" 一句。 |
| **R4-04** | 否(**可选**) | 建议在 options 矩阵里加一条 KMS 分支 vs 默认分支的**逐字段等价断言**(MTime/PreserveETag/ReplicationRequest/三个复制时间戳)。这是阻止第三次复发最廉价的护栏(2021 漏、2026 补了两个 Lock 时间戳仍漏此项)。计划第 1 条已基本覆盖,此为结构化建议。 |
| **R4-05** | 否(**可选**) | 建议加一例"KMS 目的端 + 有 tags 但无 tagging 时间戳头 → 存量不变",把兼容性主张钉在 handler 层而不仅在 options 层。 |
| **R4-06** | 否(范围外,仅登记) | 同一 KMS 字面量相对 `getDefaultOpts`(`object-api-options.go:40-44`) 还遗漏 `ProxyHeaderSet/ProxyRequest/Speedtest`;`opts.Speedtest` 在 `erasure-object.go:1625` 被读取,全局自动加密下 speedtest PUT 会丢该标志。**不要在 R4 修**,且当前也不在 R5 声明范围内,建议单列条目登记。 |
| **R4-07** | 否 | REPLACE 时 `getCpObjMetadataFromHeader:1143-1156` 会重建 map,`lastTaggingTimestamp` 为空 → handler 对 stale 事件**恒接受**,真正的拒绝来自写锁内对账。因此回归测试必须断言**最终落盘状态**(现有复现已如此),不要改为断言 handler 层行为。 |
| **R4-08** | 否(不确定性) | 本会话无 shell,无法独立复算计划 SHA-256、验证 `c4373ef290 / b2dca43fda / cfefc049c` 历史归属与 PR #184/#187。可由 `shasum -a 256 docs/investigations/r4/plan-v1.md` 与 `git log -L` 输出消解;均不影响补丁正确性。`cfefc049c` 的 KMS context 编码修复实体(`:436-444` 的 `sdkContext`)仍在,补丁不触碰。 |
### 对计划各主张的逐项裁定
- 单字段补丁**正确且对本 bounded issue 充分**:同意(依据 2/3/7)。
- 三个目的端 KMS 触发点**描述准确**:同意(R4-03 仅措辞澄清)。
- 当前 REPLACE 副本路径**描述准确**:同意(R4-01/02 属引用精度)。
- 回归矩阵与存量状态说明**充分**:同意;存量部分"不自动回填、丢失源时间不可重建、并列/更旧事件不保证修复"的表述与 `reconcileStoredObjectTags` 实际语义一致。
- 信任边界、错误行为、加密 key/context、tie 语义、兼容性:补丁均未触碰,维持不变。
### 交付
本轮为**计划共识**,非实现验收。我未作任何源码或文件修改;R4 可按 plan v1 在本地实施,实施后的差异与测试证据需另行验收。
@@ -0,0 +1,36 @@
{
"baseline": "9ebe81c1b3611f9cc73e676b5b741c2be62c467a",
"plan": "docs/investigations/r4/plan-v1.md",
"plan_sha256": "ad539f2071155de6955b583991684ed33c4bfe2e29660005840cdc97d7e1a754",
"requested_model": "claude-opus-5",
"requested_effort": "max",
"cli_version": "2.1.270",
"started_at": "2026-09-15T15:45:46.438630+00:00",
"status": "completed",
"raw_output": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/opus-review-v1.jsonl",
"raw_stderr": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/opus-review-v1.stderr.log",
"command": "/opt/homebrew/bin/claude --print --model claude-opus-5 --effort max --safe-mode --permission-mode plan --tools Read,Grep,Glob --strict-mcp-config --no-session-persistence --add-dir /Users/vonng/tmp/silo-r4-evidence-20260915-a9cb --output-format stream-json --verbose",
"completed_at": "2026-09-15T15:49:48.150062+00:00",
"assistant_models": [
"claude-opus-5"
],
"observed_model": "claude-opus-5",
"verdict": "GO_WITH_NONBLOCKING_NOTES",
"blocking_findings": 0,
"result_subtype": "success",
"is_error": false,
"session_id": "63e14a68-8565-41fa-9746-3e405fb63e9f",
"duration_ms": 161784,
"num_turns": 35,
"used_tools": {
"Read": 18,
"Glob": 4,
"Grep": 11,
"Write": 1
},
"stream_sha256": "eb0918d8a6185b180dddcfc664a96682f05502ecf3b686b08a0547f09879d57d",
"review_sha256": "e1dc12dd99326ae432623ff8de201813e6e84e7ed16a5556c21f9c514d663676",
"prompt_sha256": "07c225beff1523e056c154b3a387cf1ae345def4b4d0173065b882140ac5abdf",
"tool_scope_note": "Read/Grep/Glob allowed. Claude attempted Write to its own plan; the tool was disabled and no file was written. git diff before consensus showed no production source changes.",
"auxiliary_model_note": "assistant_models records actual reviewing assistant messages. Auxiliary usage is distinct. --effort max is explicit in the command, not inferred from model usage."
}
+65
View File
@@ -0,0 +1,65 @@
# R4 plan v1: preserve the replicated tag timestamp for SSE-KMS
## Baseline and ownership
- Baseline: `9ebe81c1b3611f9cc73e676b5b741c2be62c467a`, verified against GitHub main on 2026-09-15.
- Branch: `codex/r4-kms-tag-timestamp`; worktree: `/Users/vonng/.codex/worktrees/a9cb/silo`.
- Live open PRs at inspection: #184 and #187, neither owns this options change.
- The worktree lacks the ignored `AGENTS.md`; the parent explicitly confirms `/Users/vonng/pgsty/silo/AGENTS.md` applies. Maintain the PGSTY product graph and inexpensive compatibility.
- R4 owns only the missing field in `cmd/object-api-options.go` and its regression tests. R5 owns DELETE/empty tags, PUT/multipart receiving, sender propagation and full receiver ordering. R4 will supply a standalone source patch to R5; neither task edits the other's worktree.
## Proven defect and actual trigger
`putOptsFromHeaders` parses the trusted source tag timestamp before selecting encryption. The SSE-KMS branch constructs and returns another `ObjectOptions` carrying mtime, ETag, replication trust and both Object Lock timestamps, but omits `ReplicationSourceTaggingTimestamp`. The normal path retains it. The parser accepts and preserves RFC3339 fractional seconds even though its layout is `time.RFC3339`; the reproduction uses nanoseconds.
`CopyObjectHandler` applies local destination encryption configuration before `copyDstOpts` → `putOptsFromReq` → `putOpts` → `putOptsFromHeaders`. The omission is reached by:
1. Explicit destination SSE-KMS request headers (with or without a key ID/context).
2. A destination bucket with default SSE-KMS, when the request has no explicit SSE choice.
3. Global automatic encryption with no bucket SSE override and no explicit SSE choice.
Explicit AES256/SSE-C takes its existing branch; source encryption alone does not select the destination KMS branch. Remote federation skips local destination defaults. The relevant trigger is trusted metadata entering the destination KMS branch, not every SSE-KMS object or every tag operation.
At `CopyObjectHandler`'s tag decision, a zero source timestamp skips the tag update. Current under-lock reconciliation can preserve the stored tag/timestamp when metadata REPLACE reconstructs the map with no timestamp. In the observed same-version replica COPY, the request succeeds, destination encryption is valid, and the old tags/timestamp remain. A missing field in the options layer is not itself proof of a content-read failure.
PUT and multipart consumers' independent failure to persist a parsed tag timestamp remain R5's responsibility. R4 does not claim to fix all tag replication by correcting this constructor.
## Source and reproduction evidence
- `cmd/object-api-options.go`: trusted parsing at 383–426; KMS construction at 433–460; normal assignments at 469–475.
- `cmd/object-handlers.go`: destination default encryption at 1425–1435; `copyDstOpts` at 1454; tag timestamp consumption at 1807–1834; replica reconciliation enabled at 1851; encryption metadata merge at 1910.
- `internal/bucket/encryption/bucket-sse-config.go:135`: explicit request wins, absent config + auto encryption selects KMS, otherwise configured bucket algorithm/key ID applies.
- `cmd/erasure-server-pool-consistency.go:232`: stored valid timestamp wins over absent, older or equal incoming timestamp; writes preserve the stored tag value alongside its timestamp.
- `cmd/erasure-object.go:136` and `cmd/erasure-server-pool.go:1509`: same-version replica COPY reaches existing under-lock tag reconciliation, including object-data rewrites.
- History: the omission exists in `c4373ef290` (2021-09-18); `b2dca43fda` (2026-09-05) added the two Object Lock timestamps but not the tag timestamp. `cfefc049c` fixed KMS context encoding independently and must remain intact.
- Fresh temporary reproduction: `/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/r4_repro_test.go` and `baseline-repro.log` (overlay; no production edits).
- Command: `GOMAXPROCS=2 go test -p 2 -overlay /Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/overlay.json ./cmd -run '^TestReviewR4' -count=1 -timeout 5m -v`.
- Result: expected failure. Unencrypted and AES256 options preserve `2026-09-15T01:00:00.123456789Z`; KMS returns zero. Signed metadata REPLACE COPY on ErasureSD and Erasure (16 disks), across explicit/default/automatic KMS, returns 200 but retains `key=old` and the old timestamp for newer events. Actual encrypted metadata and plaintext GET roundtrips pass. Test deltas are 1–3 nanoseconds.
- These are in-process signed HTTP router and real local disk tests. `kms.NewStub` replaces the remote key service; the normal server encryption/decryption code still runs. Existing `tagTestCapacityDisk` avoids the host's free-space percentage threshold; it delegates all object data/metadata I/O to real test disks.
## Proposed production change
Add exactly this field to the existing KMS `ObjectOptions` literal:
```go
ReplicationSourceTaggingTimestamp: taggingtimestmp,
```
Update the neighboring explanatory comment to include tagging alongside retention/legal hold. Do not refactor the common return paths, change parsing/fallback/equal-timestamp semantics, change encryption context encoding, modify trust decisions, add SDK dependencies, or change storage/wire format. Those changes are unnecessary to restore the missing existing contract.
## Required validation after consensus
1. Add an options regression matrix covering unencrypted, SSE-S3, SSE-KMS with no context, SSE-KMS with a context, and SSE-C. Validate trusted/untrusted requests, missing/valid/malformed tag timestamps, nanosecond and timezone/whitespace handling, all three replication timestamps, mtime/ETag/trust, nonnil metadata, and unchanged SSE header serialization (including KMS key/context).
2. Promote the temporary COPY reproduction into a named, isolated regression test. Use actual signed same-version metadata COPY with REPLACE metadata and tagging directives, on single-disk and 16-disk backends. For explicit, bucket-default and automatic SSE-KMS, check newer update, older delivery, duplicate replay, and a second newer update. Verify stored tags, exact timestamp, version ID, encryption kind and plaintext GET after each operation. Include an unencrypted/SSE-S3 control if the fixture can do so without expanding implementation scope.
3. Fail the final regression tests against unmodified baseline using an overlay. Then run them on the fixed source, alongside existing replication-trust/options and bucket-KMS Object Lock tests. Check `gofmt`, `git diff --check`, and `go vet ./cmd`.
4. Run the new focused tests under `-race`. Use `GOMAXPROCS=2` and `-p 2` while sibling tasks share the host. A one-field pure option fix does not justify concurrent full-repository suites in all five tasks; full Linux CI and multi-site validation remain separate delivery gates.
5. If a test exposes a separate handler/storage defect, report evidence and coordinate with R5. Do not broaden R4's production patch to make unrelated tests pass.
## Compatibility, existing state, effort and delivery
- Public API, header names, stored key names, KMS context/key handling and supported dependencies remain unchanged. Untrusted source headers stay ignored; malformed trusted timestamps continue to fail; absent timestamp remains zero. Existing non-KMS behavior remains unchanged.
- No automatic rewrite/backfill. Lost source tag times cannot be reconstructed from the receiver alone. Upgrading permits subsequent properly timestamped events to be consumed. Review source-of-truth and target state before any targeted resync; full historical convergence also depends on R5. Repeated events subject to existing timestamp/tie semantics are not a universal repair guarantee.
- The source fix can land independently; complete deletion/empty-tag and mixed-encryption convergence needs R5 plus its integration evidence.
- Expected effort: approximately 0.5–1 engineer-day including reproduction, review and local validation; key-service deployment, multi-site failures and existing-state remediation are separate.
- After actual Opus 5.0/max agreement on this exact plan hash, implement locally without another user permission prompt. Preserve raw review, assistant model identity, request effort, baseline and plan hash, issue-by-issue disposition and explicit consensus before source edits.
- Deliver a reviewable local diff, tests and evidence. No main merge, remote publication/release, deployment or existing-state rewrite is authorized by this plan.
+18
View File
@@ -0,0 +1,18 @@
# R4 research log
## Verified baseline
2026-09-15: local clean HEAD and GitHub main both `9ebe81c1b3611f9cc73e676b5b741c2be62c467a`. Branch created as `codex/r4-kms-tag-timestamp`. Live GitHub open PRs #184 (`6addf9eb916b5a4b837480cf534cd1efa5407d3c`) and #187 (`b8f2fdde41dff3dc3b8db669c1d42d30ca5c1d3d`) concern other tasks. Claude Code reports `2.1.270`; Go reports `go1.27.1 darwin/arm64`.
## Coordination
- Parent task: `01a0a5ab-ee43-7911-bddd-1aca6f8afcc8`.
- R5: `01a0a5b9-602d-7470-9882-4817cf5fdcd1`, `/Users/vonng/.codex/worktrees/77ad/silo`.
- Parent and R5 acknowledged the ownership boundary: R4 options constructor and nonempty KMS COPY tests; R5 producer/receiver ordering and empty values. R5 will consume R4's minimal patch for combined KMS acceptance.
- Initial conservative expectation separated metadata COPY from REPLACE ordering. Inspection of current `ReplicaLockReconcile` and `reconcileStoredObjectTags` shows that stored timestamps are also reconciled under the write lock for REPLACE. The temporary reproduction therefore uses REPLACE directly; the fix must demonstrate the actual sender-shaped path without changing the handler.
## Baseline reproduction
Temporary overlay test source and output are in `/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/`. The options and HTTP/disk reproductions fail for the expected missing timestamp. Every KMS HTTP request completed with 200; newer tags remained old; encrypted object metadata and subsequent ordinary plaintext GET succeeded. The result is narrower than claiming all KMS replication fails, and stronger than merely comparing options.
The temporary source is not a production implementation. See [plan v1](plan-v1.md) for exact scope and required acceptance. The plan is frozen by SHA-256 before invoking real Opus.
+358
View File
@@ -0,0 +1,358 @@
{
"baseline": "9ebe81c1b3611f9cc73e676b5b741c2be62c467a",
"branch": "codex/r4-kms-tag-timestamp",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"source_files_match_all_test_runs": true,
"checks": [
{
"name": "baseline-final",
"command": [
"go",
"test",
"-p",
"2",
"-overlay",
"/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/baseline-final-overlay.json",
"./cmd",
"-run",
"^Test(PutOptsFromHeadersReplicationTimestamps|APICopyObjectReplicaTaggingTimestampUnderKMS)$",
"-count=1",
"-timeout=5m",
"-v"
],
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2"
},
"started_at": "2026-09-15T15:50:49.430860+00:00",
"finished_at": "2026-09-15T15:51:20.637630+00:00",
"exit_code": 1,
"log": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/baseline-final.log",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"log_sha256": "f849c2082235213764e3db7a314d52af75c4859102478a1e8f837afcaa8f1ea8",
"overlay_sha256": "f4cbc16e4ffccaf63191de2e8476162876055796adb4c38cda2d8c569a3bbabc",
"overlay_sources": {
"/Users/vonng/.codex/worktrees/a9cb/silo/cmd/object-api-options.go": {
"path": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/object-api-options.baseline.go",
"sha256": "16a560d0990ae929393f682f22b32ecd2e7f4d9390484b03b54e176fcd00cff5"
}
},
"assessment": "Expected baseline regression failure; KMS timestamp loss. Non-KMS controls pass."
},
{
"name": "focused",
"command": [
"go",
"test",
"-p",
"2",
"./cmd",
"-run",
"^Test(PutOptsFromHeadersReplicationTimestamps|APICopyObjectReplicaTaggingTimestampUnderKMS|ReplicationTrustControlsInternalOptionsAndEvents|GetAndValidateAttributesOpts.*|APICopyObjectReplicaRetentionRemovalUnderBucketKMS)$",
"-count=1",
"-timeout=5m",
"-v"
],
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2"
},
"started_at": "2026-09-15T15:51:20.638633+00:00",
"finished_at": "2026-09-15T15:51:48.382212+00:00",
"exit_code": 1,
"log": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/focused.log",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"log_sha256": "9a64cbc46e9fd6186b1d3031034b720857851ce70b1466b1da5da6d1a52b76ba",
"assessment": "New tests and options/trust pass; pre-existing KMS lock fixture blocked by host disk free-space percentage."
},
{
"name": "focused-capacity-adapted",
"command": [
"go",
"test",
"-p",
"2",
"-overlay",
"/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/capacity-overlay.json",
"./cmd",
"-run",
"^Test(PutOptsFromHeadersReplicationTimestamps|APICopyObjectReplicaTaggingTimestampUnderKMS|ReplicationTrustControlsInternalOptionsAndEvents|GetAndValidateAttributesOpts.*|APICopyObjectReplicaRetentionRemovalUnderBucketKMS)$",
"-count=1",
"-timeout=5m",
"-v"
],
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2"
},
"started_at": "2026-09-15T15:52:20.113475+00:00",
"finished_at": "2026-09-15T15:52:50.842773+00:00",
"exit_code": 0,
"log": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/focused-capacity-adapted.log",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"log_sha256": "13f7aa54564a433bef4dddcf2c5ad1fe46fa03869527fb902a255ce4c3263bc9",
"overlay_sha256": "76a7e6fb364bdaaa3469b1dc79f9ea318059f287a4888f322639920c76dfe53a",
"overlay_sources": {
"/Users/vonng/.codex/worktrees/a9cb/silo/cmd/replication-trust_test.go": {
"path": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/replication-trust-capacity_test.go",
"sha256": "c81b526ae51983881bbf464199e6b90f074fa695300fa8ff005e427e4d3c8208"
}
},
"assessment": "PASS"
},
{
"name": "race",
"command": [
"go",
"test",
"-p",
"2",
"-race",
"./cmd",
"-run",
"^Test(PutOptsFromHeadersReplicationTimestamps|APICopyObjectReplicaTaggingTimestampUnderKMS)$",
"-count=1",
"-timeout=5m",
"-v"
],
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2"
},
"started_at": "2026-09-15T15:52:50.843898+00:00",
"finished_at": "2026-09-15T15:53:43.789325+00:00",
"exit_code": 0,
"log": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/race.log",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"log_sha256": "ae66a8c9e569a5c4b57ae56e75afc85c06a8b76f1567187519da0726605a4de4",
"assessment": "PASS"
},
{
"name": "vet",
"command": [
"go",
"vet",
"-p",
"2",
"./cmd"
],
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2"
},
"started_at": "2026-09-15T15:53:43.790197+00:00",
"finished_at": "2026-09-15T15:53:51.129773+00:00",
"exit_code": 0,
"log": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/vet.log",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"log_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"assessment": "PASS"
},
{
"name": "lint",
"command": [
"/Users/vonng/pgsty/silo/.bin/golangci/v2.13.1/golangci-lint",
"run",
"--build-tags",
"kqueue",
"--timeout=10m",
"--config",
"./.golangci.yml",
"./cmd/..."
],
"cwd": "/Users/vonng/.codex/worktrees/a9cb/silo",
"env_override": {
"GOMAXPROCS": "2"
},
"started_at": "2026-09-15T15:53:51.130456+00:00",
"finished_at": "2026-09-15T15:55:39.114303+00:00",
"exit_code": 0,
"log": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/lint.log",
"source_sha256": {
"cmd/object-api-options.go": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"cmd/object-api-options-replication_test.go": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"cmd/object-copy-replication-tagging_test.go": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3"
},
"log_sha256": "e92606b0bf483111dff0a120c315ea165821348f31365020e2468a0059095c47",
"assessment": "PASS"
}
],
"format_checks": [
{
"command": [
"gofmt",
"-l",
"cmd/object-api-options.go",
"cmd/object-api-options-replication_test.go",
"cmd/object-copy-replication-tagging_test.go"
],
"exit_code": 0,
"output": ""
},
{
"command": [
"git",
"diff",
"--check"
],
"exit_code": 0,
"output": ""
}
],
"evidence_directory": "/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb",
"evidence_files": {
"baseline-final-overlay.json": {
"size": 177,
"sha256": "f4cbc16e4ffccaf63191de2e8476162876055796adb4c38cda2d8c569a3bbabc"
},
"baseline-final.json": {
"size": 1584,
"sha256": "5bf7be51e5a0eb41a40dc5fc5d3a1aba5df7733ad5dcb001f8d870a01c4233ba"
},
"baseline-final.log": {
"size": 22493,
"sha256": "f849c2082235213764e3db7a314d52af75c4859102478a1e8f837afcaa8f1ea8"
},
"baseline-identity.txt": {
"size": 1676,
"sha256": "9b21841e19a0cbb8ded18c2597488a527a27bedc65109d05d4ff598103073b68"
},
"baseline-options.log": {
"size": 11395,
"sha256": "d692f0a4bc9c58e2ac0087afa356ddf48f86e1040ea8138d68ffc4d0992bf3cb"
},
"baseline-repro.log": {
"size": 5693,
"sha256": "aa89b76f4723c6a3ce224faa7796403628d978a8707544bd97848b8887de2113"
},
"capacity-fixture.diff": {
"size": 870,
"sha256": "8d01e0b0068441f37ecee37125b81424d1f30d7c4fb37d435ea0cfe2e4617e5e"
},
"capacity-overlay.json": {
"size": 185,
"sha256": "76a7e6fb364bdaaa3469b1dc79f9ea318059f287a4888f322639920c76dfe53a"
},
"final_copy_repro_test.go": {
"size": 5942,
"sha256": "8949e07d96d2949a79f5a9e83c7a7c0473733d77b407e9c51da477d4ab74f1a8"
},
"focused-capacity-adapted.json": {
"size": 1661,
"sha256": "1a599b41caabfc5eb44db8d89c8b7e4f4f84f5f036d008369155fd337f65bdf9"
},
"focused-capacity-adapted.log": {
"size": 20384,
"sha256": "13f7aa54564a433bef4dddcf2c5ad1fe46fa03869527fb902a255ce4c3263bc9"
},
"focused.json": {
"size": 1253,
"sha256": "d9bb4979ea8aeaabb809cdc6e400a8673530bc83abf3dc2b2a06853a8523d0d9"
},
"focused.log": {
"size": 20475,
"sha256": "9a64cbc46e9fd6186b1d3031034b720857851ce70b1466b1da5da6d1a52b76ba"
},
"format-checks.json": {
"size": 352,
"sha256": "7569260900a799d5efdfb39db1f575ab1dadbbb04ace222e036968e66b6b59e7"
},
"lint.json": {
"size": 991,
"sha256": "a7144713b069f470a94b1ebe6fca6683a4b866a891a2756e15f28c280666ca14"
},
"lint.log": {
"size": 10,
"sha256": "e92606b0bf483111dff0a120c315ea165821348f31365020e2468a0059095c47"
},
"object-api-options.baseline.go": {
"size": 16653,
"sha256": "16a560d0990ae929393f682f22b32ecd2e7f4d9390484b03b54e176fcd00cff5"
},
"options-overlay.json": {
"size": 185,
"sha256": "5506b9c3b998b32f01c45af3cf01605eae9e4fb262c9ff3a6b0040abe719d4d9"
},
"options_repro_test.go": {
"size": 4192,
"sha256": "1a57a47bdd370042fa0f0d2d90efe447abedee9b9ef48a938d4bed631d83ec0b"
},
"opus-review-v1.exit": {
"size": 2,
"sha256": "9a271f2a916b0b6ee6cecb2426f0b3206ef074578be55d9bc94f6f3fe3ab86aa"
},
"opus-review-v1.jsonl": {
"size": 410466,
"sha256": "eb0918d8a6185b180dddcfc664a96682f05502ecf3b686b08a0547f09879d57d"
},
"opus-review-v1.stderr.log": {
"size": 0,
"sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
},
"overlay.json": {
"size": 158,
"sha256": "95f7c3f7e206fe36731e6d7e4a48f90c7403f07ae8155c125c6b86d2f1c2d487"
},
"r4-kms-tag-timestamp.patch": {
"size": 874,
"sha256": "2d4806d986bbd94ba4bc3951f3aeee48401ee1921c28ded0988fa09ca76ca26f"
},
"r4_repro_test.go": {
"size": 5508,
"sha256": "9bcefb6da2416b577b58085485cad60f677c2265e9dfa84d02e465e1b203766b"
},
"race.json": {
"size": 1026,
"sha256": "50b73e4acbc2426f3dcfadde78d0f0a86f10702345d2939a30204600bc750a13"
},
"race.log": {
"size": 18566,
"sha256": "ae66a8c9e569a5c4b57ae56e75afc85c06a8b76f1567187519da0726605a4de4"
},
"replication-trust-capacity_test.go": {
"size": 61271,
"sha256": "c81b526ae51983881bbf464199e6b90f074fa695300fa8ff005e427e4d3c8208"
},
"review-prompt-v1.md": {
"size": 2770,
"sha256": "07c225beff1523e056c154b3a387cf1ae345def4b4d0173065b882140ac5abdf"
},
"run-checks.py": {
"size": 2266,
"sha256": "ddecea5220bc9c286df18c0e9eca101f3d8ab731c307cd4acef937f3ecd11e65"
},
"vet.json": {
"size": 853,
"sha256": "e0964444bc91640ed6cf78229050bad5b94af4210bdf44b11d1a848f1ee930a6"
},
"vet.log": {
"size": 0,
"sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
}
},
"scope": "Darwin arm64; real signed HTTP + disk I/O; KMS service stub; single-pool single/16-disk fixtures; no remote CI, multi-site, release or deployment."
}
+52
View File
@@ -0,0 +1,52 @@
# R4 修复与本地验收
这是 2026-09-15 的本地验收快照。用户后续授权的实现级复核、提交规范调整与合并流程见 [合并前复核](merge-verification.md);以下原始测试记录及哈希保留当时状态。
## 结果
在 `putOptsFromHeaders` 的 SSE-KMS 选项构造中补齐 `ReplicationSourceTaggingTimestamp`。目的端使用显式 SSE-KMS、桶默认 KMS 或自动加密时,可信复制 COPY 现在能消费来源标签时间戳,并在现有存储锁内完成排序。
生产修改只有一个字段和相邻注释。API、存储格式、KMS key/context、信任判断和既有排序规则保持兼容。R5 的删除/空标签及 PUT/multipart 时间戳传播独立交付。
## 方案与 Opus 共识
- 基线:`9ebe81c1b3611f9cc73e676b5b741c2be62c467a`,已重新查询 GitHub main。
- 分支:`codex/r4-kms-tag-timestamp`。
- [冻结方案 v1](plan-v1.md):SHA-256 `ad539f2071155de6955b583991684ed33c4bfe2e29660005840cdc97d7e1a754`。
- 真实评审为本机 Claude Code 2.1.270,实际 assistant 模型 `claude-opus-5`,显式 `--effort max`。结论 **GO_WITH_NONBLOCKING_NOTES,0 个阻断项**。
- [逐条意见处置与双方共识](consensus.md)、[原始返回评审正文](opus-v1-review.md)、[模型与哈希记录](opus-v1.metadata.json) 已保存。先保存共识,再修改生产源码。
## 变更与测试
| 文件 | 内容 |
|---|---|
| `cmd/object-api-options.go` | 在 KMS 字面量中保留已解析的来源标签时间戳。 |
| `cmd/object-api-options-replication_test.go` | 无加密、SSE-S3、SSE-KMS、带 key/context 的 KMS、SSE-C;可信/非可信;缺失、有效、无效标签时间;纳秒、时区与空格;mtime/ETag/三个时间戳、metadata 与 SSE 序列化。 |
| `cmd/object-copy-replication-tagging_test.go` | 两种单池后端 × 五种目的端加密模式 × 五个有序事件,共 50 次签名 COPY 和 50 次普通 GET。每步检查最终标签、精确时间戳、对象版本、加密类型及明文内容。 |
COPY 使用 `metadata=REPLACE`、`tagging=REPLACE` 和可信复制身份。事件为较新更新、乱序旧更新、重复事件、再次更新,以及不带来源标签时间戳的请求。更新间隔仅 1–3 纳秒,防止时间精度退化被秒级测试掩盖。无加密与 AES256 是对照;KMS 覆盖显式、桶默认和自动加密入口。
## 验证状态
| 检查 | 结果 | 证据文件 |
|---|---|---|
| 最终测试 + 未修复基线 constructor overlay | 预期失败;只有可信 KMS 有效标签时间戳及 KMS COPY 更新失败,对照通过 | `baseline-final.log/json` |
| 修复后最终新增测试与既有 trust/options 测试 | 通过;未使用生产源码 overlay | `focused.log` |
| 既有 KMS Object Lock 回归 | 首次受宿主机磁盘余量阈值阻挡;仅适配测试容量报告后,与上述定向测试一起通过 | `focused-capacity-adapted.log/json`、`capacity-fixture.diff` |
| 新增测试 `-race` | 通过 | `race.log/json` |
| `go vet -p 2 ./cmd` | 通过 | `vet.log/json` |
| 仓库配置的 golangci-lint,范围 `./cmd/...`、`kqueue` build tag | 通过,0 issues | `lint.log/json` |
| gofmt、git diff --check | 通过 | `format-checks.json` |
原始日志和每条命令的运行记录位于 `/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/`。每份检查 JSON 都记录命令、退出码、时间和三个源码/测试文件的 SHA-256;最终交付已逐一确认文件哈希一致。[验证清单](verification.json) 另记录 overlay 的实际替换文件哈希,避免混淆基线与修复版执行代码。
测试使用真实签名 HTTP 路由、实际本地对象数据/元数据读写、服务器加解密代码;远程密钥服务由 `kms.NewStub` 代替。ErasureSD 与 16 盘 Erasure 均为单池。容量适配只使用已有 `tagTestCapacityDisk`,避免本机磁盘使用比例触发防写阈值,所有对象 I/O 仍由真实测试磁盘承担;未调整生产容量保护。
## 交付与剩余边界
- 本地实现和要求的定向验证均已完成,将源码、回归、研究、共识和验收记录作为一个本地提交交付。
- R4 的独立生产补丁已提供给 R5:`/Users/vonng/tmp/silo-r4-evidence-20260915-a9cb/r4-kms-tag-timestamp.patch`,SHA-256 `2d4806d986bbd94ba4bc3951f3aeee48401ee1921c28ded0988fa09ca76ca26f`。
- Opus 共识为方案级共识;本地测试结论来自实际运行,不把它记作 Opus 执行了测试。
- 多池/多站点故障恢复、外部 KMS 服务、完整 Linux CI、主干合并、远端发布和部署尚未执行。
- 没有改写存量。丢失的来源时间戳不能仅从接收端推导;后续重放/重同步须核对来源权威性及 R5 的全链路处理,不保证旧事件重放可以修复全部历史状态。
- 另登记 KMS 字面量缺少 Proxy/Speedtest 标志的范围外观察,已交父任务单独核验,本次未扩大修复。
+17
View File
@@ -0,0 +1,17 @@
# R5 investigation baseline
- Worktree: `/Users/vonng/.codex/worktrees/77ad/silo`
- HEAD: `9ebe81c1b3611f9cc73e676b5b741c2be62c467a`; GitHub main checked live on 2026-09-15.
- Branch: `codex/r5-tag-deletion-ordering`.
- WORKFLOW: `/Users/vonng/tmp/silo-r4-r8-20260915-01a0a5ab/WORKFLOW.md` read completely.
- This isolated worktree has no AGENTS.md. Read `/Users/vonng/pgsty/silo/AGENTS.md`: PGSTY supported stack, minimal compatible changes, separate local/merge/release gates.
- Existing tag storage reconciliation is already in HEAD; inspect and reuse it.
- No open R5 PR in live `gh pr list`; unrelated open PRs #184 and #187 belong to R6/R7.
- Parent reproduction: `/Users/vonng/tmp/silo-r4-r8-20260915-01a0a5ab/baseline-evidence/r5-handler.log`.
- Current reproduction overlay and raw output: `/Users/vonng/tmp/silo-r5-20260915-77ad/`.
- Claude Code actual version: 2.1.270 at `/opt/homebrew/bin/claude`. Required model `claude-opus-5`, effort `max`; model identity must be checked in assistant messages.
- Toolchain: go1.27.1 darwin/arm64. Targeted tests use GOMAXPROCS=2 and -p 1 to share the host.
## Ownership
R4 owns `cmd/object-api-options.go` KMS common-field preservation and option tests. R5 does not edit that file. R5 owns tag state generation, wire propagation, COPY/PUT/multipart persistence and ordered replay tests. Coordination requested through parent while R4 actual task ID is pending.
+32
View File
@@ -0,0 +1,32 @@
# R5 plan consensus
Date: 2026-09-15. Research base: `9ebe81c1b3611f9cc73e676b5b741c2be62c467a`.
## Accepted plan
- Version: **v2**, `plan-v2.md`.
- SHA256: `5a782acf3f285b23d1ae43a73481c4eb772a9a6d917fc5a550ecfc7cbf7446ca`.
- Actual reviewer: **claude-opus-5**, explicitly invoked **--effort max** through `/opt/homebrew/bin/claude` 2.1.270. All assistant messages in both reviews identify this model. The auxiliary Haiku usage in CLI bookkeeping is separately retained in modelUsage and is not the reviewer.
- Actual result: **APPROVE_WITH_NONBLOCKING_NOTES; 0 blocking items** in opus-v2-review.md.
- Codex accepts this exact v2 and its bounded per-hop scope. Plan hash was checked locally immediately before implementation. Opus's read-only tools did not run hashing; the original caveat is retained in raw review.
- Workflow permission: after this written consensus, local implementation and verification proceed without another user approval. No main merge, remote push, release, deployment or production state rewrite.
## Discussion and resolved differences
V1 was REQUEST_CHANGES with five blockers. See opus-v1-response.md for individual treatment and source evidence. V2 resolves all five. Opus explicitly withdrew its empty-only transfer proposal after the same-value re-addition counterexample, corrected its KMS COPY statement after inspecting bucket-default/auto encryption, and accepted that per-pool-only local clock guards are insufficient for ordinary source reads.
## Nonblocking notes accepted during implementation
- Extra metadata I/O occurs on scheduled metadata/heal/existing-object work with a recorded revision; ordinary object replication dispatches straight to full transfer. Completed scanner gates and incoming replication suppression avoid a feedback loop. Test the incoming no-reschedule decision.
- Pin unchanged object ModTime for local tagging changes.
- Keep a single-set monotonic guard and one uniform multi-pool candidate; direct-to-set writes outside the pool lock can transiently differ and re-converge on the next pooled mutation.
- Check actual failed COPY status/action and subsequent retry. Malformed timestamps fail both PUT and metadata COPY construction.
- Preserve scope limitations: tag-filter target selection, historical missing revisions, arbitrary unversioned content overwrites, and real multi-site/host-clock skew are not solved or production-accepted here.
## R4 dependency
Reuse reviewed local R4 commit `dbcf8dec589deb5d91e17d295cb70997635f5b55` on this isolated branch before implementation. Its only production change is the KMS options field, already examined against the provided patch SHA256 `2d4806d986bbd94ba4bc3951f3aeee48401ee1921c28ded0988fa09ca76ca26f`. R5 does not reimplement or modify that field. This makes R4+R5 tests run on actual combined source, with R5's eventual commit measured against the R4 dependency.
## Raw records
`/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v1.jsonl`, `opus-v1.stderr.log`, `opus-v1-request.json`, and matching `opus-v2.*`. In-repository review texts, prompts and metadata preserve plan hashes, model identity, usage and verdicts. V1 failure is not treated as approval.
@@ -0,0 +1,35 @@
{
"tested_dependency": "dbcf8dec589deb5d91e17d295cb70997635f5b55",
"merged_dependency": "af2b1794d38d9e70e1d2c3ee692426e4b6cab4bd",
"pr": "https://github.com/pgsty/silo/pull/193",
"files": {
"cmd/object-api-options.go": {
"tested_sha256": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"merged_sha256": "25e2e9484fafd94d1b2c857b94373758e481893ad93fb1a063edf7746277accc",
"package_body_sha256": "23fa2a25307bf1e41b665217a3f39aa7a8b860686a54209092fa9c0f1f13243a",
"package_body_identical": true
},
"cmd/object-api-options-replication_test.go": {
"tested_sha256": "1ea2a060987e32a4c76fce96ee974df475944c2d6ab482a4893e33daf7bca849",
"merged_sha256": "c21fc8889a079085d9a882499a1cbe868278a3517580651f3bed1102e2a6aef8",
"package_body_sha256": "1a57a47bdd370042fa0f0d2d90efe447abedee9b9ef48a938d4bed631d83ec0b",
"package_body_identical": true
},
"cmd/object-copy-replication-tagging_test.go": {
"tested_sha256": "5437a77e68736b4ce69de9c777675251fef24b0352dfe30bd8a836fc7ee810e3",
"merged_sha256": "73f066ed7258d430f078ecc90e551ece878bd3d4672bc762094d434ff8fec23d",
"package_body_sha256": "ef77de91cd91d1bd1c3cb10e4fee171c724c4f548749a23c7e48dcfcfb6a1585",
"package_body_identical": true
}
},
"other_changes": [
"cmd/object-api-options-replication_test.go",
"cmd/object-copy-replication-tagging_test.go",
"docs/investigations/r4/implementation-review.md",
"docs/investigations/r4/implementation-review.metadata.json",
"docs/investigations/r4/merge-verification.json",
"docs/investigations/r4/merge-verification.md",
"docs/investigations/r4/verification.md"
],
"scope": "R5 local branch will rebase onto this exact R4 merge; no R5 push or merge."
}
@@ -0,0 +1,340 @@
{
"raw_files": [
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-ack.log",
"bytes": 607,
"sha256": "ef698b8f46c850928057a31bf4e92f64f9a0dcebb8753bab00e9dcef2b44ffb6"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-extended.log",
"bytes": 2764,
"sha256": "a52bb6644b82a1986c233deeb9fb7b6cd3f4975337aa11446a1b27c459355db9"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-matrix.log",
"bytes": 20974,
"sha256": "5348c2ae4a20238ae50f70bcaea3aa55169b3479f60eab522692bdabe3420ab0"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-pools-rotation.log",
"bytes": 2249,
"sha256": "1d47acbe4d759b0f413f90589ff51b1f844f1885885d45d11c72a6295f5a4653"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-pools.log",
"bytes": 358,
"sha256": "56dd5efc0c833070576c4c7e2cb2abca8a82380060596be58871067526557c32"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/bucket-replication.go",
"bytes": 137081,
"sha256": "1e4d27c9eb2bff51eb28460d167faa3279b541d43c77ce35ad010dcab58bf7c5"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/erasure-object.go",
"bytes": 84089,
"sha256": "1012ae265e2453db125c6f2d16f866c72760b58d61c18f79dff4a551c208e3ce"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/erasure-server-pool-consistency.go",
"bytes": 14240,
"sha256": "78e63ca1117ea2d3e3e93864a4365dfa9bb303b4707de5087bdc144d0502a0db"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/erasure-server-pool.go",
"bytes": 99706,
"sha256": "2bfe0899fe3e42840fa4f078887184e4d7d2332d780ce63c31c9495af7056406"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/object-handlers-common.go",
"bytes": 19143,
"sha256": "00bf8409d25f9cf6a098a90f9e0bd7d2b237be7adc12622f0b9a66146af0a828"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/object-handlers.go",
"bytes": 146644,
"sha256": "27d47a17e12088f41b35de51da875f28a89a7e821bc27d0f88e6064ec386c993"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/cmd/object-multipart-handlers.go",
"bytes": 48935,
"sha256": "c818ce72d9115ed4f9cbe51571e7737957010a3934a4d000000895e45100edc7"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production/empty_test.go",
"bytes": 12,
"sha256": "9c78355c4da37df8f708f143fe19173dc146adcd99d1636594d265c5407755bf"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-production-overlay.json",
"bytes": 1508,
"sha256": "874d728abc9cb67c5db17ef4c3ce875d789ae4b5000dcbb593fe8dcd47094533"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-related-suite.log",
"bytes": 3521,
"sha256": "68b23b21c4de8b8252689f841376dec990257d929f0277d3087f467a246728e8"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline-resync.log",
"bytes": 116,
"sha256": "a620c0ecd112aceac9fd17b989b6283604a3866928ed51e9ed0b16079284fcbf"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline.log",
"bytes": 1629,
"sha256": "29d80ad52b0302d4eb4993db7a63c88bbc920923afa9775634d3c2fe000c066f"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/baseline_test.go",
"bytes": 18743,
"sha256": "ce609764fa53f7a86e77dc8f2c00c6d8b878c4c9b6f885fb2618bf8f932cad04"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/build-result.json",
"bytes": 729,
"sha256": "2df4873188d94c3745631b6e774e76a7646b3366fab7a3badf7ac1be136f7bd9"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/cache-reclaim.json",
"bytes": 2482128,
"sha256": "8dc7866b30bfd7fed339a3cd4a2dd40c0ec8f3b471cb004fc14f303a41642ad8"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-fixture/cmd/erasure-server-pool-consistency_test.go",
"bytes": 54157,
"sha256": "c3c1bf441e5f97fd2648c5fc9b89cb11679018e349daa4d0eef4ebbfada322db"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-fixture/cmd/post-policy_test.go",
"bytes": 33420,
"sha256": "697f8a08dceae481fd1aae7b5e7f3906b55b34fe9928688456eaa943eba79020"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-fixture/cmd/test-utils_test.go",
"bytes": 79596,
"sha256": "fb4847b10d3d59c7d62ee79a61d54e8c0bc93d6eb32cb520f5662bb7b52750f5"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-fixture-manifest.json",
"bytes": 426,
"sha256": "8940a56a6f46b7e9c236c3a8d39d927735954ae42601ca52c4a190339fb1f21f"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-fixture-overlay.json",
"bytes": 368,
"sha256": "3b8586973d426f78145aa25f0c3c9d48faf93678ffe9410fc9aa089cb6167af0"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-fixture.patch",
"bytes": 842,
"sha256": "3e3732c7fab2b95b9f2e80a6ee973a588600f84eeb2b74c2f15a09373228247f"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-post-manifest.json",
"bytes": 205,
"sha256": "34ae2c860a5cc1a9615d7b410eff781f5f01d54dba65f66edae0f724d3d232c7"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-post-overlay.json",
"bytes": 522,
"sha256": "8d0b7594481fa9028fbc4284e1e94cb853828925423d2a7b3cd6f5cbabb82e3c"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/capacity-post.patch",
"bytes": 358,
"sha256": "f4aa03d4a0a9f0bb220fa3b3b988a8dda1ad7d6764daaeb4e60eb4ee4e996674"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/discussion-baseline.log",
"bytes": 1029,
"sha256": "9581de36ec9a403d304c32192d17265b1e9c9414c60e9f2cb57321faebbe23dc"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/final-check-results.json",
"bytes": 939,
"sha256": "6652d2db1ac98d65232e37eda7738972a74f9569ac63b534f1af798ebf19f642"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/final-post-and-pools.log",
"bytes": 1280,
"sha256": "e7e5fec0761976470eafbf31bcefe4abc241525514f6c3b9a2baa035354144fa"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/fixed-resync-isolated.log",
"bytes": 116,
"sha256": "bed15b381379998bbe2a0aa1be19c2cfec1b0063886143dbe82918f9652c28ff"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/fixed-targeted-final.log",
"bytes": 25221,
"sha256": "32283f5d7eea5ce4974fefa0724a1c4de565bb0ef9a0f4fbcad68140bccc32b1"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/fixed-targeted-latest.log",
"bytes": 32772,
"sha256": "45f362258e21b631cb5ebcd15dec98bd2fa3186f509f18fe216d7cd0ebdb5a9c"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/fixed-targeted.log",
"bytes": 27709,
"sha256": "a30f7754f90cfade50ed80a066c5e2bd53faf225bd615cd07b9cf1350b2e3139"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/golangci-lint-serial",
"bytes": 103,
"sha256": "b2a00c2702468165a7851ee3a6addbef9e581833a33492d44ac66d531cfc4fff"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/implementation-files.json",
"bytes": 936,
"sha256": "7bd1279e1c99da9da4562cda6e0265d36d53a9bb546f10d42c4174c1a34b5c70"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/implementation-v1.patch",
"bytes": 48399,
"sha256": "8f6f76ee874c43b0827fde272e8a947f118efb1c1af92bf4a02ef88f93554c1b"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/make-build.log",
"bytes": 55,
"sha256": "6ba9b545236be964861749c72e7609edf12b8f470df30d1ede8fd62f497e629b"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/make-verifiers-final.log",
"bytes": 217,
"sha256": "d973482061716daa245e7d7162765dec0e6ecd5fee41249d2702a2d8bfce1c32"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/make-verifiers-serial.log",
"bytes": 410,
"sha256": "e42a5bb55f5c1ebfcf02cebebf6d82cf1ec5a2d74590cdf838deba16dd80bfdf"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/make-verifiers-success.log",
"bytes": 162,
"sha256": "b4982b6a7302e733c7bec4a5fb36b8ee8865fe95f1595e7079ce7f8455406210"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/make-verifiers.log",
"bytes": 285,
"sha256": "42f4b147ef4aac45ba91457e7932db30f9b57f285466ee3f10bbaa9a911e5bc7"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/matrix-overlay.json",
"bytes": 153,
"sha256": "a0fb5eb71ebb2bdb3374813752de5867848454794bcbf302ba0f6d7d4f6c0519"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/matrix_test.go",
"bytes": 22996,
"sha256": "c9cf527c63800fa045bdf0b8e95d740b81a3d5be3a08c74811fa5c06bed56d4f"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-implementation-request.json",
"bytes": 909,
"sha256": "9fab15ce1cfb5bad102b1880968e4731a7b5cb02d6d01e6cb2caf8bc9029a150"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-implementation.jsonl",
"bytes": 1184150,
"sha256": "fef155a7382c8f66f69b7afd5fd94559edcc6f72fc13aeb7ef01319c22c09861"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-implementation.stderr.log",
"bytes": 0,
"sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v1-request.json",
"bytes": 216,
"sha256": "175e013154c241805e00368f7841d41faf48ee847ff5251aad96ae57831e244a"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v1.jsonl",
"bytes": 1017882,
"sha256": "63c00d8362f236a18293e1a637eff3b7c7e38b0bbd11805f75d91e8774efcdb2"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v1.stderr.log",
"bytes": 0,
"sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v2-request.json",
"bytes": 216,
"sha256": "ed710eff03d3ebabab277c9e453048097a8649582df4b01f99d0ee0ad3a7c831"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v2.jsonl",
"bytes": 414265,
"sha256": "e888fbf38ba7fd49891e0757c18006875d02a6bbb325906a31fc8d98e54d0e39"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v2.stderr.log",
"bytes": 0,
"sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/overlay.json",
"bytes": 142,
"sha256": "d51934eb99e2b19d149478e090ec327ed2753a5ad2a026c8745b8e2554962a00"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/related-filter.txt",
"bytes": 5771,
"sha256": "f022bc24ae0fe391ae51a5095db1d2e415a994934327c7508e6c65edc313cc78"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/related-names.txt",
"bytes": 5767,
"sha256": "369ed4b6742d15e8ab4d790615842304a7178fbdc598e231e9205fc096c0785a"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/related-suite-capacity-final.log",
"bytes": 137077,
"sha256": "322d4854ba909bd99d5c7740abeeac05eb76cb2d39d2c7541642335ae6a1ffe9"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/related-suite-capacity.log",
"bytes": 3362,
"sha256": "1e4f1bc6be2b4339a0d9b7a2e951774fc24ec5521be9fcf53b2ce02031f5cdbe"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/related-suite-rest.log",
"bytes": 172665,
"sha256": "846d10084299a77253153c0eed8546e275c789a0289a4198052773e49a73423f"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/related-suite.log",
"bytes": 3521,
"sha256": "38e3e8e4b7ae815fce40931009a0d4755601a4f3f7f9f44a569edff024c9239b"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/sender_test.go",
"bytes": 5140,
"sha256": "a1a58fd6968b41cf6c565d9f63a1d0fa1c907f008f70acfd13c7a6c525376357"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/silo-version.log",
"bytes": 317,
"sha256": "8ce9c5d15082a78e696aa79f8ec007f72ce969ce6ebd7dab2f7db69b20b51f8b"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/targeted-race-final.log",
"bytes": 39753,
"sha256": "127cfb73f415bad43e2fd79c05150ab765322dcadb29e687b752b6174d4ad850"
},
{
"path": "/Users/vonng/tmp/silo-r5-20260915-77ad/verifiers-result.json",
"bytes": 417,
"sha256": "6faea89c420685ccae0642be88ddf86938bb25e24f066fd9e57138d16c9e9856"
}
],
"binary": {
"path": "/Users/vonng/.codex/worktrees/77ad/silo/silo",
"bytes": 93070802,
"sha256": "dd789126966d4a42bc0a9bcd8b8eab9714e3eadc7a524505a6224ea6d76c750f"
},
"scope": "Local development build and exact raw verification/review records; no publication or production acceptance."
}
@@ -0,0 +1,20 @@
{
"research_base": "9ebe81c1b3611f9cc73e676b5b741c2be62c467a",
"tested_dependency": "dbcf8dec589deb5d91e17d295cb70997635f5b55",
"reviewed_patch_sha256": "8f6f76ee874c43b0827fde272e8a947f118efb1c1af92bf4a02ef88f93554c1b",
"plan_version": "v2",
"plan_sha256": "5a782acf3f285b23d1ae43a73481c4eb772a9a6d917fc5a550ecfc7cbf7446ca",
"files": {
"cmd/bucket-replication.go": "cbab22ffe316fabc076e7f4a1fa5b1417d07b265ae9dc1b27454689355926d35",
"cmd/erasure-object.go": "4bc848685ea714d88cabbd5d1b8585fbcc06f7b19c775e1a811030e783d0e1a4",
"cmd/erasure-server-pool-consistency.go": "d2736ef6bffbb5c5758eba8df38f8d4ecb888a838ab0de8ad3cf015c051f8ad7",
"cmd/erasure-server-pool.go": "87ad0b25dfa3081d0e63d0073b788614a9c88e2498a2ce0956b93f8a0a03ef53",
"cmd/object-handlers-common.go": "101bd7d7447072d13fed50983b69b562e4725632645e623d7fdd490f388ecdec",
"cmd/object-handlers.go": "61897a260f3f5f660f41edcb50956c60e914ef98f9a987da824f16d78171fde2",
"cmd/object-multipart-handlers.go": "d9622c69c540ab32dd23916e3f534b6886473a98370c9dd17673e69a423b2a7e",
"cmd/replication-tagging-order_test.go": "c8260b4ccf82fa615e1e24b35a07f2d1aacbcf776e5c6f9dadffea4a09ad6ea8",
"cmd/replication-tagging-sender_test.go": "3770a1a48a6efe58fe8127e1e4fdf6bd7cf171e17db20f15222ea2f7b85db1af"
},
"production_unchanged_after_review": true,
"delivery_dependency": "af2b1794d38d9e70e1d2c3ee692426e4b6cab4bd"
}
@@ -0,0 +1,17 @@
{
"base_commit": "dbcf8dec589deb5d91e17d295cb70997635f5b55",
"plan_version": "v2",
"plan_sha256": "5a782acf3f285b23d1ae43a73481c4eb772a9a6d917fc5a550ecfc7cbf7446ca",
"patch_sha256": "8f6f76ee874c43b0827fde272e8a947f118efb1c1af92bf4a02ef88f93554c1b",
"files": {
"cmd/bucket-replication.go": "cbab22ffe316fabc076e7f4a1fa5b1417d07b265ae9dc1b27454689355926d35",
"cmd/erasure-object.go": "4bc848685ea714d88cabbd5d1b8585fbcc06f7b19c775e1a811030e783d0e1a4",
"cmd/erasure-server-pool-consistency.go": "d2736ef6bffbb5c5758eba8df38f8d4ecb888a838ab0de8ad3cf015c051f8ad7",
"cmd/erasure-server-pool.go": "87ad0b25dfa3081d0e63d0073b788614a9c88e2498a2ce0956b93f8a0a03ef53",
"cmd/object-handlers-common.go": "101bd7d7447072d13fed50983b69b562e4725632645e623d7fdd490f388ecdec",
"cmd/object-handlers.go": "61897a260f3f5f660f41edcb50956c60e914ef98f9a987da824f16d78171fde2",
"cmd/object-multipart-handlers.go": "d9622c69c540ab32dd23916e3f534b6886473a98370c9dd17673e69a423b2a7e",
"cmd/replication-tagging-order_test.go": "64d6dfe3436970caeafcb914157bdedac5982a2105fe72c1753a8d68cf7ed6ef",
"cmd/replication-tagging-sender_test.go": "3770a1a48a6efe58fe8127e1e4fdf6bd7cf171e17db20f15222ea2f7b85db1af"
}
}
@@ -0,0 +1,26 @@
# R5 implementation review disposition
Real reviewer: `claude-opus-5`, explicit `--effort max`, session `599b4759-add2-4a41-b5b2-865af7a2c096`.
Verdict: **GO_WITH_NONBLOCKING_NOTES; 0 blockers**. Raw review is preserved verbatim in `opus-implementation-review.md`; model usage, original plan/patch hashes and raw log location are in `opus-implementation-metadata.json`.
The accepted v2 plan remains immutable. The following implementation notes supplement it; they do not retroactively change the hash on which plan consensus was reached.
## Nonblocking notes
- **N1 accepted:** a scheduled metadata COPY can rewrite object data when the receiver applies bucket-default/automatic KMS encryption. Its cost can therefore exceed metadata I/O. The existing completed-object/scanner and incoming-replica scheduling gates still prevent a feedback loop. No new transfer optimization or HEAD protocol is introduced.
- **N2 accepted:** a malformed recorded source tag timestamp fails sender construction and remains a retry failure until an explicit correct tag mutation/repair supplies a valid revision. A missing revision is different from a present invalid/empty value. No historical time is fabricated, and no automatic production rewrite is performed.
- **N3 retained scope:** existing marker/trust/REPLICA/version predicates are preserved. Production sender requests satisfy the relevant predicates; R5 does not broaden replication trust.
- **N4 accepted compatibility change:** a trusted metadata COPY without a source tag revision preserves stored tags, including the metadata-REPLACE shape. This is the deliberate missing-revision rule in plan C, and is tested under UUID/null versions and unqualified COPY.
- **N5 accepted:** ordinary COPY records its chosen tag state, including an empty REPLACE and unchanged tags during key rotation, as a fresh local event. This is consistent with the accepted last-writer-wins scheme.
- **N6 no change:** all production writers use the lowercase reserved timestamp key. Case-insensitive sender lookup is compatible with those writers and existing lock timestamp handling.
## Coverage notes
- **L1:** the review was supplied a passing run with **13**, not 12, top-level R5 tests. Its verdict explicitly did not claim execution of the wider tests. The wider selection reproduced the same `TestReplicationResync` order-dependent initialization panic on the unmodified production baseline; that test passes in isolation on both baseline and R5. Host-capacity and actual ENOSPC failures are retained, not reported as passes. Final related, race and static/build results are recorded separately in `verification.md`. An unfiltered full `cmd` package run remains an integration check before any later merge; this task delivers a local patch and does not claim that full-package or multi-site production gate passed.
- **L2 addressed:** after every incoming multi-pool replay, the R5 test now rereads the addressed version through normal pool routing and checks its empty value and deletion revision. The per-pool checks still inspect every retained copy. This prevents a vacuous pass if all copies disappear. The test deliberately allows existing duplicate suppression to retain both identical copies; existing pool cleanup/retry tests separately exercise retirement.
- **L3 accepted boundary:** the combined KMS cases exercise destination encryption and plaintext readback; source fixtures are populated through storage APIs. They do not establish encrypted-source-to-encrypted-destination replication across two running sites. SSE-C key rotation has its own signed HTTP and decrypted GET test.
- **L4 confirmed:** both the actual SDK default metadata directive and peer metadata-REPLACE shapes are exercised.
## Changes after review
Production code is unchanged from reviewed patch SHA256 `8f6f76ee874c43b0827fde272e8a947f118efb1c1af92bf4a02ef88f93554c1b`. Test-only follow-up adds the L2 normal-routing read and applies the repository's gofumpt formatting. `implementation-manifest.json` records the exact reviewed files; the final verification manifest records the final files, so the two versions are distinguishable.
@@ -0,0 +1,46 @@
{
"model": "claude-opus-5",
"effort": "max",
"baseline": "dbcf8dec589deb5d91e17d295cb70997635f5b55",
"plan_version": "v2",
"plan_sha256": "5a782acf3f285b23d1ae43a73481c4eb772a9a6d917fc5a550ecfc7cbf7446ca",
"patch_sha256": "8f6f76ee874c43b0827fde272e8a947f118efb1c1af92bf4a02ef88f93554c1b",
"actual_assistant_models": [
"claude-opus-5"
],
"session_id": "599b4759-add2-4a41-b5b2-865af7a2c096",
"is_error": false,
"modelUsage": {
"claude-haiku-4-5-20251001": {
"inputTokens": 2125,
"outputTokens": 15,
"cacheReadInputTokens": 0,
"cacheCreationInputTokens": 0,
"webSearchRequests": 0,
"costUSD": 0.0022,
"contextWindow": 200000,
"maxOutputTokens": 32000,
"thinkingTokens": 0,
"canonicalModel": "claude-haiku-4-5",
"provider": "firstParty",
"costBasis": "list"
},
"claude-opus-5": {
"inputTokens": 106,
"outputTokens": 64414,
"cacheReadInputTokens": 7275193,
"cacheCreationInputTokens": 236959,
"webSearchRequests": 0,
"costUSD": 7.618066499999999,
"contextWindow": 1000000,
"maxOutputTokens": 64000,
"thinkingTokens": 45128,
"canonicalModel": "claude-opus-5",
"provider": "firstParty",
"costBasis": "list"
}
},
"result": "GO_WITH_NONBLOCKING_NOTES",
"blocking_items": 0,
"raw_output": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-implementation.jsonl"
}
@@ -0,0 +1,27 @@
Review the actual R5 implementation independently for correctness and regressions, using Claude Opus 5 at max effort. This is a read-only final code review after an already recorded two-round plan consensus. Do not edit files. Do not simulate tests or claim you executed them. Read the relevant source and evidence yourself; focus on material blockers and minimal compatible fixes.
Working tree: /Users/vonng/.codex/worktrees/77ad/silo
Base dependency commit: dbcf8dec589deb5d91e17d295cb70997635f5b55 (R4 KMS timestamp field, one production addition)
R5 plan v2 SHA256: 5a782acf3f285b23d1ae43a73481c4eb772a9a6d917fc5a550ecfc7cbf7446ca
R5 implementation patch SHA256: 8f6f76ee874c43b0827fde272e8a947f118efb1c1af92bf4a02ef88f93554c1b
Manifest: /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/implementation-manifest.json
Frozen review patch (7 production files + 2 new tests): /Users/vonng/tmp/silo-r5-20260915-77ad/implementation-v1.patch
Plan: /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/plan-v2.md
Prior actual review: /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/opus-v2-review.md
Consensus and disagreements: /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/consensus.md, /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/opus-v1-response.md
Baseline reproduction: /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/reproduction.md
Latest new regression run: /Users/vonng/tmp/silo-r5-20260915-77ad/fixed-targeted-latest.log (PASS, 9.161s; signed HTTP and actual storage on single/16 disks, null/UUID, COPY default and REPLACE, PUT/multipart, KMS plaintext GET, SSE-C rotation, multi-pool, retry and stale source ACK; test names and details in source.)
Additional related suite and race/static validation are ongoing and are not yet accepted. An expanded suite hit existing TestReplicationResync initialization panic before R5 tests; baseline isolation is ongoing. Do not treat that as a proven R5 regression or a passing test.
Research and implementation points to scrutinize:
1. Empty tag values are ordered states only with recorded timestamp; no fabricated tombstone for empty legacy object. Nonempty legacy sender falls back to ModTime. Malformed stored timestamp fails PUT and metadata COPY sender construction.
2. Local PUT/DELETE tagging timestamps are unconditional, advance under existing storage locks; pooled mutation computes a single revision > every copy; do not mutate caller map. Replicas keep source ordering and equal timestamp stored-wins.
3. Actual sender getCopyObjMetadata + minio Core.CopyObject sends tagging REPLACE with no metadata directive; peers may also send metadata REPLACE. Capture stored timestamp before reconstruction, accept incoming empty value, prevent stale SSE-C encMetadata snapshot overwrite, keep existing storage rechecks.
4. PUT/multipart init persist parsed trusted timestamp. Matching ETag/version no-op is relaxed only for trusted newer tag revision, retaining explicit client preconditions and SSE-C behavior. Multipart completion rechecks revision under lock.
5. Equal visible tag values can hide a newer deletion/re-addition revision, so scheduled metadata/heal work with recorded timestamp sends metadata. Preserve existing Completed/scanner and incoming no-reschedule gates.
6. Replication status ACK no longer writes old queued ri.UserTags over the current tags. Preserve current value/timestamp under metadata lock including multi-pool.
7. Scope explicitly excludes tag-filter target eligibility changes, historical tombstone recovery, arbitrary independent unversioned content replacements, clock skew consensus, and production multi-site acceptance. R4 owns object-api-options.go; it is not in this R5 diff.
There was an earlier KMS test-only fixture error: it assigned a single-PUT ETag to a multipart-layout retransmission; corrected by seeding a real multipart source and using its actual ETag/layout (full latest run passes plaintext read). Inspect the fixture rather than inferring production encryption correctness solely from prior failed test output.
Return a clear GO / GO_WITH_NONBLOCKING_NOTES / REQUEST_CHANGES verdict; blocking count; concrete source locations with cause, trigger and effect for any blockers; tests/coverage limitations if material. State which baseline, plan hash and patch hash you reviewed (you may quote supplied hash, but cannot independently compute with Read/Grep/Glob). Do not demand unrelated broad redesign. Distinguish actual source defect from nonblocking improvement or scope limitation.
@@ -0,0 +1,58 @@
Write is disabled in this session, so I'll deliver the review directly rather than to the plan file.
# R5 implementation review — tag deletion and ordered replication
## Verdict: **GO_WITH_NONBLOCKING_NOTES** — blocking count: **0**
Read-only review. No files edited, no tests run, no results simulated. Every pass I cite is quoted from the supplied log.
## What I reviewed
| Item | Value |
|---|---|
| Working tree | `/Users/vonng/.codex/worktrees/77ad/silo`, branch `codex/r5-tag-deletion-ordering` |
| Baseline dependency | `dbcf8dec589deb5d91e17d295cb70997635f5b55` (R4 SSE-KMS field) |
| Plan | `plan-v2.md`, SHA256 `5a782acf…7fc7446ca` *(quoted)* |
| Patch | `implementation-v1.patch`, SHA256 `8f6f76ee…f88f93554c1b` *(quoted)* |
**Hash caveat:** with Read/Grep/Glob I **cannot compute** either SHA256 — both are quoted from the manifest/task. What I did verify by reading is that every hunk in the frozen patch is present verbatim in the working tree across all seven production files. That is source identity by inspection, not by digest. I also confirmed R5 does not touch `cmd/object-api-options.go`; the R4 branch carrying `ReplicationSourceTaggingTimestamp` (`object-api-options.go:449-460`) is unmodified.
## Per-claim findings
**1. Empty values are ordered states; no fabricated legacy tombstone.** Confirmed. `replicationTaggingTimestamp` (`bucket-replication.go:786-794`) returns the recorded stamp even with empty tags, falls back to `ModTime` only for non-empty tags, zero otherwise. Used by both `putReplicationOpts` (`:861-870`) and the metadata sender (`:1702-1707`). The SDK omits the header for a zero time (`minio-go@…60bd07042d49/api-put-object.go:236-238`, `api-compose-object.go:286-288`), so "no revision" really travels as absence. Malformed stamps fail both constructions.
**2. Local revisions unconditional and monotonic.** Confirmed. Both handlers mint one `UTCNow()` outside the `dsc.ReplicateAny()` branch (`object-handlers.go:3773-3778`, `:3876-3881`); `getOpts` leaves `opts.UserDefined` nil (`object-api-options.go:110`,`:39`), so the unconditional map replacement drops nothing. `er.PutObjectTags` applies the guard under the existing NS lock (`erasure-object.go:2273-2282`, `:2330-2334`); an absent stamp yields `""` and preserves legacy direct-storage semantics. `z.PutObjectTags` folds one candidate strictly beyond every copy and **clones** first (`erasure-server-pool.go:3054-3062`) — `opts` is a value parameter and `er.PutObjectTags` never writes `opts.UserDefined`, so no caller map is mutated. No replica path reaches `PutObjectTags` (the only two production callers are the tagging handlers), so replicas keep strict source ordering via `reconcileStoredObjectTags`, stored-wins on ties (`erasure-server-pool-consistency.go:238-242`).
**3. COPY receiver.** Confirmed. Stored pair captured before reconstruction (`object-handlers.go:1800`); `srcInfo.UserTags` is never reassigned between the source read and the decision, so it genuinely is stored state. The decision block (`:1818-1837`) accepts an incoming empty value with a stamp and rechecks the captured state; all existing in-lock rechecks still run (`erasure-object.go:136-138`, `:1312-1315`; `erasure-multipart.go:1161-1190`; `erasure-server-pool.go:1443-1450`). The `encMetadata` fix (`:1840`) is safe and correctly placed — `encMetadata` receives reserved keys only on the SSE-C rotation path (`:1655-1659`), and the delete lands after `rotateKey`/`newEncryptReader` and before the merge at `:1910`.
**4. PUT/multipart persistence and the duplicate exception.** Confirmed. `putOptsFromHeaders` aliases `opts.UserDefined = metadata` in both branches (`object-api-options.go:451`,`:464`), so post-build writes reach storage (`object-handlers.go:2323-2325`; `object-multipart-handlers.go:315-318`, correctly *after* `maps.Copy(metadata, encMetadata)` at `:300`). The relaxation (`object-handlers-common.go:243-246`) sits below the explicit `If-Match`/`If-None-Match` checks, is gated on `isReplicaTrusted` + `olderThan` (zero source never wins, `bucket-object-lock.go:370-376`), and leaves the SSE-C exemption intact. It cannot loop: once the write lands the stamps are equal and the next attempt 412s. `completeMultipartOpts` sets neither `PreserveETag` nor a tagging timestamp (`object-api-options.go:501-550`), so completion needs no new exception and reconciles under the lock (`object-multipart-handlers.go:1201`).
**5. Equal values can hide a newer revision.** Confirmed and correctly scoped. The gate (`bucket-replication.go:1013-1018`) sits after the null-version resync exclusion and after **every** branch that can return `replicateAll`; from there only `replicateMetadata`/`replicateNone` are reachable, so it can never downgrade a needed full transfer. It is reached only from `replicationActionForTarget` → `replicateAll` (`:1608`), not from the object-replication fast path (`:1328-1343`). The Completed gate (`:3775`) and failures-only requeue (`:1316`) bound the work, and an incoming replica COPY schedules no outgoing event. Existing fixtures carry no tagging stamp (`bucket-replication_test.go:716-739`), so they are unaffected.
**6. ACK no longer overwrites current tags.** Confirmed removed (`bucket-replication.go:1276-1286`). Preservation holds on both write-backs: `er.PutObjectMetadata` copies from `ObjectInfo.UserDefined`, which `cleanMetadata` strips of `x-amz-tagging` (`object-api-utils.go:403-407`; `erasure-object.go:2254-2260`); `updatePoolMetadata` falls back to merged `UserTags` and rewrites the merged newest stamp (`erasure-server-pool-consistency.go:194-214`). Both under the object lock (`erasure-object.go:2196-2205`; `erasure-server-pool.go:3020-3029`). The sender also re-reads current state first (`bucket-replication.go:1527-1550`).
**7. Scope.** Respected — no tag-filter eligibility change, no historical tombstone invention, no clock-skew consensus, no `object-api-options.go` change.
**Trust boundary re-checked:** the reserved key cannot be injected from the wire — `containsReservedMetadata` rejects the whole `X-Minio-Internal-` class outside the SSE allowlist (`generic-handlers.go:75-85`), and `extractMetadataFromMimeWithReplication` maps only `replicationToInternalHeaders` (`handler-utils.go:258-298`).
**KMS fixture inspected directly**, not inferred from prior output: the multipart case now seeds a real multipart source and reuses its actual ETag/part layout (`replication-tagging-order_test.go:472-489`). The earlier single-PUT-ETag mismatch is gone. See L3 for what it still does not cover.
## Non-blocking notes (no change required)
- **N1 — on encrypted destinations the forced metadata COPY is not metadata-only.** The gate at `bucket-replication.go:1013-1018` yields a replica COPY; with bucket-default/auto KMS the destination applies SSE before `copyDstOpts` (`object-handlers.go:1428-1433`) and then clears `srcInfo.metadataOnly` (`:1669-1677`) — so it **rewrites object data**. Bounded to one COPY per object entering heal and one per object per explicit resync (not a loop), but the plan's "extra metadata I/O" understates this case. Worth a sentence in the cost note.
- **N2 — fail-closed on a malformed stored revision is terminal for that object** (`:786-794` → `:1702-1707`/`:867-870`, requeued by MRF at `:1316-1322`). No production writer can produce such a value, so this is a defensive tail. Note the asymmetry: storage self-heals the same corruption (invalid *stored* ⇒ incoming wins, `erasure-server-pool-consistency.go:238-242`) while the sender refuses to proceed. The minimal hardening, if ever wanted, is to treat a present-but-**empty** value as absent — I traced no reachable path producing one, and the current behavior is what plan and consensus chose, so I am not asking for it.
- **N3 — trusted-marker vs REPLICA asymmetry (pre-existing).** `object-handlers.go:2323` / `object-multipart-handlers.go:316` persist on `opts.ReplicationRequest`, while the in-lock recheck needs `isReplicaTrusted` **and** a version ID (`:2442`). Production sets both; the new precondition exception uses the stricter predicate. Accepted in plan §D.
- **N4 — a trusted metadata COPY with no source revision now ignores the request's tag value** (`object-handlers.go:1826-1833`). For a peer sending `x-amz-metadata-directive: REPLACE` without a revision, the value used to land (`X-Amz-Tagging` is in `supportedHeaders`, `handler-utils.go:271-283`). No MinIO sender produces that shape, and the R4 case `object-copy-replication-tagging_test.go:89` already expects stored-wins there, reaching it via the storage reconcile. Deliberate per plan §C.
- **N5 — ordinary COPY always writes an explicit tag value plus a fresh revision** (`:1834-1837`). (a) `x-amz-tagging-directive: REPLACE` with no tags now genuinely clears the destination, where the default-metadata path used to carry source tags forward — an S3 conformance improvement, covered by `TestAPILocalTaggingAlwaysAdvancesRevision`. (b) An in-place key-rotation COPY advances the revision without changing any value, re-asserting current tags against an older in-flight remote deletion. Both follow from last-writer-wins as specified.
- **N6 — cosmetic.** The key is read case-insensitively at `bucket-replication.go:787`/`:1016`, exactly elsewhere. `TaggingTimestamp` is lowercase (`:74`) and storage writes only lowercase, so they agree; the same mix already exists for lock timestamps in that file (`:896` vs `:1708`).
## Tests and coverage limitations (material)
- **L1 — the only established green result is the 12 R5 tests** (`fixed-targeted-latest.log`, `ok … 9.161s`): signed HTTP through real single-disk and 16-disk storage, null/UUID, COPY default and REPLACE, PUT/multipart, KMS, SSE-C rotation, multi-pool, sender retry and stale ACK. The wider `cmd` package, `-race`, `gofmt` and `git diff --check` are ongoing, and the `TestReplicationResync` panic is unattributed. I treat that as an **open verification item**, not a regression and not a pass. My static read found no existing test whose expectations R5 flips — I checked the `getReplicationAction` fixtures, the R4 KMS COPY table including its `missing-timestamp` case (R5 satisfies it via the handler instead of the storage reconcile), and confirmed no pre-existing test calls `PutObjectTags` with a supplied revision.
- **L2 — multi-pool convergence after an *incoming* replica write is asserted weakly.** `replication-tagging-order_test.go:688-698` skips any pool whose version is absent and only asserts a retained copy is correct; it pins neither which pool is retained nor that retirement ran. The case that matters for §A — one revision strictly greater than every copy, in the response *and* every stored copy — is pinned by `TestLocalTaggingCommitCannotRegressRevision`.
- **L3 — the KMS variant's source objects are plaintext.** Seeds use `obj.PutObject` (`:468`, `:477-488`), bypassing handler encryption, so encryption enters only via the incoming request and the destination bucket default. Real coverage of the R4 field on the receive side, but not encrypted-source-to-encrypted-destination end to end. The trailing plaintext GET (`:520-523`) does establish the final object is readable.
- **L4 — wire shape is well pinned.** `TestTaggingProductionCopyWireShape` asserts the real SDK request (`metadata-directive=""`, `tagging-directive=REPLACE`) against a live peer, and `r5Receive`'s `"copy"` operation covers peer metadata-REPLACE independently. Both required shapes are present.
## Recommendation
The production diff is merge-eligible as written; I found no actual source defect. The single gating action before merge is closing **L1** — a clean full-package run (plus `-race`, `gofmt`, `git diff --check`) with the `TestReplicationResync` panic isolated against the unpatched baseline. N1 and N2 deserve a sentence each in the plan's cost/limitations section; N3–N6 are already covered by plan §C/§D and need no action.
@@ -0,0 +1,44 @@
{
"model": "claude-opus-5",
"effort": "max",
"plan_version": "v1",
"plan_sha256": "fd6051527ebf19f624125bd3238da2f938420917226387c0f9373f1a98e87993",
"baseline": "9ebe81c1b3611f9cc73e676b5b741c2be62c467a",
"actual_assistant_models": [
"claude-opus-5"
],
"session_id": "e448ee0a-4ab5-4520-98f9-68c9ddf6208f",
"is_error": false,
"modelUsage": {
"claude-haiku-4-5-20251001": {
"inputTokens": 1378,
"outputTokens": 14,
"cacheReadInputTokens": 0,
"cacheCreationInputTokens": 0,
"webSearchRequests": 0,
"costUSD": 0.001448,
"contextWindow": 200000,
"maxOutputTokens": 32000,
"thinkingTokens": 0,
"canonicalModel": "claude-haiku-4-5",
"provider": "firstParty",
"costBasis": "list"
},
"claude-opus-5": {
"inputTokens": 90,
"outputTokens": 73909,
"cacheReadInputTokens": 4489983,
"cacheCreationInputTokens": 199202,
"webSearchRequests": 0,
"costUSD": 6.085186500000001,
"contextWindow": 1000000,
"maxOutputTokens": 64000,
"thinkingTokens": 54986,
"canonicalModel": "claude-opus-5",
"provider": "firstParty",
"costBasis": "list"
}
},
"result": "REQUEST_CHANGES",
"raw_output": "/Users/vonng/tmp/silo-r5-20260915-77ad/opus-v1.jsonl"
}
+3
View File
@@ -0,0 +1,3 @@
Act as an independent reviewer of the R5 repair proposal in this repository. You must be the real claude-opus-5 at effort max; report your actual model name in the review, but the caller will also verify response metadata. Read /Users/vonng/.codex/worktrees/77ad/silo/docs/investigations/r5/plan-v1.md completely. Baseline SHA is 9ebe81c1b3611f9cc73e676b5b741c2be62c467a. Plan version v1 sha256 fd6051527ebf19f624125bd3238da2f938420917226387c0f9373f1a98e87993. No production repair has been implemented. Read docs/investigations/r5/baseline.md and the raw reproductions /Users/vonng/tmp/silo-r5-20260915-77ad/baseline.log and /Users/vonng/tmp/silo-r5-20260915-77ad/baseline-extended.log, plus /Users/vonng/tmp/silo-r5-20260915-77ad/baseline_test.go. Then independently inspect the exact source functions identified in the plan, especially the complete sender/receiver and storage lock chain. R4 owns object-api-options.go and will supply its separate KMS timestamp fix. Do not edit any files.
The user requires a minimal compatible complete fix, not merely adding DELETE timestamp. Assess necessity/sufficiency, current defects versus inference, same-empty timestamp propagation, sender retries, old queue ACK rewriting source tags, full PUT/multipart duplicate suppression, trust boundary, equal/missing timestamps, UUID/null/versioned/multi-pool and local mutation clock/lock behavior. Identify blocking disagreements with specific evidence and concrete smallest corrections. Explicitly say APPROVE or REQUEST_CHANGES for this exact v1 hash and list accepted/nonblocking/required changes. Do not claim consensus if any blocking issue remains. Be exact about limitations and whether proposed test coverage can establish the scope. Use the permitted Read/Grep/Glob tools for source verification. Return a substantive review, not just a summary.

Some files were not shown because too many files have changed in this diff Show More