Compare commits

..

178 Commits

Author SHA1 Message Date
Feng Ruohang e791640dac docs: describe merged conditional PUT behavior and recovery guidance
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 11:21:16 +08:00
Feng Ruohang 9b4ae82a29 Merge pull request #207 from pgsty/codex/conditional-put-pools
fix(pools): evaluate cross-pool PUT conditions against the current object
2026-09-16 11:17:02 +08:00
Feng Ruohang 4620be394b test: synchronize conditional PUT disk fixtures
Replace unsynchronized getDisks swaps with backing disk-list updates under
erasureDisksMu, matching the existing GetDisks reader lock. Apply the same
helper to capacity and read-fault adapters while preserving nested restore
ordering.

Add a regression that overlaps fixture changes with the real IAM Walk
reader, and run the conditional PUT suite under the race detector in CI.
The regression reproduces the old fixture race; ten fixed race iterations
pass without warnings. Production conditional PUT behavior is unchanged.

Refs #199

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 10:53:24 +08:00
Feng Ruohang 5e7d603083 fix(pools): evaluate cross-pool PUT conditions against the current object
Multi-pool PUT selected a destination by capacity and evaluated
If-Match/If-None-Match only against that destination's local object
state. An empty or stale destination could accept a stale ETag or
If-None-Match:* while another pool held the current object, replacing
it; a current ETag could instead be rejected with 412 or 404.

Under PUT's existing pools-layer object lock, resolve the comparison
object with objectPoolInfos (including draining pools), treat a latest
delete marker as absence, fail closed on unreadable pool metadata, and
clear an accepted callback before destination dispatch. Replica and
data-movement callbacks keep their addressed-version semantics and
metadata reconciliation.

Reproduced on 40220bd836 and RELEASE.2026-09-03T13-18-01Z with six
signed HTTP scenarios: four defect cases failed, two controls passed.

Refs #199

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:21:21 +08:00
Feng Ruohang fb7c406ddc Merge pull request #206 from pgsty/codex/main-consolidation-20260916
test: restore valid credentials in integration fixtures
2026-09-16 08:18:56 +08:00
Feng Ruohang 416826f61f Merge pull request #205 from pgsty/codex/release-notes-followups
docs: complete September correctness and upgrade notes
2026-09-16 08:16:23 +08:00
Feng Ruohang a2e2f3ee82 test: restore valid credentials in integration fixtures
Port the shell-fixture changes from ebc9937d97b27871dcc4bb91d4b5771d3550b76a. Match the existing minimum secret length consistently across startup, aliases and helper commands.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:08:08 +08:00
Feng Ruohang 70c7ec4a9f docs: link reviewed runbook sources before website deployment
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:04:53 +08:00
Feng Ruohang 2b722b7e92 docs: complete September correctness and upgrade notes
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 07:51:11 +08:00
Feng Ruohang 40220bd836 Merge pull request #196 from pgsty/codex/merge-r4-r8
Complete R5, R6 and R8 on the main baseline containing R4 and R7.

Preserve the individual signed repairs, independent Opus 5 Max review, and integration validation evidence.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 01:14:06 +08:00
Feng Ruohang df0dfa0a34 docs: record R4-R8 integration review and validation
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:58:31 +08:00
Feng Ruohang 80684fed59 chore: align integrated repair tests with contribution checks
Use the actual author and AGPL-3.0-or-later notices for new R6/R8 tests, preserving all test bodies and the Linux build tag. Record the existing replication ARN prefix used by the new R6 fixture in the compatibility inventory; no wire behavior changes.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:44:34 +08:00
Feng Ruohang 055030ea53 fix(http): honor configured request header deadlines
Signed-off-by: Feng Ruohang <rh@vonng.com>
(cherry picked from commit 0d48d32d7e038ae1ea5966f3d7e0cb86780a6311)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:56 +08:00
Feng Ruohang aea3882c95 docs: record R6 integration verification
(cherry picked from commit d38edb2c46182d3a8fa96e040604493d20a4b478)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:56 +08:00
Feng Ruohang 0c61128d23 fix(replication): retry marker purges through persisted MRF
(cherry picked from commit cf381a7151ef25fc95ace5fedcd767fa19410de2)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:55 +08:00
Feng Ruohang 680eac66e4 fix(replication): preserve ordered tag deletions
Persist, send and reconcile empty tag states together with their revision
across COPY, PUT and multipart replication. Advance local tag mutations
under the existing locks and preserve current tags during replication ACK.

Cover signed HTTP, persistent single/multiple pool state, KMS, SSE-C key
rotation, ordering, retry and duplicate requests. Record real Opus plan
consensus, implementation review and local verification evidence.

Signed-off-by: Feng Ruohang <rh@vonng.com>
(cherry picked from commit 115fe8b12329d147adbaf817faa1737392ecbf9b)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:55 +08:00
Feng Ruohang 9f3037e941 Merge pull request #194 from pgsty/codex/r7-replication-content-encoding
fix(replication): preserve normalized replica metadata
2026-09-16 00:31:03 +08:00
Feng Ruohang 5f00f6762f Merge main after R4 validation into R7 candidate
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:21:21 +08:00
Feng Ruohang af2b1794d3 Merge pull request #193 from pgsty/codex/r4-kms-tag-timestamp
fix(replication): preserve SSE-KMS tag timestamps
2026-09-16 00:20:10 +08:00
Feng Ruohang d371f77dcc docs: record final Opus 5 Max R7 implementation review
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:10:31 +08:00
Feng Ruohang 022722a7a7 chore: align R4 contribution notices and record final review
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:08:09 +08:00
Feng Ruohang 03027727d1 fix(replication): preserve SSE-KMS tag timestamps
Carry the already-parsed source tagging timestamp through the KMS options
constructor so replica COPY can apply newer tag updates on explicitly or
automatically encrypted destinations.

Cover all option encryption modes and signed COPY persistence for newer,
stale, duplicate and timestamp-less updates, including bucket defaults.
Preserve the real Opus 5.0/max plan review, consensus and local validation.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:03:52 +08:00
Feng Ruohang 4fcdf37ce6 fix(replication): preserve normalized replica metadata
Restore only the six replication-specific metadata fields after trust
validation, so streaming uploads retain their actual content encoding and
Snowball entries do not inherit ordinary metadata from the outer archive.

Include helper, authenticated PUT/COPY/multipart and Snowball regressions,
plus the R7 investigation, actual Opus 5 consensus and local verification.
The production change is based on PR #187 by Mikhail Khadarenka.

Co-authored-by: Mikhail Khadarenka <chodorenko@gmail.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:00:17 +08:00
Feng Ruohang 9ebe81c1b3 Merge pull request #192 from pgsty/codex/iam-revision-tombstones
fix(iam): retain revocation versions through replay and recovery
2026-09-15 23:14:22 +08:00
Feng Ruohang e5f5c9e7f6 Merge pull request #191 from pgsty/codex/iam-peer-delete-reload
fix(iam): reload committed state on peer deletion notifications
2026-09-15 23:10:38 +08:00
Feng Ruohang 7b4cacc392 chore: align IAM file notices with contribution policy
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang a0dd7dae9b fix(iam): resume site healing after leadership changes
Keep one healing loop per process across replication configuration reloads. Reacquire leadership after a lease is canceled and allow shutdown while waiting, so temporary quorum loss cannot permanently stop revocation propagation.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang 709d50a916 fix(iam): persist revocations across site replay and recovery
Retain source-ordered tombstones and parent grant boundaries across both IAM backends, cache reloads, and deliberate identity recreation. Reconcile deletions through a versioned, bounded replication protocol with restart-aware acknowledgements.

Cover inherited group grants, STS retention, same-key service recreation, absolute expiration, and failures after the durable commit. Document coordinated upgrades and the remaining consistency boundaries.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang dff81f293b chore: align IAM test notice with contribution policy
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:58:52 +08:00
Feng Ruohang f653a6ea03 fix(iam): reload committed state on peer deletion notifications
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:58:52 +08:00
Feng Ruohang 47d239f84f Merge pull request #190 from pgsty/codex/fix-pool-multipart-preconditions
fix(storage): evaluate multipart preconditions across pools
2026-09-15 21:57:26 +08:00
Feng Ruohang e069fe9d92 fix(storage): evaluate multipart preconditions across pools
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 21:13:37 +08:00
Feng Ruohang d848fb52b5 Merge pull request #189 from pgsty/codex/fix-pool-tag-reconciliation
fix(storage): preserve tags during pool reconciliation
2026-09-15 19:38:19 +08:00
Feng Ruohang 3ce8319251 fix(storage): preserve tags during pool reconciliation
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 19:26:58 +08:00
Feng Ruohang 9df0f4abaf Merge pull request #188 from pgsty/codex/remove-access-tiering-final
Remove access-frequency pool tiering and preserve independent multi-pool correctness fixes. Reconcile ordinary addressed-version DELETE across pools, retaining existing quorum and compatibility boundaries.

Verified delivery head: 4d0693cb8c. All 11 final CI checks and three qualified Linux upgrade runs passed. Introduction, retirement decisions and historical unresolved observations are documented.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 15:09:48 +08:00
Feng Ruohang 4d0693cb8c docs: record controlled retirement evidence and qualified upgrade runs
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 14:56:20 +08:00
Feng Ruohang bf59e3f222 docs: record retirement execution and blocked Linux acceptance
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 12:16:45 +08:00
Feng Ruohang 41aa846097 style(storage): satisfy callback selection lint rule
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 11:31:32 +08:00
Feng Ruohang 4093fa0d78 docs: record access tiering introduction and retirement decisions
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 11:27:02 +08:00
Feng Ruohang 13bf126ebd fix(storage): reuse resolved copies for version DELETE callbacks
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 11:27:02 +08:00
Feng Ruohang 1cf529ce8a fix(storage): reconcile ordinary version DELETE across pools
Delete every copy of an explicitly addressed UUID, null version or delete
marker under the pool lock. Preserve retention and replication callbacks,
report unreadable pools and cleanup failures, and keep movement, incoming
replication, expiration and free-version cleanup on their existing paths.

Retain the separately developed general DELETE repair and replace its
access-mover-only coverage with a real interrupted rebalance copy followed
by HTTP deletion. Cover unqualified directory-marker DELETE and document
the existing pool-order-dependent 503 behavior that this makes consistent.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 11:27:01 +08:00
Feng Ruohang 9b76a21675 revert: remove access-frequency ILM tiering (#60)
Reverse the first-parent diff of a3df317ae0,
including the feature branch compatibility and mover follow-up fixes.
Retain the independent multi-pool correctness fixes from #178 and migrate
their shared test fixture away from access-tier code.

Tolerate retired ILM keys and XML, read old v9 statistics while writing v8,
and document migration without moving objects or rewriting their metadata.
Include regression coverage using a historical scanner/writer v9 fixture.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 11:27:01 +08:00
Feng Ruohang 89637554d6 Merge pull request #182 from pgsty/codex/docs-current-state-20260913
docs: align release notes and current component status
2026-09-13 10:50:22 +08:00
Feng Ruohang 2dd1e00da4 docs: align release notes and current component status
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 10:39:57 +08:00
Feng Ruohang 5d955b5b74 Merge pull request #181 from pgsty/codex/deps-release-20260913
Coordinate September dependency releases and replace vulnerable bundled curl
2026-09-13 10:14:57 +08:00
Feng Ruohang f7808a172c deps: pin the coordinated September 13 SILO stack
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 09:54:51 +08:00
Feng Ruohang acc9b514f5 Follow the curl builder rename in delivery verification
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 09:08:38 +08:00
Feng Ruohang 31ba01c5d5 Refresh Go dependencies and build current static curl for both architectures
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 09:01:09 +08:00
Feng Ruohang 48ec10312f Merge pull request #180 from pgsty/codex/issue-77-metadata-convergence
fix(replication): preserve bucket metadata source times and converge deletions
2026-09-12 18:07:36 +08:00
Feng Ruohang 114dc10529 docs: archive issue 77 design, adversarial reviews and acceptance evidence
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 17:52:25 +08:00
Feng Ruohang 461e9a7210 test(replication): satisfy diagnostic regression style checks
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 17:17:18 +08:00
Feng Ruohang fcbb93e895 fix(replication): diagnose unusable metadata without a heal source
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 17:09:34 +08:00
Feng Ruohang 62cf066ff5 fix(replication): recover physical creation time and align policy status
An independent adversarial review of the bucket metadata convergence
work found three defects it had introduced.

GetBucketInfo overwrote the physical creation probe with cached
metadata, which a bucket that never held a configuration legitimately
lacks. The new creation-time requirement then failed every policy, tag,
SSE, quota, versioning and Object Lock write on such a bucket, with no
operator recovery path, and initial synchronization skipped it silently.
Return the physical result unchanged when metadata is not requested, as
ListBuckets already does, recover the time during initial
synchronization, and pass it to MakeBucketHook so peers adopt the same
bucket generation.

Replication status compared parsed policies statement by statement while
heal compares the canonical key. An upgraded peer that stored an
equivalent statement order was therefore reported as mismatched forever,
and heal never had anything to write. Compare the key heal compares;
per-site presence counting is unchanged.

Heal diagnostics shared one log key across four conditions, so a real
peer RPC failure could be deduplicated away by an earlier message, and
they were logged at error level for the normal transient of a peer that
does not have the bucket yet. Give each reason its own key at warning
level, report only a field state that exists and still cannot be
ordered, and diagnose nothing when no site holds a state to propagate.

The recovery test now runs against the real ObjectLayer; the stub it
replaced returned the expected time and hid the defect. The policy
status test uses a statement order the canonical encoder reorders, and
adoption coverage is extended past a real field time.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 14:01:35 +08:00
Feng Ruohang 1ee64a8d89 fix(replication): gate metadata tombstone export during rolling upgrades
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 13:50:55 +08:00
Feng Ruohang 01aaef2b50 fix(replication): converge bucket metadata using deterministic source states
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 13:50:55 +08:00
Feng Ruohang bcc62afe3d fix(replication): apply bucket metadata source times under the metadata lock
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 13:50:55 +08:00
Feng Ruohang 5c57658163 Merge pull request #179 from pgsty/codex/federation-copy-fixes
fix(federation): preserve committed copy times and reject raw SSE-C replicas
2026-09-12 02:00:57 +08:00
Feng Ruohang f175e98c34 fix(federation): bind copy timestamps to committed writes
Return the committed object or part time on federation write responses and
capture it for CopyObject and UploadPartCopy without a follow-up read.
Keep successful writes compatible with targets that do not supply a time.

Reject authenticated raw SSE-C replica CopyObject across deployments before
forwarding. Extend the existing SSE fixtures to verify empty objects,
stored checksums, KMS contexts and multipart sources.

Refs: #169, #168, #171
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 01:44:22 +08:00
Feng Ruohang 12f631b502 Merge pull request #178 from pgsty/codex/multipool-correctness-20260911
fix(storage): reconcile multi-pool writes and conditional deletes
2026-09-11 20:34:58 +08:00
Feng Ruohang ccb676e60c fix(storage): preserve shared tier references during pool cleanup
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 20:23:40 +08:00
Feng Ruohang 51d41345f7 test: record Linux restart and OIDC release acceptance
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 20:05:56 +08:00
Feng Ruohang e59a3d938e fix(storage): serialize and reconcile multi-pool object updates
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 20:05:56 +08:00
Feng Ruohang b32f2d9dd0 Merge pull request #177 from pgsty/codex/release-consolidation-20260911
Fix signed payloads, Object Lock and federated copies; secure AMQP
2026-09-11 17:08:34 +08:00
Feng Ruohang d63c92e393 build(deps): secure AMQP frames and select merged Console
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:56:33 +08:00
Feng Ruohang 68127c5a63 build(deps): embed the consolidated Console source
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:40:42 +08:00
Feng Ruohang c4b5e1cb45 fix(auth): enforce header-only presigned payload checksums
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:24:42 +08:00
Feng Ruohang 9a303f5096 fix(object): preserve retention and federated copy destination state
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:24:42 +08:00
Feng Ruohang 87d8b5967f fix(auth): align signed request and policy condition semantics
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:24:42 +08:00
Feng Ruohang f760046c44 Merge remote-tracking branch 'origin/main' into codex/release-consolidation-20260911
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:13:42 +08:00
Feng Ruohang 93e7ef4bcc Merge pull request #175 from pgsty/codex/upstream-sdk-password-20260910
fix(iam)!: split self-service password permissions and update SDK stack
2026-09-10 17:51:08 +08:00
Feng Ruohang a4229b366f build(deps): align final coordinated SILO source pins
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 17:39:01 +08:00
Feng Ruohang 420340bc14 docs(iam): explain breaking password-policy semantics
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 16:58:21 +08:00
Ayush Sharma e7e87402ed fix: forward the legal hold as an Object Lock header on federated CopyObject
A cross-deployment CopyObject that requests
`x-amz-object-lock-legal-hold: ON` answered 200 while the destination
carried no hold. The resolved value reached the remote as ordinary user
metadata, `X-Amz-Meta-X-Amz-Object-Lock-Legal-Hold`, so nothing applied it.
Retention requested on the same copy survived, which is what made the loss
easy to miss.

The federation branch passes the resolved metadata map straight to
`Core.PutObject` as `PutObjectOptions.UserMetadata`. minio-go's `Header()`
writes the typed lock fields first, then prefixes every UserMetadata key it
does not recognise with `x-amz-meta-`; `supportedHeaders` covers
`x-amz-object-lock-mode` and `x-amz-object-lock-retain-until-date` but not
`x-amz-object-lock-legal-hold`, and `isAmzHeader` does not match it either.
Retention therefore arrives as real headers and the hold does not. The
high-level `validate()` that would have rejected the key never runs, because
`Core.PutObject` goes straight to the low-level PUT.

Carry the hold on the typed `LegalHold` option and forward a cloned map with
the raw key removed. The clone matters twice: typed fields are written before
the UserMetadata loop, so a leftover raw key would add a bogus `x-amz-meta-`
entry beside the correct header, and the proxy's own response and event
metadata are rebuilt from the resolved values rather than the forwarding map,
which no longer carries the hold.

Retention stays in the map deliberately. It already passes through as a
standard header, and moving it to the typed `RetainUntilDate` field would
format with `time.RFC3339` and truncate a retain-until date to whole seconds.

The new test asserts the wire: the remote must receive
`X-Amz-Object-Lock-Legal-Hold` and never the `x-amz-meta-` spelling, and the
destination version must actually store the hold. It fails without the change
with "legal hold forwarded as user metadata [ON]".

Fixes #166

Signed-off-by: Ayush Sharma <72848455+Aeirx@users.noreply.github.com>
2026-09-10 13:15:10 +05:30
Feng Ruohang a164e1dda1 fix(iam): separate password changes and refresh coordinated SDK dependencies
Use ChangeMyPassword for the authenticated user and keep CreateUser for other users. Coordinate silo-pkg b3760f56ec23, mcli fa22b40b4eb7, Console 1b95b6cec652 and upstream minio-go 78bfa91607c2. Add legacy-policy and SDK streaming regressions plus upgrade guidance.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 15:16:21 +08:00
Feng Ruohang 2f61325d4a Merge pull request #174 from pgsty/codex/contributor-orenyomtov-20260910
docs: credit @orenyomtov for the SN-2026-011 report
2026-09-10 11:34:52 +08:00
Feng Ruohang a6145e1d2e docs: credit @orenyomtov for the SN-2026-011 report
Add Oren Yomtov (github.com/orenyomtov) to the contributor wall in
README.md, README_ZH.md, and CONTRIBUTORS.md, and to the Issue reports
table, for the private disclosure of the unsigned-header CopyObject
cross-object read fixed as SN-2026-011 (#173). Community contributor
count 40 -> 41.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 11:24:12 +08:00
Feng Ruohang 5232546690 Merge pull request #173 from pgsty/codex/unsigned-amz-header-copy-20260909
fix(auth): reject unsigned x-amz-* headers to close CopyObject confused-deputy (SN-2026-011)
2026-09-10 10:10:49 +08:00
Feng Ruohang 0c46cb641b docs(security): record SN-2026-011 (unsigned x-amz-* header CopyObject)
Ledger entry for the confused-deputy fix in 123325430: an unsigned
x-amz-copy-source header turned a presigned or signed PUT into a
server-side copy of any object the signing key can read. Reported by
Oren Yomtov; inherited from upstream minio/minio; CVE requested.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 09:57:26 +08:00
Feng Ruohang 1233254309 fix(auth): reject unsigned x-amz-* headers to close CopyObject confused-deputy
A presigned or signed PUT authorized for a single object could be turned
into a server-side copy of any object the signing key can read by adding
an unsigned x-amz-copy-source header, executed as the signer. SigV4
verification only walked the signed-headers list, never the headers that
actually arrived; the meta-header check matched only X-Amz-Meta- and ran
only on the presigned path, so an unsigned x-amz-* header outside the
list was never seen while the router still dispatched the PUT to
CopyObjectHandler.

Reject any x-amz-* request header not covered by the signed headers, on
both the presigned (doesPresignedSignatureMatch) and Authorization-header
(doesSignatureMatch) paths, matching AWS S3. The check tests membership
in the signed set rather than value equality, so a header whose first
value is empty (e.g. {"", "/src/secret"}) cannot slip through.
X-Amz-Content-Sha256 is exempt (payload hash: read from the query for
presigned requests and bound into the string-to-sign for signed ones, so
it is self-protected) and X-Amz-Signature-Age is exempt (an internal
scratch header written after verification, so repeated verification of
the same request stays idempotent). The synthesized X-Amz-Tagging header
in PutObjectTagging is now injected after signature verification.

Tests that previously added x-amz-copy-source and friends after signing
(relying on the vulnerable behavior) now re-sign, mirroring real S3
clients. Adds checkUnsignedHeaders unit cases and TestPresignedVerifyIdempotent.

Reported by Oren Yomtov. Inherited unchanged from upstream minio/minio.
Tracked as SN-2026-011.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 09:57:01 +08:00
Feng Ruohang 8a2fe9b7a0 Merge pull request #164 from pgsty/codex/main-consolidation-20260909
fix: honor TLS defaults and accept valid bucket metadata reloads
2026-09-09 19:49:08 +08:00
Feng Ruohang f26ee6bf0a test: cover metadata reload with equal maximum timestamps
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 19:35:13 +08:00
Feng Ruohang d69c4ccfe4 Merge TLS default key exchange compatibility fixes
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 19:31:35 +08:00
Feng Ruohang 086e619505 Merge accepted bucket metadata reload correction
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 19:31:10 +08:00
Feng Ruohang 48e1846525 fix(tls): honor Go key exchange defaults across transports
Remove the eight explicit curve overrides so Go 1.27 honors tlsmlkem=0
across Server listeners, node links and outbound transports. Remove the
unused shared curve option and add wire-level regression coverage.

Document CA trust and TLS upgrade behavior, retain the investigation
artifacts, and exclude their synthetic routes from the rebrand guard.
The product compatibility baseline remains unchanged.

Validation: focused race tests, HTTP tests, lint, compatibility guard
positive/negative controls, and a fresh Linux build with three isolated
OIDC integration scenarios all pass.

Adversarial review: Claude Code Fable 5.1, max effort.
Final verdict: APPROVE FOR COMMIT.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 18:50:12 +08:00
Feng Ruohang bcc8871b1d Merge pull request #163 from pgsty/fix/issue-158-federated-copy-sse
fix: forward plaintext on federated CopyObject of SSE objects (#158)
2026-09-09 18:02:45 +08:00
Feng Ruohang 25cb3511c9 test: cover multipart SSE sources on federated CopyObject
A multipart SSE-S3 source is encrypted per part, so its logical size is
the sum of the parts' decrypted sizes and the decrypting reader crosses
a part boundary. Copy such a source across the federation to a plain
and to an SSE-S3 destination and check the destination plaintext and
single-encryption size.

The test router registers routes in endpoint order and the plain
PutObject route has no query matcher, so the multipart endpoints are
listed first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FodsDpa6VkghaeRE6WjmEe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 17:51:32 +08:00
Feng Ruohang cfefc049c1 fix: send the SSE-KMS context as a JSON object on federated copies
putOptsFromReq handed the parsed kms.Context straight to
encrypt.NewSSEKMS. kms.Context implements encoding.TextMarshaler, so the
SDK serialized it as a JSON string, and a request without a context
still produced one because the nil Context is a typed nil inside the
interface value and marshals to "{}". The receiving ParseHTTP rejects
both forms, so every federated CopyObject to an SSE-KMS destination
failed with InvalidArgument once the forwarded stream was correct.

Pass a plain map, or nothing when no context was requested, and cover
SSE-KMS destinations with and without an explicit context.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FodsDpa6VkghaeRE6WjmEe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 17:51:32 +08:00
Feng Ruohang af56d17630 fix: forward plaintext on federated CopyObject of SSE objects (#158)
The legacy etcd bucket-federation branch of CopyObjectHandler reads its
source through getObjectNInfo, which yields the decrypted and
decompressed bytes, but it also ran the destination encryption locally
and then forwarded that stream to the remote PutObject with the source's
stored size and the destination SSE option. SSE to plain and plain to
SSE therefore failed on a Content-Length mismatch, while SSE to SSE
matched by coincidence: the remote encrypted the ciphertext a second
time and stored an unreadable object, and a destination GET returned
the inner ciphertext with HTTP 200.

The remote write owns the destination's storage transformations, so
hand it the logical bytes at their logical size and let it encrypt
exactly once. Compression was already excluded on this branch; apply
the same rule to encryption, size the forwarded reader by actualSize,
and declare that size on the forwarded PutObject.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FodsDpa6VkghaeRE6WjmEe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 17:51:32 +08:00
Feng Ruohang d1105bbb3d Merge pull request #162 from pgsty/codex/replication-reliability-20260909
fix(replication): complete purges, expose MRF drops, and cancel resyncs reliably
2026-09-09 14:55:35 +08:00
Feng Ruohang 66fe61ff65 test: reuse replication compatibility fixtures
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:44:02 +08:00
Feng Ruohang a1141a43f2 style: format replication regression fixture
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:24 +08:00
Feng Ruohang 702f113f51 fix(replication): scope resync cancellation and drain worker lifecycle
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:02 +08:00
Feng Ruohang 63aace4099 fix(replication): expose bounded MRF queue drops
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:02 +08:00
Feng Ruohang 9d7094b770 fix(replication): complete single-object delete marker purges
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:02 +08:00
Feng Ruohang 450dcb8484 Merge pull request #161 from pgsty/codex/server-dependency-refresh-20260909
build: refresh maintained SILO stack dependencies
2026-09-09 12:04:16 +08:00
Feng Ruohang 4074d00b96 build: refresh maintained SILO stack dependencies
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 11:49:13 +08:00
Feng Ruohang 75ba0ce402 fix: drop the broken bucket-metadata reload publication guard (issue #105 T3)
The #105 T3 change (PR #156) tried to keep the resident metadata cache
monotonic by guarding peer-reload publication on lastUpdate(). But
lastUpdate() is the max of per-config timestamps and cannot order whole
records: a node caching {policy@20, CORS@10} that receives a newer CORS@15
still has lastUpdate()==20, so the guard rejects the legitimately-newer
record and the periodic refresh (same comparator) cannot repair it. A
paused reload could also resurrect deleted resident state.

Per the maintainer decision, revert the reload publication to its original
unconditional (acceptable-until-refresh) behavior:
- remove setReloaded and restore the plain Set plus notification/target
  registry updates in LoadBucketMetadataHandler;
- restore refreshBucketsMetadataLoop's own lastUpdate() staleness check and
  globalEventNotifier.set / globalBucketTargetSys.set publication;
- restore the unconditional GetConfig cache-miss publication;
- document the known freshness limitation at the reload site (the periodic
  refresh is best-effort and cannot repair an equal-maximum-timestamp
  divergence).

The T1 lifecycle merge-under-lock (UpdateExpiryLCConfig) and both T2 fixes
(DeleteBucket takes metadata.lock before deleting; saveMetadata and
loadBucketMetadataParseUnderLock recheck physical bucket existence) are
kept fully intact.

Tests:
- drop the T3 reproductions (overlapping-reload resident-cache test and the
  peer-reload-preserves-current-targets publication test);
- add lockBucketMetadataAcquireHook, a nil-in-production atomic test hook in
  the shared metadata.lock path, so tests can deterministically observe a
  caller (notably DeleteBucket, whose lock is taken through its
  erasureServerPools receiver and is invisible to an injected object layer)
  reaching the lock;
- rewrite the T2 delete-race ghost test to hold metadata.lock MID-SAVE (past
  saveMetadata's existence recheck) and synchronize on the delete's actual
  lock attempt via the hook, so it isolates the lock-before-delete fix:
  removing only DeleteBucket's metadata.lock (recheck kept) now fails it;
- rewrite the cancellation test to observe the delete's actual lock attempt,
  then cancel and await its error while still holding the lock, so a
  scheduling-delayed delete stopped by the canceled context can no longer
  pass on a broken tree.

Refs #105. Follows #156.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 16:13:54 +08:00
Feng Ruohang da142327f2 Merge pull request #159 from pgsty/fix/issue-99-100-followups
fix: repair federated CopyObject checksum edge cases (empty body, inherited, multipart-suffix)
2026-09-08 15:58:22 +08:00
Feng Ruohang a3df317ae0 Merge pull request #60 from mrjavadseydi/feat/access-based-ilm
[ILM] Relocate hot objects across server pools by GET frequency
2026-09-08 15:42:15 +08:00
Feng Ruohang 2c50d11f72 fix: correct federated CopyObject checksum edge cases (#99 follow-ups)
Three residual checksum defects in the legacy etcd federation branch of
CopyObjectHandler, found by post-merge review of #157.

1. Empty-source 500 regression. A checksum-less object gains the S3 default
   CRC-64NVME (WantServerSideChecksumType is set), but minio-go streams no
   trailing checksum for a 0-byte body (contentLength == 0), so the remote
   computed none, federatedChecksumValue was empty, hash.NewChecksumWithType
   returned nil, and the handler returned 500 -- so every empty-object
   federated copy failed. For a 0-byte source, forward the empty-content digest
   as an ordinary checksum request header instead, so the remote validates,
   persists and returns it, matching the local path (e.g. CRC32 "AAAAAA==").

2. Inherited full-object checksum dropped. When the source already carries a
   full-object checksum, the local path sets dstOpts.WantChecksum, not
   WantServerSideChecksumType (only multipart-composite sources are promoted).
   The federated branch inspected only WantServerSideChecksumType, so a
   checksum-bearing source's checksum was silently discarded on a federated
   copy that requested no algorithm. Forward WantChecksum.Encoded (always a
   plain digest) as a checksum header so the remote validates and persists it,
   and bind the returned value, matching local persistence.

3. Multipart-suffixed remote value accepted. The bind accepted a value like
   "NSRBwg==-0": NewChecksumWithType parses the "-N" as ChecksumMultipart with
   WantParts 0 and the length-only validator passes, so the destination was
   returned as COMPOSITE. A single forwarded PutObject must yield a full-object
   digest, so reject a multipart-marked parsed value in addition to the
   existing nil (missing/malformed) rejection.

A forwarded checksum request header is stripped from objInfo.UserDefined so it
is not mistaken for object metadata.

Out of scope: the SSE federated-copy corruption (srcInfo.Reader/Size mismatch
for encrypted sources) predates this work and is filed separately.

New federated regressions cover empty source with requested and default
checksum (200 + correct value + persisted), an inherited full-object checksum
preserved without a requested algorithm, and a multipart-suffixed remote value
rejected. Red/green verified for each against the merged code.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:38:11 +08:00
Feng Ruohang 5ac33e1583 test: make access move failure recovery deterministic
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:28:33 +08:00
Feng Ruohang d57c4e8407 Merge remote-tracking branch 'origin/main' into codex/access-tiering-ci-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:08:34 +08:00
Feng Ruohang 374de0fa32 fix: make access tier moves preserve versions and isolate writes
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:08:34 +08:00
Feng Ruohang 80e23dc9f2 Merge pull request #156 from pgsty/codex/bucket-metadata-merge-20260908
fix: preserve bucket metadata across concurrent updates and reloads
2026-09-08 15:03:52 +08:00
Feng Ruohang 2cc0e3c6ed test: construct notification fixtures with the event ARN type
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:52:48 +08:00
Feng Ruohang f9da3b919d fix: order bucket deletion and metadata publication safely
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:43:15 +08:00
Feng Ruohang 1309853f57 Merge remote-tracking branch 'origin/main' into codex/access-tiering-ci-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:34:51 +08:00
Feng Ruohang 39b8e6c30a Merge remote-tracking branch 'origin/main' into codex/bucket-metadata-merge-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:24:22 +08:00
Feng Ruohang 9a6e1477f4 ci: record access tiering compatibility identifiers
Record the ten documented MINIO_ILM_ACCESS settings, the internal object
metadata stamp, and the tracker storage-path suffix introduced by this PR.
The guard places the /ilm/access string in its routes set, but the value
is a component of the tracker object prefix, not a public HTTP endpoint.

Keep all existing compatibility entries. The guard, delivery rebrand check,
and Docker entrypoint compatibility tests pass with the refreshed manifest.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 13:18:50 +08:00
Feng Ruohang 49375ed2d3 Merge pull request #157 from pgsty/codex/federated-copy-merge-20260908
fix: preserve metadata and checksums in federated object copies
2026-09-08 13:16:30 +08:00
Feng Ruohang f817b5261c Merge pull request #151 from nikitapogromsky/fix/idempotent-add-system-target
logger: make AddSystemTarget idempotent
2026-09-08 13:14:30 +08:00
Feng Ruohang 885bd2c20a fix: reject missing remote checksums on federated copies
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 13:05:45 +08:00
Feng Ruohang 89b75e7913 Merge branch 'main' into feat/access-based-ilm 2026-09-08 13:04:25 +08:00
Feng Ruohang 079ebb1926 fix: serialize logger initialization and publish the console target safely
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:58:21 +08:00
Feng Ruohang 04c29aac11 Merge branch 'audit/issue-105-bucketmeta-races' into codex/bucket-metadata-merge-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:55:51 +08:00
Feng Ruohang 95e7a190b3 Merge branch 'fix/issue-99-100-federated-copyobject' into codex/federated-copy-merge-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:55:51 +08:00
Feng Ruohang 8e2392e48f Merge remote-tracking branch 'origin/main' into codex/logger-idempotency-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:55:50 +08:00
Feng Ruohang 3abe0d95a5 Merge pull request #149 from pgsty/codex/embedded-console-compat
fix: restore embedded Console proxy and WebSocket configuration
2026-09-08 10:09:21 +08:00
Feng Ruohang 9c6c9805de fix: select the validated Console mainline for embedding
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 10:00:45 +08:00
nikitapogromsky 5cb900bfad logger: make AddSystemTarget idempotent
Subscribe re-registered the console target on every console-log subscription, producing duplicate minio_logger_webhook_* series on each /minio/metrics/v3 scrape. Fixes #150

Signed-off-by: nikitapogromsky <129324283+nikitapogromsky@users.noreply.github.com>
2026-09-07 12:03:39 +03:00
Feng Ruohang 0af5d22286 fix: restore embedded Console proxy and WebSocket configuration
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 13:36:24 +08:00
Feng Ruohang f1687f402b fix: scope rebrand checks to the retired repository
Match the exact retired repository while preserving links to distinct repositories and historical issue titles in contributor credits. Continue rejecting live links to the retired repository in those credits.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 11:49:23 +08:00
Feng Ruohang ce606df2c4 fix: align contribution tooling with SILO AGPL policy
Clarify SILO contribution ownership and preserve prior copyright notices. Consolidate issue templates and route the legacy credits command through the maintained generator.

Validation: make rebrand-guard; bash -n update-credits.sh; regenerated credits match CREDITS; template and link checks.
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 11:33:21 +08:00
Feng Ruohang fd44dc4e9b docs: include the latest merged community contribution
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 11:01:04 +08:00
Feng Ruohang 479745e764 docs: credit contributors across the SILO repositories
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:58:47 +08:00
Feng Ruohang cc1c54475f Merge pull request #146 from pgsty/fix/issue-108-console-loopback-tls
fix: keep embedded Console login working over loopback TLS (#108)
2026-09-07 10:53:45 +08:00
Feng Ruohang e7654d470c fix: close residual bucket-metadata races (issue #105 audit)
Audit of the three deferred #105 follow-ups. Each reproduces with a
deterministic red test in cmd/bucket-metadata-race_test.go, and each fix is
the minimal change that turns its test green while preserving the
<bucket>.lck -> metadata.lock -> .metadata.bin lock order established by #103.

1. Lifecycle expiry merge lost update (persistent). PeerBucketLCConfigHandler
   and healBucketILMExpiry read the current lifecycle with an unlocked
   GetConfigFromDisk, merged the replicated expiry rules with the local
   transition rules, then wrote the pre-computed blob via Update. Any lifecycle
   transition change committed between the merge read and the merge write was
   silently lost. New BucketMetadataSys.UpdateExpiryLCConfig performs the read,
   merge, and save under one metadata.lock; mergeExpiryWithLCConfig now takes
   the locked snapshot and validates object-lock retention from it instead of
   re-reading (avoids a re-entrant metadata load under the lock).

2. DeleteBucket ghost .metadata.bin (persistent). DeleteBucket took only
   <bucket>.lck while config writers take only metadata.lock, so a writer that
   was mid-save could re-create .metadata.bin after the prefix purge. The purge
   now runs under metadata.lock, with a best-effort unlocked fallback so a
   delete is never blocked from completing.

3. Overlapping peer reloads publishing a stale resident cache (freshness only;
   the persisted record stays correct). LoadBucketMetadataHandler and the
   GetConfig cache-miss path published with an unconditional Set, so a reload
   that read an older revision could overwrite a newer resident record until the
   next refresh. New BucketMetadataSys.setReloaded (and a matching GetConfig
   guard) refuses to regress a newer resident record, mirroring
   refreshBucketsMetadataLoop.

Verification: go build -tags kqueue,dev ./...; go vet ./cmd; gofmt clean;
rebrand-guard baseline unchanged; go test -tags kqueue,dev ./cmd (207s) green;
new tests plus the #103 metadata suite green under -race.

Refs #105. Parent #102. Foundation #103.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:33:11 +08:00
Feng Ruohang 4b25f7e819 fix: return and persist checksum on federated CopyObject (#99)
The legacy etcd federation branch of CopyObjectHandler forwards the copied
bytes with minio-go Core.PutObject but never asked the remote for a checksum
and discarded any it returned, so a cross-deployment whole-object copy that
requested a checksum returned 200 with an empty checksum, and a checksum-less
source did not gain the S3 default CRC-64NVME that the local path assigns. The
request was neither honored nor rejected. This is the whole-object counterpart
of #72, which repaired the same class of defect for federated UploadPartCopy.

When a server-side checksum is wanted -- explicitly requested, inherited from a
multipart source, or the CRC-64NVME default for a checksum-less object, all
already captured in dstOpts.WantServerSideChecksumType -- the forwarded
PutObject now streams a trailing checksum of that type, so the remote computes
and persists it and echoes it in the response. The value the remote reports for
that exact write is bound into objInfo.Checksum, matching how the local
CopyObject path carries checksums into the CopyObjectResult. Reading the value
from the same UploadInfo that produced the ETag keeps the pair bound to one
write.

Only the requested algorithm is returned; a malformed or absent remote value
leaves objInfo.Checksum unset, so an ordinary copy that wanted no checksum
still returns none. Two small mapping helpers convert between the server's
hash.ChecksumType and the minio-go request type and response field.

New end-to-end tests drive the real federation branch through
getRemoteInstanceClient and minio-go into a second in-process deployment and
assert that CRC32/CRC32C/SHA256/CRC64NVME and the no-algorithm default are all
returned in the CopyObjectResult and persisted on the destination, and that a
requested algorithm never leaks other algorithms into the response.

Fixes #99

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:32:59 +08:00
Feng Ruohang 711b092f86 fix: keep embedded Console login working over loopback TLS (#108)
The embedded Console reaches the S3/STS API at https://127.0.0.1:<port>
(minioConfigToConsoleFeatures), a loopback endpoint whose TLS certificate is
not expected to carry a 127.0.0.1 SAN. silo-console v2.3.x began verifying
every outbound TLS peer, so the Console's STS AssumeRole handshake to that
loopback endpoint now fails certificate validation and BOTH local and LDAP
logins fail with a generic "invalid login". The failure happens in the Console
HTTP client before any request reaches a server auth/STS/LDAP handler, so no
server-side auth error is logged, matching the report.

Restore the documented loopback bypass by opting the embedded Console into its
endpoint-scoped CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY switch whenever the server
falls back to the 127.0.0.1 endpoint under TLS. The exemption is scoped to that
single loopback origin inside Console; every other HTTPS peer (IdP, Prometheus,
webhooks) stays verified, preserving the v2.3.x hardening. An explicitly
configured endpoint is reached under its own verified name and is never
exempted. initConsoleServer unsets CONSOLE_* before re-deriving them, so the
switch cannot be supplied by the operator on the embedded path; the server must
assert it.

Fixes #108

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:32:59 +08:00
Feng Ruohang 2bc103b80c fix: strip all reserved metadata on federated CopyObject (#100)
The legacy etcd federation branch of CopyObjectHandler forwards the copied
source metadata to the remote deployment with minio-go Core.PutObject, after
removing only two reserved keys (compression and actual-size). Every small
object is stored inline, so its stored metadata also carries
X-Minio-Internal-inline-data; the remote's setRequestLimitMiddleware rejects
any request bearing a reserved-prefix header (containsReservedMetadata), so
the forwarded write failed with 400 InvalidArgument "Your metadata headers
are not supported." for the default COPY metadata directive.

A plain federated PutObject must not carry any internal storage metadata, so
strip the whole reserved-prefix class before forwarding instead of an
enumerated subset. Enumerating a third key would only defer the next leak:
besides inline-data, replication bookkeeping (replica/replication status and
timestamps) is added to the same map earlier in the handler and would be
rejected just the same. None of these keys is required by the remote for a
correct plain PutObject; they are internal storage details the remote sets
for itself. The stripping uses stringsHasPrefixFold, matching the remote's
own case-insensitive detection. Ordinary user metadata (x-amz-meta-*) is
untouched and still copied.

The pre-existing UUID ETag on the federated write (no Content-MD5 is sent) is
out of scope and left unchanged, as recorded in the issue.

A new end-to-end test drives the real federation branch through
getRemoteInstanceClient and minio-go into a second in-process deployment,
copying an inline source with the default COPY directive. It asserts the copy
now succeeds, that no reserved-prefix header reaches the remote on any
forwarded request, and that copied user metadata survives.

Fixes #100

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:32:59 +08:00
Feng Ruohang ad873c7357 Merge pull request #132 from mrjavadseydi/fix/issue-106-bucket-quota-metrics
fix: report effective bucket quotas in metrics
2026-09-07 09:54:01 +08:00
Feng Ruohang 0af0907eff Merge pull request #134 from pgsty/fix/issue-120-ssec-replica-retransmit
fix: retransmit and re-order Object Lock for SSE-C replicas (single erasure set)
2026-09-07 00:24:45 +08:00
Feng Ruohang e27ba2bc14 Merge pull request #131 from pgsty/fix/issue-117-lock-resend-compare
fix: stop re-replicating an object whose retention was removed
2026-09-07 00:22:21 +08:00
Feng Ruohang 236e163c0b fix: repair an undecodable SSE-C replica on retransmit
PutObjectHandler's precondition callback ran DecryptObjectInfo on the
stored object before checkPreconditionsPUT, so an authenticated raw
SSE-C replica overwrite was rejected when the stored version could not
decrypt. A replica a pre-fix destination (issue #109) left as
compress(ciphertext) or a re-encrypted body has an invalid decrypted
length, so DecryptObjectInfo returned errObjectTampered and the
retransmission that repairs it never ran -- the version stayed damaged
through resync. #134's raw-replica exemption only covered the
version/ETag duplicate check inside checkPreconditionsPUT, one step too
late.

Skip the stored object's decryption precondition only for a PURE raw
SSE-C replica overwrite (a trusted SSE-C replica write with no public
precondition), keyed on the incoming request's restored SSE-C metadata,
the same predicate checkPreconditionsPUT uses. Such a write fully
replaces the object, so requiring the damaged stored version to decrypt
is both wrong and unnecessary. A conditional request keeps the check:
DecryptObjectInfo also normalizes the stored sealed ETag to the
client-visible one, and If-Match/If-None-Match must compare against
that, not the sealed ETag -- skipping it for every replica inverted both
conditions. Ordinary writes and non-SSE-C replicas are unchanged.

Adds red/green regressions: a raw retransmit over a version staged as an
undecodable body returns 500 XMinioObjectTampered before this change and
200 with full customer-key recovery after; and a conditional replica PUT
(If-Match / If-None-Match) on the client-visible ETag is honoured rather
than inverted. Fixes the single-PUT compression-damage recovery gap in

Signed-off-by: Feng Ruohang <rh@vonng.com>
#120.
2026-09-07 00:11:04 +08:00
Feng Ruohang 7220210e8d test: reconcile #134 fixtures with #119 and refresh the rebrand baseline
Rebased onto current main. #119 made PutObjectPart derive an encrypted
part's plaintext length and reject a part that cannot be a valid sio
stream, so TestReplicaLockReconcileNullVersion's completeNullMPU fixture
(a 4-byte plaintext part under SSE-C metadata) no longer stores; build it
with sio.Encrypt like the other encrypted-part fixtures. Also regenerate
the rebrand-guard baseline for the replication SSE header the retransmit
path reintroduces (headers 84 -> 85). Mechanical integration only.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 34cbca97ea docs: point the multi-pool lock follow-up at pgsty/silo#133
Fill the tracked-issue number into the scope comments of the single
erasure set Object Lock reconcile added for SSE-C replica retransmit.
No behaviour change.

Refs pgsty/silo#120
Refs pgsty/silo#133

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 109d824e5f fix: retransmit and re-order Object Lock for SSE-C replicas (single erasure set)
Issue #120 routes an existing SSE-C replica through PutObjectHandler and
NewMultipartUploadHandler. On main those handlers assigned the incoming retention
and legal hold directly, without the source-timestamp ordering #111 added to
CopyObjectHandler and without persisting the ordering timestamps, so in
active-active replication a retransmit carrying an older value could overwrite a
destination version's newer lock state.

Share #111's ordering decision as applyReplicatedObjectLock in
cmd/bucket-object-lock.go and call it from CopyObject, PUT and multipart
initiation. A request that is not an actual trusted replica keeps ordinary write
semantics (a validated value is applied and stamped now); only a real replica
update is ordered against the stored version, so a marker-only peer write no
longer drops a validated hold or default retention. CopyObject keeps its SSE-C
key-rotation encMetadata reconciliation inline. putReplicationOpts now emits a
stored retention ordering timestamp even when the value keys are absent, so a
removal recorded on the retransmit PUT path still replicates onward. replicateAll
marks Failed and carries the error when putReplicationOpts fails.

The handler decision is made against the version as it stands then, which a
concurrent lock update can outrun before the write commits, and for multipart
across the whole initiation-to-completion span. Close that window under the
object write lock the receiving erasure set holds: a trusted SSE-C replica full
write sets ObjectOptions.ReplicaLockReconcile, and erasureObjects.PutObject and
CompleteMultipartUpload re-run the ordering (reconcileStoredObjectLock, which
orders retention and legal hold independently by their reserved timestamps)
against the destination version read on that set before committing. Persisted
upload metadata records the null version as an empty VersionID, so completion
looks that up as the null version rather than the latest. The reconcile runs only
against an existing version; a not-found destination keeps the write's own
accepted lock, including a pre-upgrade upload that persisted values without
ordering timestamps, and a non-not-found read error fails the write. Scoped to
the SSE-C paths this issue enables; CopyObject is left as #111 wrote it.

Scope: this orders Object Lock against the destination version under the write
lock and is correct for a single erasure set. A multi-pool deployment -- where a
version can have duplicate copies across pools, object ModTime ties do not track
per-field lock timestamps, and the object namespace lock is per-pool -- needs a
cross-pool lock-safe reconcile and is deliberately out of scope here, tracked in
pgsty/silo#TBD-multipool-lock.

Tests: TestAPISSECReplicaRetransmitObjectLockOrdering and its multipart sibling;
TestAPIReplicaMultipartNewerHoldSurvivesCompletion and
TestReplicaPutObjectLockReconcileUnderWriteLock (a hold or retention reaching the
version after the handler decision, or after multipart initiation, survives the
commit; a pre-upgrade upload on an absent version keeps its lock);
TestReplicaLockReconcileNullVersion (a null-version completion reconciles the null
version, not a coexisting UUID version, and an absent null version keeps its
accepted lock); TestAPIReplicaMarkerOnlyAppliesObjectLock; TestReplicaStoredLock;
the timestamp-only putReplicationOpts round trip; and the retransmit, exemption
and target-head tests. The #111 CopyObject replica suite and the existing #120
suite stay green, as do the PUT/multipart handler and object-layer regression
suites. Compatibility: the shared helper preserves #111's CopyObject behavior; a
non-replica PUT or multipart initiation that sets Object Lock now also stamps the
reserved ordering timestamp, matching CopyObject since #111; only trusted SSE-C
replica writes take the in-lock reconcile.

Refs pgsty/silo#120

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 87746913fc fix: retransmit existing SSE-C replicas instead of metadata-copying them
The replication sender's target HEAD carries no SSE-C customer key, so for
an SSE-C object the target answers 400 and replicateAll fell into a
metadata-only CopyObject that fails on any non-empty SSE-C object (the
undecryptable source checksum makes the target recompute one and rewrite
the data with a plaintext-sized reader). Once a non-empty SSE-C replica
existed, tag, retention and legal-hold changes never reached it, a heal
never retransmitted, and a resync neither repaired the replica nor
counted it correctly. Forcing a full retransmit alone was not enough:
checkPreconditionsPUT rejects a write whose PreserveETag and VersionID
match the stored version, only the single-part sealed ETag is truncated
before that comparison, so a multipart SSE-C retransmit answered 412,
which the sender turns into success. Inherited from upstream ad04afe38.

Select replicateAll when the SSE-C HEAD cannot answer (the two previous
assignments were dead: rAction still forced the metadata path), exempt an
authenticated replica write that carries an SSE-C seal from the duplicate
version and ETag rejection (the predicate is the incoming write's
restored SSE-C metadata, not what the destination holds), and send the
internal replication marker on the resync accounting HEAD for SSE-C
objects so a peer answers with the replica metadata instead of 400.

Tests: TestAPISSECReplicaRetransmitOverExistingVersion (multipart replica
initiation over the same version and ETag answered 412 on main, now 200
with parts sent and plaintext readback; single-part and zero-byte writes
unchanged), TestAPISSECReplicaWriteExemptionIsKeyedOnTheIncomingWrite
(plaintext replica over an SSE-C version still 412; SSE-C replica over a
plaintext version exempted and readable), and
TestAPISSECReplicationTargetHead (keyless HEAD 400, missing key 404,
marked HEAD 200 with metadata, metadata CopyObject ExcessData on a
non-empty object) on ErasureSD and Erasure. Compatibility: every update
of an SSE-C object now retransmits its bytes; a peer that rejects the
internal marker fails the accounting HEAD as before; the #109 destination
fix must be deployed first or a retransmitted replica is transformed
again.

Fixes pgsty/silo#120

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 7935c84f9a fix: recognize timestamp-only retention-removal tombstone in resend compare
retentionRemovedAtSource only recognized representation (1) of a removed
retention: the object lock key present with an empty value. But a removal
that arrived by replication persists representation (2): restoreRetention
(and the receiver's replica update path) writes only the retention ordering
timestamp when the mode is empty, leaving the mode and retain-until-date keys
absent. For that shape the helper returned false, so replicationActionForTarget
skipped the GetObjectRetention confirmation and let getReplicationAction's
replicateNone stand, silently dropping a needed removal when the destination
HEAD hides retention behind a permission-filtered credential.

Recognize representation (2) as well: a present retention ordering timestamp
with the mode value absent or empty is a removal. A present timestamp paired
with a non-empty mode is a retention that was set, not removed, and still
returns false.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:06:47 +08:00
Feng Ruohang c185635b43 fix: treat empty object lock values as absent when comparing
getReplicationAction builds its source map from oi1.UserDefined, where a
removed retention is a present key with an empty value, and its target map
from the destination's HEAD headers, which can never carry those keys because
setObjectHeaders skips empty lock values and FilterObjectLockMetadata drops
both keys when the mode is invalid. The comparison then always reports a
difference, the replicateNone fast path is dead for such versions, and an
otherwise matching version re-copies its metadata on every evaluation.

Skip an entry whose value is empty and whose key is x-amz-object-lock-mode or
x-amz-object-lock-retain-until-date, case-insensitively, in both comparison
loops, using the joined value on the target side. Normalizing only the source
would regress the case where both sides hold the empty pair.

HEAD also omits a real retention from a credential without
s3:GetObjectRetention, which the documented target policy does not grant, so
that normalization alone would read a destination hiding a retention as in
sync and drop the removal. replicationActionForTarget therefore confirms with
the destination before skipping the resend: only an explicit answer, no
retention on the version, clears it. Everything else keeps today's metadata
resend, including a denied or unreachable destination, a mode the SDK does not
recognize, and InvalidRequest, which names a bucket without Object Lock but is
also what a destination answers when its own read of that configuration fails.
The null version an existing object resync excludes is never reopened.

Tests: TestGetReplicationActionEmptyObjectLockValues (eight cases, red on 2
and 3 before this change), TestRetentionRemovedAtSource,
TestTargetRetentionConfirmedAbsent,
TestReplicationActionForTargetRetentionRemoval,
TestReplicationActionForTargetNullVersionResync and
TestEmptyRetentionValuesAreOmittedFromObjectResponseHeaders.
Compatibility: sender side only, no wire or storage change, so a fixed source
converges against any destination version.

Fixes pgsty/silo#117

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-07 00:06:47 +08:00
Feng Ruohang c201148738 Merge pull request #145 from pgsty/fix/issue-10-conditional-delete
fix: support conditional DeleteObject (If-Match) with atomic precondition
2026-09-06 23:58:42 +08:00
Feng Ruohang 53adb21c52 Merge pull request #142 from pgsty/fix/issue-141-resync-dispatch-scope
fix: scope resync worker dispatch to the target being resynced
2026-09-06 23:57:17 +08:00
Feng Ruohang 58b0ee36ca Merge pull request #140 from pgsty/fix/issue-139-resync-classification
fix: count resync success by replication outcome, not target existence
2026-09-06 23:56:57 +08:00
Feng Ruohang 8a1f594add Merge pull request #138 from pgsty/fix/issue-136-resync-counter-flush
fix: persist an honest resync terminal status about object counts
2026-09-06 23:55:59 +08:00
Feng Ruohang 5b9959617e Merge pull request #130 from pgsty/docs/issue-116-startup-readiness
docs: describe the startup readiness window of the health probes
2026-09-06 23:55:37 +08:00
Feng Ruohang da19d91b64 Merge pull request #143 from pgsty/fix/issue-107-chunked-checksum
fix: honor a header-delivered checksum advertised as a chunked trailer
2026-09-06 23:55:33 +08:00
Feng Ruohang 40bee4b7ba fix(delete): honor If-Match precondition on DeleteObject (#10)
DeleteObject ignored the If-Match request header and always deleted the
object (204). AWS S3 conditional deletes require that when If-Match is
provided and does not match the object's current ETag, the delete is
refused with 412 Precondition Failed and the object is left intact.

The precondition is evaluated in erasureServerPools.DeleteObject, while
the server-pool delete lock is held, before the delete-marker short-circuit
and before any version is removed. It runs against the version that will
actually be deleted: pinfo.ObjInfo for a normal delete, or the specifically
addressed version (read under the held lock) for a version-scoped delete,
since getPoolInfoExistingWithOpts strips VersionID. The check is a pure
function (no ResponseWriter writes) and returns PreConditionFailed, which
toAPIError maps to 412; CheckPrecondFn is cleared before lower layers run
so the precondition is evaluated exactly once.

Semantics:
- If-Match mismatch on a live object -> 412, object preserved.
- If-Match "*" requires a live object; a delete-marker-latest -> 412, and
  an explicitly addressed delete-marker version -> 412 (getObjectInfo
  returns the marker with MethodNotAllowed; the marker is the precondition
  target, not a 405).
- SSE-C/SSE-KMS: compared against the public ETag derived without the
  customer key, so a satisfiable condition is never falsely rejected.
- Explicit versionId -> evaluated against that version; a missing addressed
  version -> NoSuchVersion whether or not the key exists; a missing object
  (no versionId) -> NoSuchKey; no If-Match -> unchanged (including the
  unconditional version-scoped delete's error behavior).

Scope: atomicity is guaranteed for a single erasure set (the default
deployment). Multi-pool conditional-delete atomicity (concurrent writers
across pools, cross-pool version selection) is tracked as a follow-up.

Tests: pure-helper unit test (delete marker, "*", SSE-C without key);
object-layer tests (unversioned match/mismatch/missing, versioned
delete-marker-latest and addressed delete-marker version, explicit-version
selection, missing version on present and absent keys, read-quorum loss);
handler tests (412/204/wildcard/404) across both backends, with red/green
demonstrated per guard.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 23:05:27 +08:00
Feng Ruohang 6a9b5d6763 fix: honor header checksum when x-amz-trailer is advertised on non-trailer chunked PUT (#107)
The AWS Java SDK v2, with chunked encoding enabled (its default), sends a
PutObject as a non-trailer signed aws-chunked stream
(x-amz-content-sha256: STREAMING-AWS4-HMAC-SHA256-PAYLOAD). When a checksum
algorithm is set it puts the precomputed value in the x-amz-checksum-crc32
header, yet still advertises the checksum in x-amz-trailer even though no
trailer chunk is ever sent.

GetContentChecksum treated any x-amz-trailer-advertised checksum as trailing
with an empty value, deferring it to a trailer. For the non-trailer auth type
the handler sets req.Trailer = nil, so at EOF the hash.Reader looked the value
up in a nil trailer, got "", and returned XAmzContentChecksumMismatch (HTTP
400) even though the correct value sat in the request header. Real S3 accepts
the request, and disabling chunked encoding removed the trailer advertisement,
matching the reported symptom.

Honor the header value directly when a trailer-advertised checksum is already
present in the request headers; fall back to trailing delivery only when the
header is absent. When the header carries the checksum but it does not parse,
reject the request with ErrInvalidChecksum instead of falling through to a
no-validation path, so a malformed client-supplied checksum is never silently
dropped. This also restores the checksum echo on the response and the stored
value, while keeping genuine trailer uploads and wrong-checksum rejection
intact.

Fixes #107.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 22:48:29 +08:00
Feng Ruohang 0720ed477e fix(replication): scope resync dispatch to the target being resynced
The resync worker pool runs for a single target (opts.arn), but the dispatch
loop admitted any object whose ExistingObjResync.mustResync() was true for ANY
target. On a bucket with per-target rules (A and B), a resync of A would pull
in objects that only qualify for B - even with a single active resync, since
qualification is any-target. After the outcome-based classification (previous
change) such a cross-target object leaves A absent from its per-object result
and is counted as an A failure - an object A was never responsible for.

Scope admission to the resync's own target: dispatch an object only if it must
resync for opts.arn specifically (mustResyncTarget), via a small pure helper
objectNeedsResyncForARN. Only opts.arn carries this resync's ResetID, and that
reset is already folded into its per-target decision, so the per-target check
both scopes dispatch and honors the reset. Each target has its own resyncBucket,
so no cross-target object is dropped - it is handled by that target's resync.
The classifier's absent-ARN failure path is now unreachable for normally
dispatched objects and remains only as defense-in-depth (e.g. a config/lock
error before replication is attempted).

Delete-marker/version-purge handling, the null-version exclusion, the
finalization ordering and the outcome-based classification are unchanged.

Adds a table-driven regression for the predicate: with A/B rules and a resync
of A, an object qualifying only for B is not admitted; any-target scoping
admits it and fails the test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 21:57:52 +08:00
Feng Ruohang 46e82eb54d fix(replication): count resync success by outcome, not target existence
The resync worker classified each object by whether the target version
merely existed (a tgt.StatObject HEAD), ignoring the outcome of the
replicateObject/replicateDelete call it had just made. A quota-rejected
update leaves the old version in place, so StatObject succeeded and the
resync recorded a false success - reported as Completed / N success /
0 failed and persisted across restart (issue #139). #134's SSE-C HEAD
marker made StatObject succeed for SSE-C too, exposing it there. The delete
path had the mirror flaw (a failed delete leaves the object, so the HEAD
succeeded), and FailedSize was never incremented (a failed 196,608-byte
object counted as 1 failed / 0 bytes).

replicateObject and replicateDelete already build the per-target
replicatedInfos (each replicatedTargetInfo carries Arn, ReplicationStatus
and Err) but discarded it. Return it (callers that only trigger replication
ignore the value - a Go call statement discards it, so the queue paths are
unchanged) and classify the resync from the target whose Arn == opts.arn via
a small pure helper:

- Completed without error -> replicated (+ that target's size, falling back
  to the object size).
- Failed or errored -> failed (+ the object size, fixing FailedSize).
- opts.arn absent from the result (not attempted) -> failed; a resync that
  cannot confirm the object reached the target is not a success.

The StatObject-existence block (including the delete-marker/MethodNotAllowed
special case, now subsumed by the delete outcome) is removed. #134's SSE-C
HEAD marker is left intact - it is needed for genuine SSE-C success.

Adds a table-driven regression for the classifier covering a completed
update, a failed update over an existing version, an errored-but-Completed
result, a delete failure, a delete-marker success (zero bytes) and an
un-attempted ARN. Classifying by existence makes the failed cases count
success and fails the test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 18:48:15 +08:00
Feng Ruohang f8ca4a8656 fix(replication): keep resync Completed status honest about object counts
resyncBucket could publish and persist a Completed resync status that did
not actually cover every object, in two ways:

1. It joined only the producer workers before the deferred markStatus ran,
   not the goroutine that folds each worker result into the status, so a
   Completed status could omit the last object (or a failed object) until the
   periodic ~1m flush (issue #136). The same finalization also closed the
   result channel on early-return paths while workers were still in flight,
   risking a send-on-closed-channel panic and a lost result.
2. markStatus persists under its own background context, so if the parent
   context was cancelled during the drain - workers then return without
   sending their computed result - or a worker dropped a result on the
   resync-cancel signal, a bare Completed was still recorded with counts that
   no longer matched the objects seen.

Fixes (count integrity only; the inherited cancellation deadlock, walker leak,
and single-token routing are tracked as separate follow-ups):

- Centralize shutdown in a resyncResults helper whose finish() stops the
  workers (closes inputs, waits for them to exit) before closing the result
  channel and waiting for the consumer to drain, then lets the deferred
  markStatus persist the final counts. finish() now runs on every exit path.
- Record a dropped result via sendResyncResult (a worker consuming the
  resync-cancel token returns without sending), and in the finalizer downgrade
  a Completed status to Failed via finalResyncStatus when the parent context
  was cancelled or a worker aborted - so a persisted Completed never
  misrepresents an incomplete resync.

Deterministic tests: an on-disk round-trip of the terminal status (complete
counts stay Completed; parent-cancel-during-drain and worker-abort each
downgrade to Failed), and testing/synctest drain/worker-order assertions that
fail deterministically if a finish() wait is removed. The inherited
cancellation structure (inline Walk, the dispatch send, the worker cancel
branches) is left unchanged for the follow-ups.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 15:28:55 +08:00
Feng Ruohang e7d0e62f16 Merge pull request #135 from pgsty/fix/attributes-encrypted-parts-test-reconcile
test: reconcile encrypted-parts attributes test with #119 write validation
2026-09-06 09:59:54 +08:00
Feng Ruohang aee290fc34 test: reconcile encrypted-parts attributes test with #119 write validation
TestAPIGetObjectAttributesEncryptedPartLengths (from #128) built its
fixtures by PutObjectPart-ing plaintext bodies under encrypted-object
metadata with per-part sizes 5245473 and 1. Since #119, PutObjectPart
always derives an encrypted part's plaintext length from the bytes
written and rejects a part that cannot be a valid sio stream, so those
fixtures can no longer be created through a normal write and both
variants failed at write time.

Such an on-disk shape now only exists as pre-#119 data or from an old
peer, which is exactly the state the GetObjectAttributes per-part
tamper check (#128) defends. Inject that ObjectInfo directly through a
stub object layer (the setObjectLayer pattern used by the #110 tamper
test) and exercise the handler, which is what this test pins. The
handler path, the crafted part sizes, and both assertions
(separately-encrypted-parts -> ErrObjectTampered; legacy-single-stream
-> stored fragment sizes) are unchanged. Test-only; reconciles two
already-merged correct changes (#119 and #128).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 09:51:09 +08:00
Feng Ruohang 62cce2b152 Merge pull request #121 from pgsty/fix/issue-110-tampered-status
fix: return 500 for unreadable objects instead of 206
2026-09-06 09:26:28 +08:00
Feng Ruohang 65d4806a7b Merge pull request #127 from pgsty/fix/issue-112-bucket-metadata-cors
fix: include per-bucket CORS in bucket metadata export and import
2026-09-06 09:26:21 +08:00
Feng Ruohang 765757473a Merge pull request #126 from pgsty/fix/issue-118-no-compressed-ssec
fix: exclude SSE-C objects from compression
2026-09-06 09:26:12 +08:00
Feng Ruohang 425bd7fff1 Merge pull request #124 from pgsty/fix/issue-119-ssec-part-actual-size
fix: record plaintext part sizes for replicated SSE-C multipart parts
2026-09-06 09:26:04 +08:00
Feng Ruohang e12e739a53 Merge pull request #128 from pgsty/fix/issue-114-115-object-attributes-parts
fix: report logical part sizes and end pagination correctly in GetObjectAttributes
2026-09-06 09:25:28 +08:00
Feng Ruohang 8b736dee34 Merge pull request #123 from pgsty/fix/issue-113-rotation-checksum-algorithm
fix: honor a requested checksum algorithm on SSE-C key rotation
2026-09-06 09:25:02 +08:00
Feng Ruohang 32e75c27bf Merge pull request #122 from pgsty/fix/issue-109-raw-ssec-replica
fix: store raw SSE-C replicas verbatim on the destination
2026-09-06 09:24:26 +08:00
Feng Ruohang 885ca604a1 Merge pull request #129 from pgsty/fix/issue-111-object-lock-replica-ordering
fix: order value-less replicated Object Lock updates by timestamp
2026-09-06 09:23:34 +08:00
Feng Ruohang d10382d0dc build: refresh rebrand compatibility baseline for the SSE-C replica helper
The raw SSE-C seal headers moved from an inline slice in
cmd/object-multipart-handlers.go to the shared isRawSSECReplica helper
in cmd/replication-trust.go (already recorded), so regenerate the
compat-baseline allowlist to drop the three stale object-multipart
entries. No behaviour change; make rebrand-guard is green.

Refs pgsty/silo#109

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 09:08:22 +08:00
mr javad seydi 0db4bf3b00 fix: report effective bucket quotas in metrics
Signed-off-by: mr javad seydi <seydi.birjand@gmail.com>
2026-09-05 21:23:46 +03:30
Feng Ruohang 87c621965d docs: describe the startup readiness window of the health probes
After a restart a node's remote erasure drives that could not be connected
during startup stay uninstalled until the next connectDisks pass, about 15
seconds later. During that window the liveness, readiness and both cluster
probes answer 200, admin info shows every drive online and mcli ready agrees,
because the cluster probes aggregate each peer's report of its own local
drives rather than the drives this node has installed. A PUT through that
node can still fail with 503 SlowDownWrite and a cross-node GET can answer
404 NoSuchKey until the window closes. The behaviour is inherited from
upstream and reproduced on the 0806 and 0903 releases alike.

Document the window and the bounded data-path check (PUT through each node,
read each object through every node, fixed deadline, re-read acknowledged
objects) that automation should use instead of the probes, and correct the
readiness probe description, which also fails on request-queue overload and
an unreachable KMS. No product change: the probes keep their documented
purpose, and changing them or the reconnect cadence was judged unproven
tuning in the agreed plan.

Refs pgsty/silo#116

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 17:17:05 +08:00
Feng Ruohang b2dca43fda fix: order value-less replicated Object Lock updates by timestamp
CopyObjectHandler rebuilt the destination metadata with the public Object Lock keys stripped (cmd/object-handlers.go:1708) and then restored a value only inside retentionMode.Valid() and legalHold.Status.Valid() (cmd/object-handlers.go:1715 and :1732), so a replica update that carried no retention or legal-hold value never reached the ordering comparison and silently erased whatever the destination held, however new it was; a retention removal that did win recorded no ordering timestamp either, so cmd/bucket-object-lock.go:370 later read an unparseable stored timestamp and let an older retained value back in. Each replica field is now decided on its source timestamp first and its incoming value second, and both restore helpers write the stored timestamp back before returning early on an empty stored value, which is the only way a removal timestamp survives the REPLACE metadata directive. Legal hold stays deliberately asymmetric: S3 has no legal-hold removal, an explicitly empty status is already rejected as invalid, and an absent status conveys no change even when an orphaned timestamp arrives with it, so only a valid ON or OFF can win.

Three inherited defects would have defeated that ordering, so they are fixed here too. The SSE-KMS branch of putOptsFromHeaders built its own ObjectOptions and dropped the parsed lock timestamps, leaving every replicated lock update unordered on a bucket with default KMS encryption; it now carries them. The in-place SSE-C key rotation snapshots the stored reserved metadata into encMetadata before the lock decision exists and merges it back afterwards to preserve the encryption headers, reinstating the ordering timestamp the decision had just replaced; the snapshot is now reconciled with the decision for a trusted replica. Finally, the value-less handling applies only to an actual replica: a trusted peer that sends the replication marker without REPLICA status keeps the previous behaviour, so a REPLACE copy carrying no lock headers still writes a version with no retention and no hold.

Tests: TestAPICopyObjectReplicaAbsentLockFieldsPreserveNewerState, TestAPICopyObjectReplicaRetentionRemovalKeepsOrderingTimestamp, TestAPICopyObjectReplicaObjectLockOrdering, TestAPICopyObjectReplicaRetentionRemovalUnderBucketKMS, TestAPICopyObjectReplicaLockTimestampSurvivesSSECKeyRotation and TestAPICopyObjectMarkerOnlyLeavesObjectLockUnchanged, all on ErasureSD and Erasure. Compatibility: no API, wire or stored-field change, and a field arriving with no source timestamp is unordered and now preserves destination state, so an un-upgraded 0806 peer keeps replicating safely while it still runs the old erasing receiver.

Fixes pgsty/silo#111

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 16:58:11 +08:00
Feng Ruohang 33a91d972f fix: carry per-bucket CORS through bucket metadata export/import
The admin bucket-metadata handlers enumerate every bucket config by name, and
per-bucket CORS was never added to that enumeration: export omitted cors.xml
(cmd/admin-bucket-handlers.go:414 cfgFiles) and import ignored the entry
outright, with no case in applyImportedBucketMetadata (:598) or SetStatus
(:629), so a CORS-only archive reported 0/0 buckets imported and a restored
bucket silently lost its configuration. Export now writes the stored document
verbatim and import validates it with the same parser and validator as
PutBucketCorsHandler, merging it under the existing bucket metadata lock and
announcing it through the dedicated SRBucketMetaTypeCorsConfig event; the local
CORS timestamp rule is extracted into localCORSUpdatedAt and reused so an
imported document always lands strictly above bucket creation, which matters
because the import stamps its fields before creating any missing bucket and a
CORS event below Created is dropped as an older bucket incarnation.

The import reads one byte past the declared entry size so archive/zip reaches
EOF and verifies the entry checksum, otherwise a corrupt or over-long entry
carrying well formed XML would overwrite the stored document; and the CORS
event is sent even when the shared bucket metadata hook failed, so an
unreachable peer cannot withhold an already committed CORS document from the
reachable ones.

Tests: TestAdminBucketMetadataCORSRoundTrip and
TestAdminBucketMetadataCORSImportReplicatesPastPeerFailure (new, ErasureSD and
Erasure).
Compatibility: the ZIP gains one entry, older archives stay importable and
leave CORS untouched; no mcli or madmin-go change is needed because
madmin.BucketStatus already carries Cors and mcli copies the export ZIP
verbatim.

Fixes pgsty/silo#112

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 16:21:12 +08:00
Feng Ruohang 2cbd48a3c3 fix: end GetObjectAttributes part pagination correctly
GetObjectAttributes decided truncation by comparing the last returned
part number with the part count (cmd/object-handlers.go:720), which is
only a coincidence of contiguous numbering. Sparse parts 1/3 reported a
complete page as truncated with a marker that loops, and parts 1/3/5
with max-parts=1 stopped after part 3 and silently dropped part 5. Set
IsTruncated in the break that proves an eligible part was left
unreturned, and zero NextPartNumberMarker when the listing is complete,
as ListObjectParts already does. Also reject negative x-amz-max-parts
and x-amz-part-number-marker in getAndValidateAttributesOpts with the
same API errors ListObjectParts uses, instead of answering an invalid
request with an empty parts listing; an absent or zero max-parts still
means the default page size.

Tests: TestAPIGetObjectAttributesPartsPagination (sparse 1/3/5 and
contiguous 1/2 walks on ErasureSD and Erasure),
TestGetAndValidateAttributesOptsPartsRange, and the sparse variants of
TestAPIGetObjectAttributesMultipartLogicalPartSize. Compatibility: no
field is added or removed; IsTruncated and NextPartNumberMarker change
only where they were wrong, and negative pagination values that no SDK
sends now fail fast.

Fixes pgsty/silo#115

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 16:17:39 +08:00
Feng Ruohang b5409ca112 fix: report logical part sizes in GetObjectAttributes
GetObjectAttributes filled ObjectPart.Size from the on-disk part length
(cmd/object-handlers.go:715), so every compressed or encrypted multipart
object reported transformed sizes that do not sum to the logical
ObjectSize the same response returns from objInfo.GetActualSize().
Report each part's uploaded plaintext length instead: a compressed part
uses its recorded ActualSize, and a separately encrypted part derives the
plaintext length with sio.DecryptedSize, because ActualSize is the
ciphertext length for a replicated SSE-C part and zero for parts written
before actualSize existed.

Only parts of an encrypted multipart object are streams of their own. A
legacy encrypted object carries no multipart marker and is one continuous
stream that the erasure writer split into storage fragments, so those
fragments keep their stored size. Where a part is a stream, one whose
length cannot be a valid encrypted stream has no logical length, and the
request now fails with XMinioObjectTampered rather than reporting the
ciphertext length; DecryptObjectInfo does not catch that case, because
ObjectInfo.isMultipart gives up on the first bad part and only the object
total is then validated.

Tests: TestAPIGetObjectAttributesMultipartLogicalPartSize (plain,
compressed, SSE-C and compressed+SSE-C, consecutive and sparse part
numbers), TestAPIGetObjectAttributesCompressedEmptyTrailingPart,
TestAPIGetObjectAttributesEncryptedPartLengths, and a part-size assertion
added to TestAPISSECMultipartReplicationTrust. Compatibility: the XML
shape is unchanged and nothing is written to disk, only the value of the
existing Size element is corrected.

Fixes pgsty/silo#114

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 16:17:39 +08:00
Feng Ruohang 35bd75948a fix: exclude SSE-C objects from compression
With compression allow_encryption=on an SSE-C object is stored as
encrypt(s2(plaintext)), while replication reads it raw (NoDecryption at
cmd/erasure-object.go:257) and putReplicationOpts drops the internal
compression and actual-size headers (cmd/bucket-replication.go:786). The
replica keeps the source seal with no compression marker, so a GET with the
correct customer key returns HTTP 200 and the raw S2 stream instead of the
object, and the source records the transfer as COMPLETED.

Widen the one condition in excludeForCompression (cmd/object-api-utils.go:613)
so SSE-C is never compressed, whatever allow_encryption says. This covers all
four producers at once, PutObject, NewMultipartUpload, CopyObject and
PutObjectExtract, plus any future caller of isCompressible.
crypto.SSEC.IsRequested ignores copy-source headers, so a copy is judged on its
destination key only, and a raw SSE-C replica write is unaffected because it
carries no public SSE-C headers. allow_encryption keeps its meaning for SSE-S3
and SSE-KMS, where the server owns the key and decompresses before replicating.

Tests: TestAPISSECCompressionReplicaStaysReadable (single PUT and multipart),
TestAPISSECCompressionProducerMatrix, TestAPISSECCompressionSkippedOnCopyObject,
TestAPISSECCompressionSkippedOnSnowballExtract and the control
TestSSECBatchReplicationCannotRead in cmd/compression-ssec_test.go. Two existing
expectations pinned the removed shape and are updated:
TestAPICopyObjectSSECKeyRotationNullVersionCompressesRewrite is renamed
TestAPICopyObjectSSECKeyRotationNullVersionSkipsCompression and now expects an
uncompressed rewrite, keeping its body, checksum, version and ETag assertions;
the SSE-C compressed-encrypted variant of
TestAPICopyObjectServerSideChecksumEncryption becomes compressible-extension and
expects an uncompressed destination, its SSE-S3 sibling keeping the compressed
coverage.

Compatibility: a deliberate behaviour change. Deployments with
allow_encryption=on no longer store new or rewritten SSE-C data compressed, so
those writes cost more space; objects already stored compressed keep working on
the source and are the concern of pgsty/silo#109, which rejects them at
replication time. Multipart uploads initiated before this change keep
compressing their parts from the metadata saved at initiation. Upstream
468a9fae8 refused this combination at PUT time and a2cab0255 removed the guard;
upstream master is still unguarded, so this is a deliberate divergence.

Fixes pgsty/silo#118

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 15:51:24 +08:00
Feng Ruohang 0c8d74205b fix: record plaintext part sizes for replicated SSE-C multipart parts
Trusted SSE-C replication uploads parts as raw ciphertext with the
ciphertext length as Content-Length, and erasureObjects.PutObjectPart only
derived the plaintext length when the caller passed a negative size, so
each replicated part persisted the ciphertext length as ActualSize (the
field defined as the uploaded size without encryption bytes). On the
replica, partNumberToRangeSpec turned those lengths into a plaintext range,
so GET/HEAD ?partNumber=N returned the wrong bytes and shifted
Content-Range (2560, 2560 and 5120 bytes for 5 MiB, 5 MiB and 1 MiB
parts), and a later decommission or rebalance re-uploaded the parts with
the stale value and recomputed the object-level actual-size from their sum,
after which a whole-object GET advertised a Content-Length larger than the
body it wrote.

Derive the plaintext length of an encrypted, uncompressed part from the
bytes actually written (sio.DecryptedSize) in PutObjectPart, the single
place a part is persisted, rejecting a length that cannot be a valid
stream before the part is committed; and derive part lengths from
part.Size in partNumberToRangeSpec for encrypted, uncompressed objects,
returning an error instead of a nil range, so replicas already on disk
read correctly without a resync. Compressed parts keep ActualSize.

Tests: TestAPISSECReplicaPartNumberReads (three-part SSE-C replica,
?partNumber=N bytes, Content-Length and Content-Range equal the source)
and TestSSECReplicaPartActualSizeDataMovement (replay through the data
movement path leaves part and object sizes at plaintext values) fail on
main and pass with the fix on ErasureSD and Erasure;
TestAPIGetObjectWithPartNumberHandler, TestAPISSECMultipartReplicationTrust
and TestAPIListObjectPartsHandler stay green. Compatibility: no wire or
API change; objects an unfixed server already moved carry a poisoned
object-level actual-size and need a rewrite or resync.

Fixes pgsty/silo#119

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 15:33:23 +08:00
Feng Ruohang fcc4d77895 fix: honor a requested checksum algorithm on SSE-C key rotation
An in-place SSE-C key rotation takes the fast path at cmd/object-handlers.go:1523
that only rewraps the object key, while every line that turns
x-amz-checksum-algorithm into a stored checksum lives in the re-encrypting else
branch at 1571-1609, so a requested algorithm was silently dropped and the stale
source checksum was kept and reported. Extend the canRotateKeyInPlace guard so a
client request carrying the header falls through to the copy that recomputes,
stores and reports it. Replica-trusted requests keep the fast path: getOpts
leaves their source reader encrypted, so a rewrite would hash ciphertext, and a
replica has to keep the checksum its source assigned.

Tests: TestAPICopyObjectSSECKeyRotationChecksumAlgorithm (new, red before the
guard), TestAPICopyObjectSSECKeyRotationKeepsChecksumAbsence (new, pins the
accepted limitation that a headerless rotation preserves the stored checksum
state including absence, gaining no default CRC64NVME) and
TestAPICopyObjectSSECKeyRotationReplicaKeepsFastPath (new, pins the replica
carve-out on a non-empty and on a zero byte source).
Compatibility: no API or wire change; a rotation without the header and every
replica-trusted rotation are unchanged, while a client rotation carrying the
header now rewrites the object data, so the ETag changes, a multipart source
collapses to a single part object, the copy replicates as an object rather than
as metadata, and the rewritten bytes are compressed if compression is enabled for
that object, as AWS CopyObject documents. Upstream MinIO carries the same
defect from 2718d9a43 (minio/minio#21399); this is a deliberate divergence.

Fixes pgsty/silo#113

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 15:25:08 +08:00
Feng Ruohang c52acc1a5d fix: store raw SSE-C replicas verbatim on the destination
A raw SSE-C replica write carries the source ciphertext and the source seal
in X-Minio-Replication-Server-Side-Encryption-* headers but no public SSE-C
request headers, so crypto.Requested() was false and PutObjectHandler and
NewMultipartUploadHandler applied the destination's default encryption and
compression to bytes that were already ciphertext (upstream 468a9fae8,
"Enable replication of SSE-C objects", never exempted the raw path). With
destination default SSE-S3 the replica's IV and seal were overwritten and
GET returned 400; with destination compression the replica stored
compress(ciphertext) and GET failed, while the source reported COMPLETED.

Recognize a validated raw SSE-C replica (replicaTrusted plus a seal header,
shared helper isRawSSECReplica) and skip bucket default encryption,
compression and the encryption branch on the single PUT path, and default
encryption plus compression on the multipart initiation path, which already
skipped key generation. On the sender, reject replication of an object that
is both compressed and SSE-C, since the wire carries no compression state
and the destination would otherwise store an undetectable S2 stream, and
make replicateObject/replicateAll report a putReplicationOpts failure as
Failed instead of Completed.

Tests: TestAPISSECReplicaSkipsDestinationTransforms (single PUT and
multipart under destination default SSE-S3, compression and an explicit SSE
header, plus an untrusted control), TestAPISSECMultipartReplicaRoundTripWith
Compression, and TestPutReplicationOptsRejectsCompressedSSEC fail on main
and pass with the fix on ErasureSD and Erasure; the replication-trust,
multipart and PutObject suites stay green. Compatibility: no wire, API or
metadata change; the destination change applies only to trusted replica
writes carrying a source seal; replicas already transformed must be
rewritten from an intact source (see pgsty/silo#120 for why a resync does
not do that yet).

Fixes pgsty/silo#109

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 15:07:22 +08:00
Feng Ruohang 5703426b3c fix: return 500 for unreadable objects instead of 206
ErrObjectTampered was mapped to http.StatusPartialContent since upstream
ca6b4773e (2017), so a GET or HEAD of an object the server cannot decode
(invalid encrypted size, malformed actual-size, bad multipart ETag shape)
answered with a success status and an XML error document that SDKs handed
back as object content; boto3 returned the XML as Body and a zero-length
HEAD as success. Every origin of errObjectTampered is a stored-state
defect, not caller input, so map the entry to 500 Internal Server Error
and keep the XMinioObjectTampered code and message. The comment records
the deliberate divergence from upstream.

Tests: TestObjectTamperedGETHEADStatus (signed GET and HEAD, ErasureSD and
Erasure) fails with 206 on main and passes with 500; TestAPIErrCode,
TestAPIErrCodeDefinition, TestAPIHeadObjectHandler,
TestAPIHeadObjectHandlerWithEncryption and TestAPIGetObjectHandler stay
green. Compatibility: only the status line of one MinIO-specific error
changes; clients now retry damaged-object reads per their 5xx policy.

Fixes pgsty/silo#110

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 14:50:00 +08:00
Feng Ruohang f0bd164b92 docs: refresh community contributors
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 19:21:26 +08:00
Feng Ruohang 9936a69d89 ci: recover container publication from verified component pins
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 15:10:43 +08:00
Feng Ruohang ce2326c946 build: pin the published mcli 20260903 archives
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 14:38:54 +08:00
Feng Ruohang 4c164907f5 build: align Helm defaults with the published release tag
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 14:36:12 +08:00
mr javad seydi 7a060cab1e feat(ilm): relocate hot objects across server pools by GET frequency
Keep NVMe/HDD pool pairs useful without remote tiering: promote objects that
are read often, demote previously moved objects once they go idle, and leave
the feature off until operators set a two-pool topology.

Signed-off-by: mr javad seydi <seydi.birjand@gmail.com>
2026-08-15 13:48:22 +03:30
458 changed files with 54611 additions and 2440 deletions
-54
View File
@@ -1,54 +0,0 @@
---
name: Bug report
about: Create a report to help us improve
title: ''
labels: community, triage
assignees: ''
---
## NOTE
Silo issues are handled by community maintainers on a best-effort basis. There
is no SLA, SLO, or emergency production-support channel. Follow the local
[Code of Conduct](../code_of_conduct.md) when participating. Report suspected
vulnerabilities through the private process in [SECURITY.md](../SECURITY.md),
not in a public issue.
<!--- Provide a general summary of the issue in the Title above -->
## Expected Behavior
<!--- If you're describing a bug, tell us what should happen -->
<!--- If you're suggesting a change/improvement, tell us how it should work -->
## Current Behavior
<!--- If describing a bug, tell us what happens instead of the expected behavior -->
<!--- If suggesting a change/improvement, explain the difference from current behavior -->
## Possible Solution
<!--- Not obligatory, but suggest a fix/reason for the bug, -->
<!--- or ideas how to implement the addition or change -->
## Steps to Reproduce (for bugs)
<!--- Provide a link to a live example, or an unambiguous set of steps to -->
<!--- reproduce this bug. Include code to reproduce, if relevant -->
<!--- and include relevant Silo logs with secrets and credentials removed -->
1.
2.
3.
4.
## Context
<!--- How has this issue affected you? What are you trying to accomplish? -->
<!--- Providing context helps us come up with a solution that is most useful in the real world -->
## Regression
<!-- Is this issue a regression? (Yes / No) -->
<!-- If Yes, optionally include the Silo version, commit id, or PR that caused this regression. -->
## Your Environment
<!--- Include as many relevant details about the environment you experienced the bug in -->
* Version used (`silo --version`):
* Server setup and configuration:
* Operating System and version (`uname -a`):
+8
View File
@@ -7,6 +7,14 @@ assignees: ''
---
Report bugs in the PGSTY SILO server (`pgsty/silo`) here. Community maintainers
handle reports on a best-effort basis. There is no SLA, SLO, or emergency
production-support channel. Follow the
[Code of Conduct](https://github.com/pgsty/silo/blob/main/code_of_conduct.md).
Report suspected vulnerabilities privately through
[SECURITY.md](https://github.com/pgsty/silo/blob/main/SECURITY.md).
For patches, see the [contribution guide](https://github.com/pgsty/silo/blob/main/CONTRIBUTING.md).
<!--- Provide a general summary of the issue in the Title above -->
## Expected Behavior
@@ -7,6 +7,9 @@ assignees: ''
---
Suggest improvements to the PGSTY SILO server (`pgsty/silo`) here.
For patches, see the [contribution guide](https://github.com/pgsty/silo/blob/main/CONTRIBUTING.md).
**Is your feature request related to a problem? Please describe.**
A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]
+7 -3
View File
@@ -1,10 +1,14 @@
## Contribution Licensing (no CLA, inbound=outbound, DCO required)
This project does not use a CLA; contributions are accepted inbound=outbound.
This pull request contributes to PGSTY SILO (`pgsty/silo`). Code contributions
are accepted under AGPL-3.0-or-later, the same license as the server.
This project does not use a CLA or require a separate Apache-2.0 license grant.
By submitting this pull request I represent that I have the right to contribute
the changes, which are licensed under this repository's
the code changes under this repository's
[GNU Affero General Public License v3.0 or later](https://www.gnu.org/licenses/agpl-3.0.html)
and remain my copyright. Every commit must carry a DCO `Signed-off-by` trailer
and retain copyright in my original work. Existing copyright and license
notices remain intact; separately licensed material keeps its applicable terms.
Every commit must carry a DCO `Signed-off-by` trailer
(`git commit -s`) certifying the
[Developer Certificate of Origin](https://developercertificate.org/) — see
[CONTRIBUTING.md](https://github.com/pgsty/silo/blob/main/CONTRIBUTING.md).
+60 -3
View File
@@ -7,6 +7,11 @@ on:
description: "Published RELEASE.* tag to package as pgsty/silo"
required: true
type: string
recovery:
description: "Run the current main workflow against an already-published tag"
required: false
default: false
type: boolean
permissions:
contents: read
@@ -75,12 +80,18 @@ jobs:
fetch-depth: 0
- name: Verify workflow identity matches release source
env:
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
RECOVERY: ${{ inputs.recovery }}
run: |
set -euo pipefail
CHECKED_OUT_REVISION="$(git rev-parse HEAD)"
if [ "${CHECKED_OUT_REVISION}" != "${GITHUB_SHA}" ]; then
echo "Checked out ${CHECKED_OUT_REVISION}, but workflow identity is ${GITHUB_SHA}. Dispatch this workflow from ${RELEASE_TAG}." >&2
exit 1
if [ "${RECOVERY}" != "true" ] || [ "${GITHUB_REF}" != "refs/heads/${DEFAULT_BRANCH}" ]; then
echo "Checked out ${CHECKED_OUT_REVISION}, but workflow identity is ${GITHUB_SHA}. Dispatch from ${RELEASE_TAG}, or use recovery from ${DEFAULT_BRANCH}." >&2
exit 1
fi
echo "Recovery workflow ${GITHUB_SHA} is packaging published source ${CHECKED_OUT_REVISION}."
fi
- name: Prepare verified Docker contexts
@@ -129,10 +140,50 @@ jobs:
mkdir -p "${context}/dockerscripts"
tar -xzf "${archive}" -C "${context}" silo
cp Dockerfile.goreleaser Dockerfile.distroless LICENSE NOTICE CREDITS "${context}/"
cp dockerscripts/docker-entrypoint.sh dockerscripts/download-static-curl.sh \
cp dockerscripts/docker-entrypoint.sh dockerscripts/build-static-curl.sh \
"${context}/dockerscripts/"
done
# The classic image bundles mcli. Resolve its two archive digests
# from the immutable published release instead of trusting defaults
# copied into an older Server tag. This also gives a recovery run a
# narrow override when a tag selected the right mcli release but
# accidentally retained stale archive pins.
MC_REPO="$(awk -F= '/^ARG MC_REPO=/{print $2; exit}' Dockerfile.goreleaser)"
MC_VERSION="$(awk -F= '/^ARG MC_VERSION=/{print $2; exit}' Dockerfile.goreleaser)"
test -n "${MC_REPO}"
test -n "${MC_VERSION}"
MC_VERSION_HYPHEN="${MC_VERSION#RELEASE.}"
MC_PKG_VERSION="$(echo "${MC_VERSION_HYPHEN}" | sed -E 's/^([0-9]{4})-([0-9]{2})-([0-9]{2})T([0-9]{2})-([0-9]{2})-([0-9]{2})Z$/\1\2\3\4\5\6.0.0/')"
if [ "${MC_PKG_VERSION}" = "${MC_VERSION_HYPHEN}" ]; then
echo "Invalid bundled mcli tag: ${MC_VERSION}" >&2
exit 1
fi
if [ "$(gh release view "${MC_VERSION}" --repo "${MC_REPO}" --json isDraft --jq .isDraft)" != false ] || \
[ "$(gh release view "${MC_VERSION}" --repo "${MC_REPO}" --json isPrerelease --jq .isPrerelease)" != false ] || \
[ "$(gh release view "${MC_VERSION}" --repo "${MC_REPO}" --json isImmutable --jq .isImmutable)" != true ]; then
echo "Bundled mcli ${MC_REPO}@${MC_VERSION} must be a published immutable release" >&2
exit 1
fi
mc_checksums="mcli_${MC_PKG_VERSION}_checksums.txt"
gh release download "${MC_VERSION}" --repo "${MC_REPO}" \
--dir "${assets_dir}" --pattern "${mc_checksums}"
gh attestation verify "${assets_dir}/${mc_checksums}" \
--repo "${MC_REPO}" \
--signer-workflow "${MC_REPO}/.github/workflows/release.yml" \
--source-ref "refs/tags/${MC_VERSION}" >/dev/null
MC_AMD64_SHA256="$(awk -v name="mcli_${MC_PKG_VERSION}_linux_amd64.tar.gz" '$2 == name {print $1}' "${assets_dir}/${mc_checksums}")"
MC_ARM64_SHA256="$(awk -v name="mcli_${MC_PKG_VERSION}_linux_arm64.tar.gz" '$2 == name {print $1}' "${assets_dir}/${mc_checksums}")"
[[ "${MC_AMD64_SHA256}" =~ ^[0-9a-f]{64}$ ]]
[[ "${MC_ARM64_SHA256}" =~ ^[0-9a-f]{64}$ ]]
{
echo "MC_AMD64_SHA256=${MC_AMD64_SHA256}"
echo "MC_ARM64_SHA256=${MC_ARM64_SHA256}"
} >> "${GITHUB_ENV}"
echo "RELEASE_REVISION=$(git rev-parse HEAD)" >> "${GITHUB_ENV}"
- name: Set up QEMU
@@ -162,6 +213,9 @@ jobs:
file: docker-release/amd64/Dockerfile.goreleaser
platforms: linux/amd64
push: true
build-args: |
MC_AMD64_SHA256=${{ env.MC_AMD64_SHA256 }}
MC_ARM64_SHA256=${{ env.MC_ARM64_SHA256 }}
tags: |
pgsty/silo:${{ env.RELEASE_TAG }}-amd64
pgsty/silo:latest-amd64
@@ -178,6 +232,9 @@ jobs:
file: docker-release/arm64/Dockerfile.goreleaser
platforms: linux/arm64
push: true
build-args: |
MC_AMD64_SHA256=${{ env.MC_AMD64_SHA256 }}
MC_ARM64_SHA256=${{ env.MC_ARM64_SHA256 }}
tags: |
pgsty/silo:${{ env.RELEASE_TAG }}-arm64
pgsty/silo:latest-arm64
+3
View File
@@ -90,6 +90,9 @@ jobs:
- name: Run S3 Select tests under race detector
run: go test -race ./internal/s3select/... -count=1
- name: Run conditional PUT tests under race detector
run: go test -race ./cmd -run '^Test(PoolsConditionalPut|SinglePoolConditionalPutHTTP)' -count=1 -timeout=5m
crosscompile:
name: Cross Compile
runs-on: ubuntu-latest
+24 -2
View File
@@ -10,7 +10,7 @@ on:
- "Dockerfile.distroless"
- "cmd/healthcheck-main.go"
- "cmd/main.go"
- "dockerscripts/download-static-curl.sh"
- "dockerscripts/build-static-curl.sh"
- "dockerscripts/docker-entrypoint.sh"
- "dockerscripts/docker-entrypoint_test.sh"
- "silo.service"
@@ -41,6 +41,28 @@ permissions:
contents: read
jobs:
curl:
name: Static curl (${{ matrix.arch }})
runs-on: ${{ matrix.runner }}
strategy:
matrix:
include:
- arch: amd64
runner: ubuntu-latest
- arch: arm64
runner: ubuntu-24.04-arm
steps:
- uses: actions/checkout@v7
- name: Build and exercise curl in an empty runtime
run: |
docker build --build-arg TARGETARCH=${{ matrix.arch }} \
--target curl-runtime -f Dockerfile.goreleaser -t silo-curl-test .
docker run --rm silo-curl-test --version | tee curl-version.txt
grep -F 'curl 8.22.0 ' curl-version.txt
grep -F 'HTTP2' curl-version.txt
docker run --rm silo-curl-test --fail --silent --show-error \
--connect-timeout 15 --max-time 60 https://curl.se/robots.txt
validate:
runs-on: ubuntu-latest
steps:
@@ -490,7 +512,7 @@ jobs:
bash -n buildscripts/package/lifecycle_test.sh
buildscripts/package/lifecycle_test.sh
bash -n dockerscripts/docker-entrypoint_test.sh
bash -n dockerscripts/download-static-curl.sh
bash -n dockerscripts/build-static-curl.sh
dockerscripts/docker-entrypoint_test.sh
go run ./buildscripts/rebrand-guard
buildscripts/verify-rebrand.sh
+1 -1
View File
@@ -29,7 +29,7 @@ jobs:
- name: Install govulncheck
run: |
go install golang.org/x/vuln/cmd/govulncheck@v1.7.0
go install golang.org/x/vuln/cmd/govulncheck@v1.8.0
echo "$(go env GOPATH)/bin" >> "${GITHUB_PATH}"
- name: Run govulncheck
+162
View File
@@ -0,0 +1,162 @@
# Changelog
## Unreleased
The entries below describe source changes on main since the latest published Server.
**The latest published Server remains 20260903.** These changes are not in its
binaries, packages or images. See the [component matrix](https://silo.pgsty.com/compatibility/versions/)
and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-09-03T13-18-01Z...main).
### Authorization and security
- Persist IAM deletion revisions and parent revocation boundaries so stale site
events cannot restore deleted identities, policies or their older grants
(#191, #192). Peer deletion notifications reload committed storage; deliberate
recreation requires a newer revision, and credentials issued before the
parent's revocation remain invalid.
**Coordinated upgrade required:** upgrade every participating node and site.
Mixed old/new nodes sharing an IAM backend and rolling downgrade are
unsupported. Back up complete IAM storage and encryption material; an admin
export of live records omits deletion history. Reissue credentials for
recreated parents and explicitly reconcile pre-upgrade revocations whose
history is already lost. Restoring an older backup can lose later revocations;
keep affected sites isolated until reconciliation/rekeying is complete. See
[the operator runbook](https://github.com/pgsty/silo.pgsty.com/blob/7bd2d57c2ce5aaa804d0b1a2fe0e5eed69d15235/content/operations/replication/iam-upgrade.md).
- Enforce an absolute HTTP/1 request-header deadline through the connection
wrapper (#196). Repeated small reads no longer extend that deadline, and
`--read-header-timeout` / `MINIO_READ_HEADER_TIMEOUT` now reaches the HTTP
server. HTTP/1 request bodies retain the rolling idle timeout; this does not
impose a total upload/download duration. A shorter setting also constrains
TLS handshake reads. The wrapper's strict header mode is not applied to HTTP/2.
- Reject unsigned `x-amz-*` request headers that could turn a signed PUT into a
copy of another object accessible to the signer (SN-2026-011). The latest
public Server is affected; the fix is on main. See [the advisory ledger](docs/security/advisories.md).
- Align signed request fields with policy conditions and enforce header-only
presigned payload checksums. See [the signed-header review](https://silo.pgsty.com/blog/design/signed-header-coverage/).
- **Breaking policy semantics:** separate self-service `admin:ChangeMyPassword`
from `admin:CreateUser`. Built-in read-only policies follow the split. Preserve
both denies if the previous combined restriction must survive upgrades or
rollback. Saved policies are not rewritten. Deploy with the matching Console
and pkg; see [the migration guide](docs/iam/password-permissions.md).
### Object storage and replication
- Preserve object tags during multi-pool metadata reconciliation by reading the
resolved tag field together with its revision (#189). Previously, reconciliation
could replace existing tags with an empty value.
- Preserve the tag revision on SSE-KMS metadata replication (#193), and advance
tag revisions monotonically on local PUT/DELETE tagging (#196). Empty tags
participate in reconciliation as an ordered deletion, preventing older
events from restoring removed tags. SSE-C key rotation also retains the tag
revision. Malformed historical revisions can fail and retry; their missing
history is not reconstructed by the upgrade.
- Complete delete-marker version purges and preserve their identity and retry
state through MRF recovery (#196). Recovery accepts a 405 marker response only
when its version, bucket, object name and modification time match the task.
Purge audit status is normalized from `COMPLETE` to `COMPLETED`.
Thanks to Julien Laurenceau (@julienlau) for the investigation and proposed
fix in #184 that helped shape this follow-up.
- Restore only the six replication-specific metadata fields after ordinary
request metadata extraction (#194). This prevents transport-only `aws-chunked`
from being stored as Content-Encoding while preserving the signed-header
protections. Trusted Snowball entries no longer inherit the outer archive's
ordinary metadata. Thanks to Mikhail Khadarenka (@chodorenko) for the fix in #187.
**Existing data:** these repairs prevent new errors; they do not scan or rewrite
historical object metadata, recover lost tags or prove that old purge work has
converged. Follow the [read-only audit procedure](https://github.com/pgsty/silo.pgsty.com/blob/7bd2d57c2ce5aaa804d0b1a2fe0e5eed69d15235/content/operations/replication/replica-metadata-audit.md)
before planning any repair of stored state.
- Evaluate conditional multipart completion against the logical current object
across all pools while holding the existing object lock. A stale `If-Match`
can no longer replace newer data in another pool, and the current ETag is no
longer rejected because the upload resides next to an older copy. Conditions
are evaluated once; a current delete marker counts as an absent object.
**Availability change:** if any pool's metadata cannot be read, conditional
completion fails even when another pool can still serve GET/HEAD. This also
applies when the unreadable pool may not hold the object: absence cannot be
verified. Retry after the pool recovers. Unconditional completion and the
single-pool path retain their existing behavior.
- Evaluate ordinary multi-pool conditional PUT against the logical current
object across all pools, including draining pools, under the existing object
lock (#207). A stale destination copy no longer accepts a stale ETag or rejects
the current one; a current delete marker is treated as absence.
**Availability change:** if any pool's object metadata cannot be verified,
the condition fails even when GET can use another pool; read-quorum failures
return 503. Restore readability or heal before retrying. Unconditional PUT,
single-pool conditions and internal replication retain their existing behavior.
A public condition with a destination `versionId` compares the current object
while preserving the requested write version. This change does not retire
stale copies in other pools, undo historical accepted overwrites or provide
a new global clock-ordering guarantee. The multipart-completion repair in #190
neither introduced nor repaired this separate PUT defect.
- Reconcile ordinary single-object version DELETE across all pools, including
null versions, delete markers and unqualified directory-marker DELETE. This
applies the deletion to every resolved pool copy under existing quorum
rules. Pending outbound delete replication retains versions until the
existing replication worker completes their purge; a successful response
does not imply immediate physical removal from every drive. Unreadable
pools now consistently return 503 instead of depending on pool traversal
order; insufficient read quorum returns `SlowDownRead`. This extends the
existing failure surface. Retry after recovery.
Cleanup failures also return an error. Batch deletion already fans out across
pools; replication and scanner cleanup keep their existing contracts. See
[scope and limitations](docs/bucket/lifecycle/access-tiering-removal.md#version-deletion-scope).
- Remove the opt-in GET-frequency pool-tiering feature from PR #60, including
its tracker, mover, scanner hooks, configuration, XML actions and metrics.
Accept and ignore retired configuration/XML and preserve ordinary statistics
when reading v9 caches. See [migration notes](docs/bucket/lifecycle/access-tiering-removal.md).
The [decision record](docs/investigations/access-tiering-revert.md) preserves
the feature's introduction, subsequent fixes, rollback scope and review history.
- Preserve the independent multi-pool write, metadata, healing and conditional
deletion fixes from PR #178, including shared remote-tier reference protection.
- Enforce `If-Match` on DELETE, preserve retention and independently ordered
Object Lock/tag updates, and correctly retransmit encrypted replicas.
- Preserve plaintext part sizes and raw SSE-C replicas; prevent SSE-C
compression, honor key-rotation checksums, and complete attributes pagination.
- Repair federated CopyObject checksums, destination timestamps, reserved
metadata, encrypted-object forwarding, legal hold and KMS context.
- Make resync counters, target selection, cancellation and worker lifetimes
reflect actual work, and report bounded MRF drops.
- Converge bucket metadata with deterministic source state, deletion tombstones,
creation time recovery and diagnostics. The mixed-version export gate requires
coordinated upgrades before tombstones are exported. See [the #77 record](docs/investigations/issue-77-current.md).
- Include per-bucket CORS in metadata export/import, close metadata publication
and logger races, and report effective bucket quotas in metrics.
### Console, dependencies and delivery
- Restore embedded Console login over loopback TLS, trusted-proxy handling and
all four WebSocket connection limits. Preserve Go TLS defaults across transports.
- Directly require `github.com/pgsty/silo-pkg/v3` v3.14.0; select Console
`v0.0.0-20260913015128-417559bb2c97` and MC
`v0.0.0-20260913012246-4f609a4da3bb` with explicit PGSTY replacements.
- Pin upstream minio-go `v7.3.1-0.20260910142817-60bd07042d49`; refresh Go x/*
modules and security fixes including bounded AMQP frame handling. Keep Go
1.27.1 and go-systemd v22.6.0's NetBSD compatibility replacement.
- Refresh container base digests and build static curl 8.22.0 from verified
source for both Linux architectures. Pin the actual mcli 20260913 archives and
hashes. Helm's client image follows that release; its Server image still names
the latest published Server 20260903.
The dependency update passed the final candidate's Go, vulnerability and Test
Release workflows; native curl builds passed on both architectures. A local
ARM64 image passed startup, health, S3 transfer and embedded Console checks.
These checks do not publish a Server tag or production image and do not replace
cluster upgrade/rollback acceptance for the next release. Dated investigations
retain the exact source and runtime boundaries they tested.
## RELEASE.2026-09-03T13-18-01Z
Published source: `9b11dc9469e650815b775cb47b039610644f5da4`.
[Complete release notes](https://silo.pgsty.com/blog/release/silo-20260903/) ·
[GitHub release](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-09-03T13-18-01Z)
This release ships Go 1.27.1, silo-pkg v3.13.2, upstream minio-go `0e78d3f18efe`,
mcli 20260903 and embedded Console source `464a59d73ada` (v2.3.0 version identity).
Installing the newer standalone mcli or Console does not replace components
inside this existing Server binary or image.
Earlier releases: [release archive](https://github.com/pgsty/silo/releases).
+14 -8
View File
@@ -79,9 +79,10 @@ documentation is owned by the separate
## Licensing of Contributions
Silo is licensed under the [GNU AGPL v3.0 or later](LICENSE). Its core is
Copyright (c) MinIO, Inc.; the combined work can never be relicensed, and this
fork does not try to.
Code contributions to PGSTY SILO (`pgsty/silo`) are accepted under the
[GNU AGPL v3.0 or later](LICENSE), the same license as the server. Submit issues
and pull requests to this repository's maintainers. No separate Apache-2.0
license grant to SILO or upstream MinIO maintainers is required.
* **No CLA.** We do not ask you to sign a Contributor License Agreement and we
do not take your copyright. Contributions are accepted inbound=outbound: you
@@ -112,15 +113,20 @@ fork does not try to.
`Signed-off-by` trailers) and add your own sign-off as the person passing it
along. Never import code from a proprietary distribution.
* **File headers.** Files derived from upstream keep the original MinIO
copyright header unchanged. New files added by this fork use the dual
header, followed by the standard AGPL boilerplate:
* **File headers.** Preserve existing copyright and license notices in inherited
and third-party files. New original files name their actual copyright holders
and use AGPL-3.0-or-later. Use a header such as the following, then append the
standard AGPL boilerplate:
```
// Copyright (c) 2015-2025 MinIO, Inc.
// Copyright (c) 2025-2026 PGSTY
// Copyright (c) 2026 Your Name
```
* **Separately licensed material.** Documentation contributions in `docs/`
follow its existing [CC BY 4.0 license](docs/LICENSE). Third-party components
and earlier Apache-2.0 contributions retain their original licenses and
attribution; this policy does not relicense earlier work.
* **Squash merges** must keep the `Signed-off-by:` trailers in the resulting
commit message.
+170 -65
View File
File diff suppressed because one or more lines are too long
+19 -11
View File
@@ -1,4 +1,15 @@
FROM golang:1.27.1-alpine AS build
FROM golang:1.27.1-alpine@sha256:cf6fca6641884b8433441b2b0652976f975e1d0fdd26d177eaaf8596087f3125 AS curl-build
ARG TARGETARCH
COPY dockerscripts/build-static-curl.sh /build/build-static-curl
RUN /bin/sh /build/build-static-curl
# Exercise the exact shipped curl without a dynamic loader or shared libraries.
FROM scratch AS curl-runtime
COPY --from=curl-build /go/bin/curl /curl
COPY --from=curl-build /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/ca-certificates.crt
ENTRYPOINT ["/curl"]
FROM golang:1.27.1-alpine@sha256:cf6fca6641884b8433441b2b0652976f975e1d0fdd26d177eaaf8596087f3125 AS build
ARG TARGETARCH
@@ -6,9 +17,9 @@ ENV GOPATH=/go
ENV CGO_ENABLED=0
ARG MC_REPO=pgsty/mc
ARG MC_VERSION=RELEASE.2026-09-03T07-13-05Z
ARG MC_AMD64_SHA256=6387cbeebb17c4bd52b447ee332ba22777e129034befa6598308c3aa04f02d09
ARG MC_ARM64_SHA256=5fc434c7e416bb4787e92306a8165d55284aac639a29744160b2a198fdd7573a
ARG MC_VERSION=RELEASE.2026-09-13T00-00-00Z
ARG MC_AMD64_SHA256=9d2a92de9c7b887d9b944fe9ddce68d23f1b6df3415092e737594e56593f5e2b
ARG MC_ARM64_SHA256=3d82e9ea6c601c4cb44fe5dd5f2ad1b7d7d64369378110f9ada9c524688a452a
RUN apk add -U --no-cache \
ca-certificates \
@@ -56,18 +67,14 @@ RUN apk add -U --no-cache \
chmod +x /go/bin/mcli && \
ln -sf mcli /go/bin/mc
COPY dockerscripts/download-static-curl.sh /build/download-static-curl
RUN chmod +x /build/download-static-curl && \
/build/download-static-curl
FROM registry.access.redhat.com/ubi9/ubi:latest AS certs
FROM registry.access.redhat.com/ubi9/ubi:latest@sha256:206b65b8ee0f04b992818c9a51b29081b14974630d4850bc358097d0c44ea156 AS certs
RUN dnf -y install ca-certificates && \
update-ca-trust && \
cp /etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem /tmp/ca-certificates.crt && \
dnf clean all && \
rm -rf /var/cache/dnf
FROM registry.access.redhat.com/ubi9/ubi-micro:latest
FROM registry.access.redhat.com/ubi9/ubi-micro:latest@sha256:f332c99eb8f798a8486821c91937f10ad64ee83d7e739303be2df051040918f6
LABEL org.opencontainers.image.title="Silo" \
org.opencontainers.image.description="S3-Interface Libre Object Storage" \
@@ -88,7 +95,8 @@ ENV MINIO_ACCESS_KEY_FILE=access_key \
COPY --from=certs /tmp/ca-certificates.crt /etc/ssl/certs/ca-certificates.crt
COPY silo /usr/bin/silo
COPY --from=build /go/bin/mcli /usr/bin/mcli
COPY --from=build /go/bin/curl* /usr/bin/
COPY --from=curl-build /go/bin/curl /usr/bin/curl
COPY --from=curl-build /go/share/curl /licenses/curl
COPY dockerscripts/docker-entrypoint.sh /usr/bin/docker-entrypoint.sh
COPY LICENSE /licenses/LICENSE
COPY NOTICE /licenses/NOTICE
+1 -1
View File
@@ -216,7 +216,7 @@ docker: checks build-debugging ## builds the local Linux Silo container image
--ldflags "$(LDFLAGS)" -o "$$context/silo"; \
mkdir -p "$$context/dockerscripts"; \
cp Dockerfile.goreleaser LICENSE NOTICE CREDITS "$$context/"; \
cp dockerscripts/docker-entrypoint.sh dockerscripts/download-static-curl.sh \
cp dockerscripts/docker-entrypoint.sh dockerscripts/build-static-curl.sh \
"$$context/dockerscripts/"; \
docker build -q --no-cache --platform linux/$(GOARCH) -t $(TAG) --build-arg TARGETARCH=$(GOARCH) \
-f "$$context/Dockerfile.goreleaser" "$$context"
+73 -45
View File
@@ -35,6 +35,14 @@
> [!NOTE]
> Renamed from `pgsty/minio` to `pgsty/silo`, default branch `master` → `main`, on 2026-08-06. Artifacts under the original MinIO identity stay published on the archived [`minio`](https://github.com/pgsty/silo/tree/minio) branch and in releases up to [`RELEASE.2026-08-04T00-00-00Z`](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-08-04T00-00-00Z).
## Current release and main branch
The latest published Server is [20260903](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-09-03T13-18-01Z).
As of 2026-09-13, the main branch has newer security, storage, Console and
shared-package changes that have not shipped in a Server release. See
[CHANGELOG.md](CHANGELOG.md) and the [component version matrix](https://silo.pgsty.com/compatibility/versions/)
for the exact release/source boundary, including SN-2026-011 and password-policy migration.
## Overview
PGSTY SILO keeps one maintained release line of the open-source MinIO server alive after upstream ended community distribution: builds, packages, multi-arch images, security fixes, and the full web console. Pigsty runs it in production as its PostgreSQL backup repository.
@@ -89,59 +97,79 @@ The S3 API, `MINIO_*` variables, `minio_*` metrics, `x-minio-*` headers, `/minio
Every divergence from upstream is listed in the code-verified [compatibility audit](https://silo.pgsty.com/compatibility/server/). Treat each release as a downstream upgrade: pin versions, read the [release notes](https://silo.pgsty.com/tags/silo/), and keep a rollback path.
### TLS and Go upgrades
TLS key exchange follows Go's defaults across the S3 listener, node links,
replication, identity providers, etcd, and external HTTP services. If an endpoint
cannot accept ML-KEM, `GODEBUG=tlsmlkem=0` disables the default hybrid exchanges
for the process; certificate verification remains enabled. This option does not
disable ML-DSA signatures or resolve every TLS reset. Prefer updating the
incompatible endpoint before removing the temporary setting.
If only the new SecP hybrids cause problems, `GODEBUG=tlssecpmlkem=0` disables
those groups while retaining X25519MLKEM768.
For builds targeting Go 1.27, setting either `SSL_CERT_FILE` or `SSL_CERT_DIR`
on macOS replaces Keychain trust with on-disk roots and Go's verifier. Stale or
incomplete CA paths can break previously trusted connections; unset inherited
values to restore Keychain trust. Explicit certificates in the configured `CAs`
directory remain additive to the selected root pool.
Go 1.27 binaries require macOS 13 or later. See the
[Go release notes](https://go.dev/doc/go1.27) and the
[SILO stack investigation](docs/investigations/go127-stack.md).
## Security & Contributing
Report vulnerabilities privately as described in [`SECURITY.md`](SECURITY.md); every fix ships with a public [advisory](https://silo.pgsty.com/blog/security/). Contributions are accepted inbound=outbound under AGPL-3.0-or-later with no CLA — only DCO sign-off (`git commit -s`) is required; see [`CONTRIBUTING.md`](CONTRIBUTING.md).
## Contributors
<table>
<tr>
<td align="center" width="150">
<a href="https://github.com/ZouhairCharef"><img src="https://github.com/ZouhairCharef.png?size=100" width="72" alt="ZouhairCharef"><br><sub><b>@ZouhairCharef</b></sub></a><br><sub>CVE-2026-34986</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/mfredenhagen"><img src="https://github.com/mfredenhagen.png?size=100" width="72" alt="mfredenhagen"><br><sub><b>@mfredenhagen</b></sub></a><br><sub>CVE-2026-39883</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/pinginfo"><img src="https://github.com/pinginfo.png?size=100" width="72" alt="pinginfo"><br><sub><b>@pinginfo</b></sub></a><br><sub>Notification streaming</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/waterkip"><img src="https://github.com/waterkip.png?size=100" width="72" alt="waterkip"><br><sub><b>@waterkip</b></sub></a><br><sub>Documentation links</sub>
</td>
</tr>
</table>
<p>
<a href="https://github.com/magicxor"><img src="https://github.com/magicxor.png?size=64" width="44" alt="magicxor" title="@magicxor"></a>
<a href="https://github.com/ycjlin"><img src="https://github.com/ycjlin.png?size=64" width="44" alt="ycjlin" title="@ycjlin"></a>
<a href="https://github.com/h5vx"><img src="https://github.com/h5vx.png?size=64" width="44" alt="h5vx" title="@h5vx"></a>
<a href="https://github.com/Dansyuqri"><img src="https://github.com/Dansyuqri.png?size=64" width="44" alt="Dansyuqri" title="@Dansyuqri"></a>
<a href="https://github.com/davinkevin"><img src="https://github.com/davinkevin.png?size=64" width="44" alt="davinkevin" title="@davinkevin"></a>
<a href="https://github.com/lem21h"><img src="https://github.com/lem21h.png?size=64" width="44" alt="lem21h" title="@lem21h"></a>
<a href="https://github.com/sulin37392"><img src="https://github.com/sulin37392.png?size=64" width="44" alt="sulin37392" title="@sulin37392"></a>
<a href="https://github.com/mosesdd"><img src="https://github.com/mosesdd.png?size=64" width="44" alt="mosesdd" title="@mosesdd"></a>
<a href="https://github.com/Xavier-777"><img src="https://github.com/Xavier-777.png?size=64" width="44" alt="Xavier-777" title="@Xavier-777"></a>
<a href="https://github.com/jiadzh"><img src="https://github.com/jiadzh.png?size=64" width="44" alt="jiadzh" title="@jiadzh"></a>
<a href="https://github.com/TLINDEN"><img src="https://github.com/TLINDEN.png?size=64" width="44" alt="TLINDEN" title="@TLINDEN"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://github.com/AntonOfTheWoods.png?size=64" width="44" alt="AntonOfTheWoods" title="@AntonOfTheWoods"></a>
<a href="https://github.com/zylpsrs"><img src="https://github.com/zylpsrs.png?size=64" width="44" alt="zylpsrs" title="@zylpsrs"></a>
<a href="https://github.com/nsanitate"><img src="https://github.com/nsanitate.png?size=64" width="44" alt="nsanitate" title="@nsanitate"></a>
<a href="https://github.com/makinikm"><img src="https://github.com/makinikm.png?size=64" width="44" alt="makinikm" title="@makinikm"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://github.com/spaceg00se-r.png?size=64" width="44" alt="spaceg00se-r" title="@spaceg00se-r"></a>
<a href="https://github.com/heroes1412"><img src="https://github.com/heroes1412.png?size=64" width="44" alt="heroes1412" title="@heroes1412"></a>
<a href="https://github.com/vampywiz17"><img src="https://github.com/vampywiz17.png?size=64" width="44" alt="vampywiz17" title="@vampywiz17"></a>
<a href="https://github.com/chalukyaj"><img src="https://github.com/chalukyaj.png?size=64" width="44" alt="chalukyaj" title="@chalukyaj"></a>
<a href="https://github.com/cbornet"><img src="https://github.com/cbornet.png?size=64" width="44" alt="cbornet" title="@cbornet"></a>
<a href="https://github.com/jvasile"><img src="https://github.com/jvasile.png?size=64" width="44" alt="jvasile" title="@jvasile"></a>
<a href="https://github.com/Kesavaambati"><img src="https://github.com/Kesavaambati.png?size=64" width="44" alt="Kesavaambati" title="@Kesavaambati"></a>
<a href="https://github.com/redfoxfox"><img src="https://github.com/redfoxfox.png?size=64" width="44" alt="redfoxfox" title="@redfoxfox"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://github.com/kuldeep-link11.png?size=64" width="44" alt="kuldeep-link11" title="@kuldeep-link11"></a>
<a href="https://github.com/meesudzu"><img src="https://github.com/meesudzu.png?size=64" width="44" alt="meesudzu" title="@meesudzu"></a>
<a href="https://github.com/pmezhuev"><img src="https://github.com/pmezhuev.png?size=64" width="44" alt="pmezhuev" title="@pmezhuev"></a>
<a href="https://github.com/kh0mka"><img src="https://github.com/kh0mka.png?size=64" width="44" alt="kh0mka" title="@kh0mka"></a>
**41 community contributors** build SILO, Console, mcli, shared packages, and related projects. The list includes maintainers and every human Issue or PR author, ordered by merged PRs, other PRs, then issue reports. Gold rings highlight significant contributions.
<p align="center">
<a href="https://github.com/Vonng"><img src="https://silo.pgsty.com/images/contributors/Vonng.svg" width="60" height="60" alt="@Vonng" title="@Vonng — Maintains SILO, Console, mcli, shared packages, releases, and documentation"></a>
<a href="https://github.com/h5vx"><img src="https://silo.pgsty.com/images/contributors/h5vx.svg" width="60" height="60" alt="@h5vx" title="@h5vx — Implemented per-bucket CORS configuration and enforcement"></a>
<a href="https://github.com/mrjavadseydi"><img src="https://silo.pgsty.com/images/contributors/mrjavadseydi.svg" width="60" height="60" alt="@mrjavadseydi" title="@mrjavadseydi — Fixed effective bucket quota metrics; proposed access-frequency ILM"></a>
<a href="https://github.com/Dansyuqri"><img src="https://silo.pgsty.com/images/contributors/Dansyuqri.svg" width="60" height="60" alt="@Dansyuqri" title="@Dansyuqri — Added ChecksumType to multipart completion responses"></a>
<a href="https://github.com/ycjlin"><img src="https://silo.pgsty.com/images/contributors/ycjlin.svg" width="60" height="60" alt="@ycjlin" title="@ycjlin — Fixed missing-bucket ListObjects semantics"></a>
<a href="https://github.com/pinginfo"><img src="https://silo.pgsty.com/images/contributors/pinginfo.svg" width="60" height="60" alt="@pinginfo" title="@pinginfo — Repaired bucket notification streaming"></a>
<a href="https://github.com/ZouhairCharef"><img src="https://silo.pgsty.com/images/contributors/ZouhairCharef.svg" width="60" height="60" alt="@ZouhairCharef" title="@ZouhairCharef — Patched CVE-2026-34986 in go-jose"></a>
<a href="https://github.com/mfredenhagen"><img src="https://silo.pgsty.com/images/contributors/mfredenhagen.svg" width="60" height="60" alt="@mfredenhagen" title="@mfredenhagen — Patched CVE-2026-39883 in OpenTelemetry"></a>
<a href="https://github.com/waterkip"><img src="https://silo.pgsty.com/images/contributors/waterkip.svg" width="60" height="60" alt="@waterkip" title="@waterkip — Repointed documentation links to the SILO portal"></a>
<a href="https://github.com/mikemikimike"><img src="https://silo.pgsty.com/images/contributors/mikemikimike.svg" width="60" height="60" alt="@mikemikimike" title="@mikemikimike — Contributed the replicated SSE-C plaintext part-size fix"></a>
<a href="https://github.com/metaneutrons"><img src="https://silo.pgsty.com/images/contributors/metaneutrons.svg" width="60" height="60" alt="@metaneutrons" title="@metaneutrons — Reported and proposed explicit-version delete authorization"></a>
<a href="https://github.com/magicxor"><img src="https://silo.pgsty.com/images/contributors/magicxor.svg" width="60" height="60" alt="@magicxor" title="@magicxor — Reported and proposed conditional DELETE support for If-Match"></a>
<a href="https://github.com/davinkevin"><img src="https://silo.pgsty.com/images/contributors/davinkevin.svg" width="60" height="60" alt="@davinkevin" title="@davinkevin — Proposed the distroless container image and dependency automation"></a>
<a href="https://github.com/lem21h"><img src="https://silo.pgsty.com/images/contributors/lem21h.svg" width="48" height="48" alt="@lem21h" title="@lem21h — Proposed robustness and goroutine improvements"></a>
<a href="https://github.com/sulin37392"><img src="https://silo.pgsty.com/images/contributors/sulin37392.svg" width="48" height="48" alt="@sulin37392" title="@sulin37392 — Proposed dependency updates"></a>
<a href="https://github.com/cbornet"><img src="https://silo.pgsty.com/images/contributors/cbornet.svg" width="60" height="60" alt="@cbornet" title="@cbornet — Reported multipart and streaming checksum defects and missing-bucket semantics"></a>
<a href="https://github.com/vampywiz17"><img src="https://silo.pgsty.com/images/contributors/vampywiz17.svg" width="60" height="60" alt="@vampywiz17" title="@vampywiz17 — Reported LDAP TLS and Console login regressions"></a>
<a href="https://github.com/orenyomtov"><img src="https://silo.pgsty.com/images/contributors/orenyomtov.svg" width="60" height="60" alt="@orenyomtov" title="@orenyomtov — Reported the unsigned-header CopyObject cross-object read (SN-2026-011)"></a>
<a href="https://github.com/mumu-lab"><img src="https://silo.pgsty.com/images/contributors/mumu-lab.svg" width="48" height="48" alt="@mumu-lab" title="@mumu-lab — Reported bucket quota metrics reading a deprecated field"></a>
<a href="https://github.com/jvasile"><img src="https://silo.pgsty.com/images/contributors/jvasile.svg" width="48" height="48" alt="@jvasile" title="@jvasile — Reported missing user, group, and defaults in Debian packages"></a>
<a href="https://github.com/pmezhuev"><img src="https://silo.pgsty.com/images/contributors/pmezhuev.svg" width="48" height="48" alt="@pmezhuev" title="@pmezhuev — Reported missing RPM package signatures"></a>
<a href="https://github.com/TLINDEN"><img src="https://silo.pgsty.com/images/contributors/TLINDEN.svg" width="48" height="48" alt="@TLINDEN" title="@TLINDEN — Reported the missing client in release tarballs"></a>
<a href="https://github.com/makinikm"><img src="https://silo.pgsty.com/images/contributors/makinikm.svg" width="48" height="48" alt="@makinikm" title="@makinikm — Reported the missing client in the container image"></a>
<a href="https://github.com/meesudzu"><img src="https://silo.pgsty.com/images/contributors/meesudzu.svg" width="48" height="48" alt="@meesudzu" title="@meesudzu — Requested the migration guide from upstream MinIO"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://silo.pgsty.com/images/contributors/kuldeep-link11.svg" width="48" height="48" alt="@kuldeep-link11" title="@kuldeep-link11 — Reported NATS JWT credentials and target reload issues"></a>
<a href="https://github.com/sargarass"><img src="https://silo.pgsty.com/images/contributors/sargarass.svg" width="48" height="48" alt="@sargarass" title="@sargarass — Reported ListMultipartUploads prefix and pagination semantics"></a>
<a href="https://github.com/liuhaodongliu990-cmyk"><img src="https://silo.pgsty.com/images/contributors/liuhaodongliu990-cmyk.svg" width="48" height="48" alt="@liuhaodongliu990-cmyk" title="@liuhaodongliu990-cmyk — Reported indeterminate progress for prefix downloads"></a>
<a href="https://github.com/Xavier-777"><img src="https://silo.pgsty.com/images/contributors/Xavier-777.svg" width="48" height="48" alt="@Xavier-777" title="@Xavier-777 — Reported Console lifecycle management and file preview gaps"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://silo.pgsty.com/images/contributors/spaceg00se-r.svg" width="48" height="48" alt="@spaceg00se-r" title="@spaceg00se-r — Requested cpuv1 support and reported a workflow token failure"></a>
<a href="https://github.com/kh0mka"><img src="https://silo.pgsty.com/images/contributors/kh0mka.svg" width="48" height="48" alt="@kh0mka" title="@kh0mka — Reported inter-node I/O timeouts in ReadFileStreamHandler"></a>
<a href="https://github.com/bagutzu"><img src="https://silo.pgsty.com/images/contributors/bagutzu.svg" width="48" height="48" alt="@bagutzu" title="@bagutzu — Requested KES-compatible external KMS and OpenBao support"></a>
<a href="https://github.com/DestroyLee"><img src="https://silo.pgsty.com/images/contributors/DestroyLee.svg" width="48" height="48" alt="@DestroyLee" title="@DestroyLee — Reported the missing documentation navigation"></a>
<a href="https://github.com/mosesdd"><img src="https://silo.pgsty.com/images/contributors/mosesdd.svg" width="48" height="48" alt="@mosesdd" title="@mosesdd — Requested a maintained Helm chart"></a>
<a href="https://github.com/zylpsrs"><img src="https://silo.pgsty.com/images/contributors/zylpsrs.svg" width="48" height="48" alt="@zylpsrs" title="@zylpsrs — Reported missing Console tiering and site replication"></a>
<a href="https://github.com/heroes1412"><img src="https://silo.pgsty.com/images/contributors/heroes1412.svg" width="48" height="48" alt="@heroes1412" title="@heroes1412 — Reported the unusable profiling option"></a>
<a href="https://github.com/redfoxfox"><img src="https://silo.pgsty.com/images/contributors/redfoxfox.svg" width="48" height="48" alt="@redfoxfox" title="@redfoxfox — Reported Chinese documentation availability"></a>
<a href="https://github.com/jiadzh"><img src="https://silo.pgsty.com/images/contributors/jiadzh.svg" width="48" height="48" alt="@jiadzh" title="@jiadzh — Requested Windows build guidance"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://silo.pgsty.com/images/contributors/AntonOfTheWoods.svg" width="48" height="48" alt="@AntonOfTheWoods" title="@AntonOfTheWoods — Asked for clarity on Helm chart and operator options"></a>
<a href="https://github.com/chalukyaj"><img src="https://silo.pgsty.com/images/contributors/chalukyaj.svg" width="48" height="48" alt="@chalukyaj" title="@chalukyaj — Proposed making the SILO Operator easier to discover"></a>
<a href="https://github.com/nsanitate"><img src="https://silo.pgsty.com/images/contributors/nsanitate.svg" width="48" height="48" alt="@nsanitate" title="@nsanitate — Proposed CNCF Sandbox governance"></a>
<a href="https://github.com/Kesavaambati"><img src="https://silo.pgsty.com/images/contributors/Kesavaambati.svg" width="48" height="48" alt="@Kesavaambati" title="@Kesavaambati — Asked about community support and image maintenance"></a>
</p>
GitHub does not generate a contributor graph for forks, so [`CONTRIBUTORS.md`](CONTRIBUTORS.md) — not the Insights page — is this project's attribution record. It names everyone alongside the change or report they contributed.
[View the full contribution record](CONTRIBUTORS.md) for each person's proposals, fixes, and reports.
## Background
+52 -45
View File
@@ -35,6 +35,13 @@
> [!NOTE]
> 2026-08-06,本仓库由 `pgsty/minio` 更名为 `pgsty/silo`,默认分支由 `master` 更名为 `main`。以原 MinIO 形态维持的归档构件仍位于归档的 [`minio`](https://github.com/pgsty/silo/tree/minio) 分支,以及截止 [`RELEASE.2026-08-04T00-00-00Z`](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-08-04T00-00-00Z) 的历次发布中。
## 当前发行版与主分支
最新已发布的 Server 仍为 [20260903](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-09-03T13-18-01Z)。
截至 2026-09-13,主分支已合入更新的安全、存储、Console 与共享包改动,但尚未发布新 Server。
准确的已发布/源码边界见 [CHANGELOG.md](CHANGELOG.md) 与[组件版本矩阵](https://silo.pgsty.com/zh/compatibility/versions/),
其中包括 SN-2026-011 修复状态与密码权限迁移要求。
## 概述
上游停止社区发行后,Silo 为开源 MinIO 服务端维护一条持续可用的版本线:构建、软件包、多架构镜像、安全修复与完整 Web 控制台。Pigsty 在生产环境中用它承载 PostgreSQL 备份存储。
@@ -95,53 +102,53 @@ S3 API、`MINIO_*` 环境变量、`minio_*` 指标、`x-minio-*` 头、`/minio/*
## 贡献者
<table>
<tr>
<td align="center" width="150">
<a href="https://github.com/ZouhairCharef"><img src="https://github.com/ZouhairCharef.png?size=100" width="72" alt="ZouhairCharef"><br><sub><b>@ZouhairCharef</b></sub></a><br><sub>CVE-2026-34986</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/mfredenhagen"><img src="https://github.com/mfredenhagen.png?size=100" width="72" alt="mfredenhagen"><br><sub><b>@mfredenhagen</b></sub></a><br><sub>CVE-2026-39883</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/pinginfo"><img src="https://github.com/pinginfo.png?size=100" width="72" alt="pinginfo"><br><sub><b>@pinginfo</b></sub></a><br><sub>桶通知流式输出</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/waterkip"><img src="https://github.com/waterkip.png?size=100" width="72" alt="waterkip"><br><sub><b>@waterkip</b></sub></a><br><sub>文档链接修正</sub>
</td>
</tr>
</table>
<p>
<a href="https://github.com/magicxor"><img src="https://github.com/magicxor.png?size=64" width="44" alt="magicxor" title="@magicxor"></a>
<a href="https://github.com/ycjlin"><img src="https://github.com/ycjlin.png?size=64" width="44" alt="ycjlin" title="@ycjlin"></a>
<a href="https://github.com/h5vx"><img src="https://github.com/h5vx.png?size=64" width="44" alt="h5vx" title="@h5vx"></a>
<a href="https://github.com/Dansyuqri"><img src="https://github.com/Dansyuqri.png?size=64" width="44" alt="Dansyuqri" title="@Dansyuqri"></a>
<a href="https://github.com/davinkevin"><img src="https://github.com/davinkevin.png?size=64" width="44" alt="davinkevin" title="@davinkevin"></a>
<a href="https://github.com/lem21h"><img src="https://github.com/lem21h.png?size=64" width="44" alt="lem21h" title="@lem21h"></a>
<a href="https://github.com/sulin37392"><img src="https://github.com/sulin37392.png?size=64" width="44" alt="sulin37392" title="@sulin37392"></a>
<a href="https://github.com/mosesdd"><img src="https://github.com/mosesdd.png?size=64" width="44" alt="mosesdd" title="@mosesdd"></a>
<a href="https://github.com/Xavier-777"><img src="https://github.com/Xavier-777.png?size=64" width="44" alt="Xavier-777" title="@Xavier-777"></a>
<a href="https://github.com/jiadzh"><img src="https://github.com/jiadzh.png?size=64" width="44" alt="jiadzh" title="@jiadzh"></a>
<a href="https://github.com/TLINDEN"><img src="https://github.com/TLINDEN.png?size=64" width="44" alt="TLINDEN" title="@TLINDEN"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://github.com/AntonOfTheWoods.png?size=64" width="44" alt="AntonOfTheWoods" title="@AntonOfTheWoods"></a>
<a href="https://github.com/zylpsrs"><img src="https://github.com/zylpsrs.png?size=64" width="44" alt="zylpsrs" title="@zylpsrs"></a>
<a href="https://github.com/nsanitate"><img src="https://github.com/nsanitate.png?size=64" width="44" alt="nsanitate" title="@nsanitate"></a>
<a href="https://github.com/makinikm"><img src="https://github.com/makinikm.png?size=64" width="44" alt="makinikm" title="@makinikm"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://github.com/spaceg00se-r.png?size=64" width="44" alt="spaceg00se-r" title="@spaceg00se-r"></a>
<a href="https://github.com/heroes1412"><img src="https://github.com/heroes1412.png?size=64" width="44" alt="heroes1412" title="@heroes1412"></a>
<a href="https://github.com/vampywiz17"><img src="https://github.com/vampywiz17.png?size=64" width="44" alt="vampywiz17" title="@vampywiz17"></a>
<a href="https://github.com/chalukyaj"><img src="https://github.com/chalukyaj.png?size=64" width="44" alt="chalukyaj" title="@chalukyaj"></a>
<a href="https://github.com/cbornet"><img src="https://github.com/cbornet.png?size=64" width="44" alt="cbornet" title="@cbornet"></a>
<a href="https://github.com/jvasile"><img src="https://github.com/jvasile.png?size=64" width="44" alt="jvasile" title="@jvasile"></a>
<a href="https://github.com/Kesavaambati"><img src="https://github.com/Kesavaambati.png?size=64" width="44" alt="Kesavaambati" title="@Kesavaambati"></a>
<a href="https://github.com/redfoxfox"><img src="https://github.com/redfoxfox.png?size=64" width="44" alt="redfoxfox" title="@redfoxfox"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://github.com/kuldeep-link11.png?size=64" width="44" alt="kuldeep-link11" title="@kuldeep-link11"></a>
<a href="https://github.com/meesudzu"><img src="https://github.com/meesudzu.png?size=64" width="44" alt="meesudzu" title="@meesudzu"></a>
<a href="https://github.com/pmezhuev"><img src="https://github.com/pmezhuev.png?size=64" width="44" alt="pmezhuev" title="@pmezhuev"></a>
<a href="https://github.com/kh0mka"><img src="https://github.com/kh0mka.png?size=64" width="44" alt="kh0mka" title="@kh0mka"></a>
**41 位社区贡献者**共同建设 SILO、Console、mcli、公共包与相关项目。名单包含维护者,以及所有提出 Issue 或 PR 的真人作者;按已合并 PR、其他 PR、Issue 报告排序,黄圈标记显著贡献。
<p align="center">
<a href="https://github.com/Vonng"><img src="https://silo.pgsty.com/images/contributors/Vonng.svg" width="60" height="60" alt="@Vonng" title="@Vonng — 维护 SILO、Console、mcli、公共包、发行与文档"></a>
<a href="https://github.com/h5vx"><img src="https://silo.pgsty.com/images/contributors/h5vx.svg" width="60" height="60" alt="@h5vx" title="@h5vx — 实现单桶 CORS 配置与请求执行"></a>
<a href="https://github.com/mrjavadseydi"><img src="https://silo.pgsty.com/images/contributors/mrjavadseydi.svg" width="60" height="60" alt="@mrjavadseydi" title="@mrjavadseydi — 修复有效桶配额指标,并提交按访问频率分层的 ILM 方案"></a>
<a href="https://github.com/Dansyuqri"><img src="https://silo.pgsty.com/images/contributors/Dansyuqri.svg" width="60" height="60" alt="@Dansyuqri" title="@Dansyuqri — 为分片上传完成响应补充 ChecksumType"></a>
<a href="https://github.com/ycjlin"><img src="https://silo.pgsty.com/images/contributors/ycjlin.svg" width="60" height="60" alt="@ycjlin" title="@ycjlin — 修复缺失桶的 ListObjects 语义"></a>
<a href="https://github.com/pinginfo"><img src="https://silo.pgsty.com/images/contributors/pinginfo.svg" width="60" height="60" alt="@pinginfo" title="@pinginfo — 修复桶通知的流式输出"></a>
<a href="https://github.com/ZouhairCharef"><img src="https://silo.pgsty.com/images/contributors/ZouhairCharef.svg" width="60" height="60" alt="@ZouhairCharef" title="@ZouhairCharef — 修复 go-jose 中的 CVE-2026-34986"></a>
<a href="https://github.com/mfredenhagen"><img src="https://silo.pgsty.com/images/contributors/mfredenhagen.svg" width="60" height="60" alt="@mfredenhagen" title="@mfredenhagen — 修复 OpenTelemetry 中的 CVE-2026-39883"></a>
<a href="https://github.com/waterkip"><img src="https://silo.pgsty.com/images/contributors/waterkip.svg" width="60" height="60" alt="@waterkip" title="@waterkip — 将文档链接指向 SILO 门户"></a>
<a href="https://github.com/mikemikimike"><img src="https://silo.pgsty.com/images/contributors/mikemikimike.svg" width="60" height="60" alt="@mikemikimike" title="@mikemikimike — 提交 SSE-C 复制分片明文尺寸修复"></a>
<a href="https://github.com/metaneutrons"><img src="https://silo.pgsty.com/images/contributors/metaneutrons.svg" width="60" height="60" alt="@metaneutrons" title="@metaneutrons — 报告并提交显式版本删除鉴权方案"></a>
<a href="https://github.com/magicxor"><img src="https://silo.pgsty.com/images/contributors/magicxor.svg" width="60" height="60" alt="@magicxor" title="@magicxor — 报告并提交 DELETE If-Match 条件请求支持方案"></a>
<a href="https://github.com/davinkevin"><img src="https://silo.pgsty.com/images/contributors/davinkevin.svg" width="60" height="60" alt="@davinkevin" title="@davinkevin — 提交 distroless 容器镜像与依赖自动更新方案"></a>
<a href="https://github.com/lem21h"><img src="https://silo.pgsty.com/images/contributors/lem21h.svg" width="48" height="48" alt="@lem21h" title="@lem21h — 提交健壮性与 goroutine 改进"></a>
<a href="https://github.com/sulin37392"><img src="https://silo.pgsty.com/images/contributors/sulin37392.svg" width="48" height="48" alt="@sulin37392" title="@sulin37392 — 提交依赖更新"></a>
<a href="https://github.com/cbornet"><img src="https://silo.pgsty.com/images/contributors/cbornet.svg" width="60" height="60" alt="@cbornet" title="@cbornet — 报告分片与流式校验和缺陷及缺失桶语义问题"></a>
<a href="https://github.com/vampywiz17"><img src="https://silo.pgsty.com/images/contributors/vampywiz17.svg" width="60" height="60" alt="@vampywiz17" title="@vampywiz17 — 报告 LDAP TLS 与 Console 登录回归"></a>
<a href="https://github.com/orenyomtov"><img src="https://silo.pgsty.com/images/contributors/orenyomtov.svg" width="60" height="60" alt="@orenyomtov" title="@orenyomtov — 报告未签名头导致的 CopyObject 跨对象读取(SN-2026-011)"></a>
<a href="https://github.com/mumu-lab"><img src="https://silo.pgsty.com/images/contributors/mumu-lab.svg" width="48" height="48" alt="@mumu-lab" title="@mumu-lab — 报告桶配额指标读取已弃用字段的问题"></a>
<a href="https://github.com/jvasile"><img src="https://silo.pgsty.com/images/contributors/jvasile.svg" width="48" height="48" alt="@jvasile" title="@jvasile — 报告 Debian 包缺少用户、用户组与默认配置"></a>
<a href="https://github.com/pmezhuev"><img src="https://silo.pgsty.com/images/contributors/pmezhuev.svg" width="48" height="48" alt="@pmezhuev" title="@pmezhuev — 报告 RPM 包缺少 GPG 签名"></a>
<a href="https://github.com/TLINDEN"><img src="https://silo.pgsty.com/images/contributors/TLINDEN.svg" width="48" height="48" alt="@TLINDEN" title="@TLINDEN — 报告发布压缩包缺少客户端"></a>
<a href="https://github.com/makinikm"><img src="https://silo.pgsty.com/images/contributors/makinikm.svg" width="48" height="48" alt="@makinikm" title="@makinikm — 报告容器镜像缺少客户端"></a>
<a href="https://github.com/meesudzu"><img src="https://silo.pgsty.com/images/contributors/meesudzu.svg" width="48" height="48" alt="@meesudzu" title="@meesudzu — 提出从上游 MinIO 迁移的指南需求"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://silo.pgsty.com/images/contributors/kuldeep-link11.svg" width="48" height="48" alt="@kuldeep-link11" title="@kuldeep-link11 — 报告 NATS JWT 凭据与通知目标重载问题"></a>
<a href="https://github.com/sargarass"><img src="https://silo.pgsty.com/images/contributors/sargarass.svg" width="48" height="48" alt="@sargarass" title="@sargarass — 报告 ListMultipartUploads 前缀与分页语义问题"></a>
<a href="https://github.com/liuhaodongliu990-cmyk"><img src="https://silo.pgsty.com/images/contributors/liuhaodongliu990-cmyk.svg" width="48" height="48" alt="@liuhaodongliu990-cmyk" title="@liuhaodongliu990-cmyk — 报告前缀下载进度显示异常"></a>
<a href="https://github.com/Xavier-777"><img src="https://silo.pgsty.com/images/contributors/Xavier-777.svg" width="48" height="48" alt="@Xavier-777" title="@Xavier-777 — 报告 Console 生命周期管理与文件预览缺失"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://silo.pgsty.com/images/contributors/spaceg00se-r.svg" width="48" height="48" alt="@spaceg00se-r" title="@spaceg00se-r — 提出 cpuv1 支持需求并报告工作流令牌错误"></a>
<a href="https://github.com/kh0mka"><img src="https://silo.pgsty.com/images/contributors/kh0mka.svg" width="48" height="48" alt="@kh0mka" title="@kh0mka — 报告 ReadFileStreamHandler 节点间 I/O 超时"></a>
<a href="https://github.com/bagutzu"><img src="https://silo.pgsty.com/images/contributors/bagutzu.svg" width="48" height="48" alt="@bagutzu" title="@bagutzu — 提出兼容 KES 的外部 KMS 与 OpenBao 支持需求"></a>
<a href="https://github.com/DestroyLee"><img src="https://silo.pgsty.com/images/contributors/DestroyLee.svg" width="48" height="48" alt="@DestroyLee" title="@DestroyLee — 报告文档目录导航缺失"></a>
<a href="https://github.com/mosesdd"><img src="https://silo.pgsty.com/images/contributors/mosesdd.svg" width="48" height="48" alt="@mosesdd" title="@mosesdd — 提出维护 Helm Chart 的需求"></a>
<a href="https://github.com/zylpsrs"><img src="https://silo.pgsty.com/images/contributors/zylpsrs.svg" width="48" height="48" alt="@zylpsrs" title="@zylpsrs — 报告 Console 缺少分层与站点复制"></a>
<a href="https://github.com/heroes1412"><img src="https://silo.pgsty.com/images/contributors/heroes1412.svg" width="48" height="48" alt="@heroes1412" title="@heroes1412 — 报告性能分析选项不可用"></a>
<a href="https://github.com/redfoxfox"><img src="https://silo.pgsty.com/images/contributors/redfoxfox.svg" width="48" height="48" alt="@redfoxfox" title="@redfoxfox — 报告中文文档站点不可用"></a>
<a href="https://github.com/jiadzh"><img src="https://silo.pgsty.com/images/contributors/jiadzh.svg" width="48" height="48" alt="@jiadzh" title="@jiadzh — 提出 Windows 构建指导需求"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://silo.pgsty.com/images/contributors/AntonOfTheWoods.svg" width="48" height="48" alt="@AntonOfTheWoods" title="@AntonOfTheWoods — 提出明确 Helm Chart 与 Operator 选项的需求"></a>
<a href="https://github.com/chalukyaj"><img src="https://silo.pgsty.com/images/contributors/chalukyaj.svg" width="48" height="48" alt="@chalukyaj" title="@chalukyaj — 提出改善 SILO Operator 可发现性的建议"></a>
<a href="https://github.com/nsanitate"><img src="https://silo.pgsty.com/images/contributors/nsanitate.svg" width="48" height="48" alt="@nsanitate" title="@nsanitate — 提出加入 CNCF Sandbox 的治理建议"></a>
<a href="https://github.com/Kesavaambati"><img src="https://silo.pgsty.com/images/contributors/Kesavaambati.svg" width="48" height="48" alt="@Kesavaambati" title="@Kesavaambati — 提出社区支持与容器镜像维护问题"></a>
</p>
GitHub 不为 fork 仓库生成贡献者图表,因此 [`CONTRIBUTORS.md`](CONTRIBUTORS.md)(而非 Insights 页面)才是本项目的署名记录,其中逐一记录了每个人对应的改动或报告。
[查看完整贡献记录](CONTRIBUTORS.md),了解每位贡献者的提案、修复与问题报告。
## 背景
+1 -1
View File
@@ -40,7 +40,7 @@ if [ -n "${MCLI_BIN:-}" ]; then
exit 0
fi
release=${MCLI_RELEASE:-RELEASE.2026-09-03T07-13-05Z}
release=${MCLI_RELEASE:-RELEASE.2026-09-13T00-00-00Z}
version_hyphen=${release#RELEASE.}
package_version=$(printf '%s\n' "${version_hyphen}" | sed -E 's/^([0-9]{4})-([0-9]{2})-([0-9]{2})T([0-9]{2})-([0-9]{2})-([0-9]{2})Z$/\1\2\3\4\5\6.0.0/')
if [ "${package_version}" = "${version_hyphen}" ]; then
@@ -502,6 +502,7 @@
"MINIO_SITE_COMMENT",
"MINIO_SITE_NAME",
"MINIO_SITE_REGION",
"MINIO_SITE_REPLICATION_METADATA_TOMBSTONES",
"MINIO_STORAGE_CLASS_COMMENT",
"MINIO_STORAGE_CLASS_INLINE_BLOCK",
"MINIO_STORAGE_CLASS_OPTIMIZE",
@@ -594,6 +595,7 @@
"x-minio-internal-encrypted-multipart",
"x-minio-internal-encryptedmultipart",
"x-minio-internal-erasure-upgraded",
"x-minio-internal-ilm-atier",
"x-minio-internal-inline-data",
"x-minio-internal-objectlock-legalhold-timestamp",
"x-minio-internal-replica-status",
@@ -616,6 +618,7 @@
"x-minio-internal-transitioned-object",
"x-minio-internal-xyz",
"x-minio-key",
"x-minio-last-modified",
"x-minio-lifecycleconfig-updatedat",
"x-minio-meta",
"x-minio-meta-appid",
@@ -631,6 +634,7 @@
"x-minio-replication-encrypted-multipart",
"x-minio-replication-ready",
"x-minio-replication-reset-status",
"x-minio-replication-server-side-encryption",
"x-minio-replication-server-side-encryption-iv",
"x-minio-replication-server-side-encryption-seal-algorithm",
"x-minio-replication-server-side-encryption-sealed-key",
@@ -825,6 +829,7 @@
"/site-replication/peer/bucket-ops",
"/site-replication/peer/edit",
"/site-replication/peer/iam-item",
"/site-replication/peer/iam-revisions",
"/site-replication/peer/idp-settings",
"/site-replication/peer/join",
"/site-replication/peer/remove",
@@ -873,6 +878,7 @@
"/v2/metrics/cluster",
"/v2/metrics/node",
"/v2/metrics/resource",
"/v3/site-replication/peer/iam-revisions",
"/var/vcap/bosh",
"/verifybinary",
"/version",
@@ -900,6 +906,7 @@
".minio.sys/config/config.json",
".minio.sys/config/hello.txt",
".minio.sys/config/iam/${username}/identity.json",
".minio.sys/config/ilm/access",
".minio.sys/format.json",
".minio.sys/multipart",
".minio.sys/multipart/bucket/object/uploads.json",
@@ -927,6 +934,7 @@
"arn:minio:kms:::this-is-disregarded",
"arn:minio:kms:::xyz-test-key",
"arn:minio:replication:",
"arn:minio:replication::",
"arn:minio:replication::8320b6d18f9032b4700f1f03b50d8d1853de8f22cab86931ee794e12f190852c:destinationbucket",
"arn:minio:replication:::",
"arn:minio:replication:::dest-bucket",
@@ -1058,12 +1066,10 @@
"cmd/metrics.go=\"Version of current MinIO server instance\"",
"cmd/metrics.go=\"minio\"",
"cmd/object-api-utils.go=\".minio.sys\"",
"cmd/object-handlers-common.go=\"X-Minio-Last-Modified\"",
"cmd/object-handlers.go=\"minio-federated\"",
"cmd/object-handlers.go=\"minio.metadata.\"",
"cmd/object-handlers.go=\"minio.versionId\"",
"cmd/object-multipart-handlers.go=\"X-Minio-Replication-Server-Side-Encryption-Iv\"",
"cmd/object-multipart-handlers.go=\"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm\"",
"cmd/object-multipart-handlers.go=\"X-Minio-Replication-Server-Side-Encryption-Sealed-Key\"",
"cmd/replication-trust.go=\"X-Minio-Replication-Encrypted-Multipart\"",
"cmd/replication-trust.go=\"X-Minio-Replication-Server-Side-Encryption-Iv\"",
"cmd/replication-trust.go=\"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm\"",
+2
View File
@@ -115,7 +115,9 @@ func collect(repo string) (manifest, error) {
fset := token.NewFileSet()
for _, rel := range files {
// Investigation artifacts contain synthetic routes and archived configurations.
if rel == "SILO_REBRANDING_MIGRATION.md" ||
strings.HasPrefix(rel, "docs/investigations/") ||
strings.HasPrefix(rel, "buildscripts/rebrand-guard/") ||
strings.HasPrefix(rel, "buildscripts/helm-migration-guard/") {
continue
+11 -9
View File
@@ -36,7 +36,7 @@ for file in \
buildscripts/verify-helm-migration.sh \
Dockerfile.goreleaser \
Dockerfile.distroless \
dockerscripts/download-static-curl.sh \
dockerscripts/build-static-curl.sh \
dockerscripts/docker-entrypoint.sh \
helm/silo/Chart.yaml \
helm/silo/values.yaml \
@@ -101,7 +101,7 @@ require_text Dockerfile.goreleaser "Published checksum drift"
require_text Dockerfile.distroless 'COPY --chmod=0755 silo /usr/bin/silo'
require_text Dockerfile.distroless 'ENTRYPOINT ["/usr/bin/silo"]'
require_text Dockerfile.distroless '"/usr/bin/silo", "healthcheck", "ready"'
require_text dockerscripts/download-static-curl.sh "sha256sum -c"
require_text dockerscripts/build-static-curl.sh "sha256sum -c"
require_text helm/silo/Chart.yaml "name: silo"
require_text helm/silo/values.yaml "repository: pgsty/silo"
require_text helm/silo/templates/deployment.yaml "/usr/bin/docker-entrypoint.sh silo server"
@@ -159,11 +159,13 @@ fi
# The repository and its default branch are pgsty/silo and main. The invariant
# is that the old name is never a live target, not that it is never spoken: the
# READMEs have to name it to explain the rename and to point at the archived
# artifacts, which is the opposite of stranding a reader on it.
# artifacts, which is the opposite of stranding a reader on it. CONTRIBUTORS.md
# also quotes historical issue titles.
#
# So two rules. First, no live URL may resolve to the old repository anywhere,
# READMEs included.
stale_repo_url="$(rg -n -e 'github\.com/pgsty/minio' -e 'hub\.docker\.com/r/pgsty/minio' \
# READMEs and CONTRIBUTORS.md included.
old_repo_pattern='pgsty/minio(\.git)?([^[:alnum:]_.-]|$)'
stale_repo_url="$(rg -n -e "github\.com/${old_repo_pattern}" -e "hub\.docker\.com/r/${old_repo_pattern}" \
--glob '!.git/**' --glob '!dist/**' \
--glob '!SILO_REBRANDING_MIGRATION.md' \
--glob '!buildscripts/rebrand-guard/compat-baseline.json' . |
@@ -175,10 +177,10 @@ fi
# Second, the bare name may only appear where it is deliberate: the pinned
# pre-rebrand image digest in the upgrade test, the two guards that refuse a
# legacy image, and the two READMEs that document the rename and the archived
# minio branch.
repo_guard_allowlist='^(buildscripts/minio-upgrade\.sh|buildscripts/verify-rebrand\.sh|buildscripts/helm-migration-guard/main\.go|README\.md|README_ZH\.md):'
stale_repo="$(rg -n 'pgsty/minio' --glob '!.git/**' --glob '!dist/**' \
# legacy image, the two READMEs that document the rename and the archived
# minio branch, and historical issue titles in CONTRIBUTORS.md.
repo_guard_allowlist='^(buildscripts/minio-upgrade\.sh|buildscripts/verify-rebrand\.sh|buildscripts/helm-migration-guard/main\.go|README\.md|README_ZH\.md|CONTRIBUTORS\.md):'
stale_repo="$(rg -n "${old_repo_pattern}" --glob '!.git/**' --glob '!dist/**' \
--glob '!SILO_REBRANDING_MIGRATION.md' \
--glob '!buildscripts/rebrand-guard/compat-baseline.json' . |
sed 's#^\./##' | grep -Ev "${repo_guard_allowlist}" || true)"
+348
View File
@@ -0,0 +1,348 @@
package cmd
import (
"archive/zip"
"bytes"
"encoding/base64"
"encoding/json"
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
"github.com/minio/mux"
)
func corsAdminRequest(t *testing.T, cred auth.Credentials, method, path string, body []byte) *httptest.ResponseRecorder {
t.Helper()
router := mux.NewRouter()
registerAdminRouter(router, true)
req, err := newTestSignedRequestV4(method, adminPathPrefix+adminAPIVersionPrefix+path,
int64(len(body)), bytes.NewReader(body), cred.AccessKey, cred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("admin %s: %d: %s", path, rec.Code, rec.Body.String())
}
return rec
}
func corsImportReport(t *testing.T, rec *httptest.ResponseRecorder) madmin.BucketMetaImportErrs {
t.Helper()
var rpt madmin.BucketMetaImportErrs
if err := json.Unmarshal(rec.Body.Bytes(), &rpt); err != nil {
t.Fatalf("import report %q: %v", rec.Body.String(), err)
}
return rpt
}
func corsZip(t *testing.T, entries map[string][]byte) []byte {
t.Helper()
var buf bytes.Buffer
zw := zip.NewWriter(&buf)
for name, data := range entries {
w, err := zw.Create(name)
if err != nil {
t.Fatal(err)
}
if _, err = w.Write(data); err != nil {
t.Fatal(err)
}
}
if err := zw.Close(); err != nil {
t.Fatal(err)
}
return buf.Bytes()
}
// corsCorruptedZip builds an archive holding a stored (uncompressed) cors.xml
// whose payload is altered after the checksum is computed, plus the given
// companion entries. The altered document stays well formed, so only the zip
// checksum tells the two apart.
func corsCorruptedZip(t *testing.T, name string, doc []byte, others map[string][]byte) []byte {
t.Helper()
var buf bytes.Buffer
zw := zip.NewWriter(&buf)
w, err := zw.CreateHeader(&zip.FileHeader{Name: name, Method: zip.Store})
if err != nil {
t.Fatal(err)
}
if _, err = w.Write(doc); err != nil {
t.Fatal(err)
}
for other, data := range others {
ow, err := zw.Create(other)
if err != nil {
t.Fatal(err)
}
if _, err = ow.Write(data); err != nil {
t.Fatal(err)
}
}
if err = zw.Close(); err != nil {
t.Fatal(err)
}
raw := buf.Bytes()
at := bytes.Index(raw, []byte("app.example.com"))
if at < 0 {
t.Fatalf("stored CORS payload not found in archive")
}
raw[at] = 'A'
return raw
}
// TestAdminBucketMetadataCORSRoundTrip covers the export/import round trip for
// per-bucket CORS, per-file error reporting for an invalid document, and that
// an archive without cors.xml leaves an existing configuration alone.
func TestAdminBucketMetadataCORSRoundTrip(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, cred auth.Credentials, t *testing.T) {
corsXML := []byte(testSiteReplicationCORSDoc)
if _, err := updateLocalBucketCORSMetadata(t.Context(), obj, bucket, corsXML); err != nil {
t.Fatal(err)
}
// Export must carry the stored document verbatim.
rec := corsAdminRequest(t, cred, http.MethodGet, "/export-bucket-metadata?bucket="+bucket, nil)
archive := rec.Body.Bytes()
zr, err := zip.NewReader(bytes.NewReader(archive), int64(len(archive)))
if err != nil {
t.Fatal(err)
}
var exported []byte
for _, f := range zr.File {
if f.Name != bucket+"/"+bucketCorsConfig {
continue
}
r, err := f.Open()
if err != nil {
t.Fatal(err)
}
exported, err = io.ReadAll(r)
r.Close()
if err != nil {
t.Fatal(err)
}
}
if !bytes.Equal(exported, corsXML) {
t.Fatalf("%s: exported CORS = %q, want %q", instanceType, exported, corsXML)
}
// Drop the configuration: the archive must then omit the entry.
if _, err = updateLocalBucketCORSMetadata(t.Context(), obj, bucket, nil); err != nil {
t.Fatal(err)
}
if _, _, err = globalBucketMetadataSys.GetCorsConfigXML(bucket); err == nil {
t.Fatalf("%s: CORS still present before restore", instanceType)
}
rec = corsAdminRequest(t, cred, http.MethodGet, "/export-bucket-metadata?bucket="+bucket, nil)
empty := rec.Body.Bytes()
zr, err = zip.NewReader(bytes.NewReader(empty), int64(len(empty)))
if err != nil {
t.Fatal(err)
}
for _, f := range zr.File {
if f.Name == bucket+"/"+bucketCorsConfig {
t.Fatalf("%s: export emitted %s for a bucket without CORS", instanceType, f.Name)
}
}
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata", archive)
if st := corsImportReport(t, rec).Buckets[bucket]; !st.Cors.IsSet || st.Cors.Err != "" {
t.Fatalf("%s: import report cors = %+v", instanceType, st.Cors)
}
stored, storedAt, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: restored CORS = %q, err = %v", instanceType, stored, err)
}
created, err := globalBucketMetadataSys.CreatedAt(bucket)
if err != nil {
t.Fatal(err)
}
if !storedAt.After(created) {
t.Fatalf("%s: restored CORS timestamp %v is not after bucket creation %v", instanceType, storedAt, created)
}
// An archive without cors.xml must not remove the configuration.
corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{bucket + "/quota.json": []byte(`{"quota":0}`)}))
if stored, _, err = globalBucketMetadataSys.GetCorsConfigXML(bucket); err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: import without cors.xml changed CORS: %q, err = %v", instanceType, stored, err)
}
// A bucket the import itself creates must still land above its own
// creation time, otherwise CORS replication would drop the restore.
fresh := "cors-import-created-bucket"
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{fresh + "/" + bucketCorsConfig: corsXML}))
if st := corsImportReport(t, rec).Buckets[fresh]; !st.Cors.IsSet || st.Cors.Err != "" {
t.Fatalf("%s: fresh bucket import report cors = %+v", instanceType, st.Cors)
}
freshStored, freshAt, err := globalBucketMetadataSys.GetCorsConfigXML(fresh)
if err != nil || !bytes.Equal(freshStored, corsXML) {
t.Fatalf("%s: fresh bucket CORS = %q, err = %v", instanceType, freshStored, err)
}
freshCreated, err := globalBucketMetadataSys.CreatedAt(fresh)
if err != nil {
t.Fatal(err)
}
if !freshAt.After(freshCreated) {
t.Fatalf("%s: fresh bucket CORS timestamp %v is not after creation %v", instanceType, freshAt, freshCreated)
}
// An invalid document must fail loudly for that bucket and change nothing.
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{bucket + "/" + bucketCorsConfig: []byte("<CORSConfiguration><CORSRule>")}))
if st := corsImportReport(t, rec).Buckets[bucket]; st.Cors.Err == "" {
t.Fatalf("%s: invalid CORS import reported no error: %+v", instanceType, st)
}
if stored, _, err = globalBucketMetadataSys.GetCorsConfigXML(bucket); err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: invalid CORS import changed stored config: %q, err = %v", instanceType, stored, err)
}
// A well formed document carried by a corrupt zip entry must be
// rejected too, leaving the stored document and its timestamp alone
// while the other configs in the same archive still apply.
_, corsAt, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil {
t.Fatal(err)
}
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsCorruptedZip(t, bucket+"/"+bucketCorsConfig, corsXML,
map[string][]byte{bucket + "/quota.json": []byte(`{"quota":4096,"quotatype":"hard"}`)}))
st := corsImportReport(t, rec).Buckets[bucket]
if st.Cors.Err == "" {
t.Fatalf("%s: corrupt CORS entry reported no error: %+v", instanceType, st)
}
if !st.Quota.IsSet || st.Quota.Err != "" {
t.Fatalf("%s: corrupt CORS entry blocked the neighboring quota: %+v", instanceType, st.Quota)
}
stored, storedAt, err = globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil || !bytes.Equal(stored, corsXML) || !storedAt.Equal(corsAt) {
t.Fatalf("%s: corrupt CORS entry changed stored config: %q at %v (was %v), err = %v", instanceType, stored, storedAt, corsAt, err)
}
quota, _, err := globalBucketMetadataSys.GetQuotaConfig(t.Context(), bucket)
if err != nil || quota == nil || quota.Quota != 4096 {
t.Fatalf("%s: neighboring quota not applied: %+v, err = %v", instanceType, quota, err)
}
}})
}
// corsPeerStub is a stand-in site-replication peer. It records every
// SRBucketMeta it is asked to apply and answers with status.
func corsPeerStub(t *testing.T, applied chan<- madmin.SRBucketMeta, status int) *httptest.Server {
t.Helper()
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Method == http.MethodPut && applied != nil {
var item madmin.SRBucketMeta
if err := json.NewDecoder(r.Body).Decode(&item); err != nil {
t.Errorf("decode peer apply: %v", err)
w.WriteHeader(http.StatusBadRequest)
return
}
applied <- item
}
w.WriteHeader(status)
}))
}
// TestAdminBucketMetadataCORSImportReplicatesPastPeerFailure pins that an
// imported CORS document reaches the reachable peers even when the shared
// bucket metadata hook failed against an unreachable one, and that both
// failures are still reported for the bucket.
func TestAdminBucketMetadataCORSImportReplicatesPastPeerFailure(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, cred auth.Credentials, t *testing.T) {
ctx := t.Context()
corsXML := []byte(testSiteReplicationCORSDoc)
healthyApplies := make(chan madmin.SRBucketMeta, 4)
healthy := corsPeerStub(t, healthyApplies, http.StatusOK)
defer healthy.Close()
broken := corsPeerStub(t, nil, http.StatusBadRequest)
defer broken.Close()
// With site replication on, admin requests resolve their token signing
// key through the site replicator account, so it has to exist.
serviceCred, err := auth.CreateCredentials(siteReplicatorSvcAcc, "cors-import-service-secret")
if err != nil {
t.Fatal(err)
}
serviceCred.ParentUser = cred.AccessKey
if _, err = globalIAMSys.store.AddServiceAccount(ctx, serviceCred); err != nil {
t.Fatal(err)
}
defer globalIAMSys.DeleteServiceAccount(ctx, serviceCred.AccessKey, false)
globalSiteReplicatorCred.Set(serviceCred.SecretKey)
defer globalSiteReplicatorCred.Set("")
globalSiteReplicationSys.Lock()
oldEnabled, oldState := globalSiteReplicationSys.enabled, globalSiteReplicationSys.state
globalSiteReplicationSys.enabled = true
globalSiteReplicationSys.state = srState{
Name: "cors-import-test",
ServiceAccountAccessKey: serviceCred.AccessKey,
Peers: map[string]madmin.PeerInfo{
globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()},
"peer-healthy": {Name: "healthy", DeploymentID: "peer-healthy", Endpoint: healthy.URL},
"peer-broken": {Name: "broken", DeploymentID: "peer-broken", Endpoint: broken.URL},
},
}
globalSiteReplicationSys.Unlock()
defer func() {
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled, globalSiteReplicationSys.state = oldEnabled, oldState
globalSiteReplicationSys.Unlock()
}()
rec := corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{
bucket + "/" + bucketCorsConfig: corsXML,
bucket + "/quota.json": []byte(`{"quota":8192,"quotatype":"hard"}`),
}))
st := corsImportReport(t, rec).Buckets[bucket]
if !st.Cors.IsSet || st.Cors.Err != "" {
t.Fatalf("%s: import report cors = %+v", instanceType, st.Cors)
}
stored, storedAt, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: stored CORS = %q, err = %v", instanceType, stored, err)
}
// The reachable peer must have been told about the CORS document,
// carrying exactly the timestamp that was saved locally.
var corsSeen, sharedSeen bool
for range 2 {
select {
case item := <-healthyApplies:
if item.Type != madmin.SRBucketMetaTypeCorsConfig {
sharedSeen = item.Bucket == bucket && item.Quota != nil
continue
}
if item.Bucket != bucket || item.Cors == nil || !item.UpdatedAt.Equal(storedAt) {
t.Fatalf("%s: peer CORS event = %#v, want %s at %v", instanceType, item, bucket, storedAt)
}
payload, decErr := base64.StdEncoding.Strict().DecodeString(*item.Cors)
if decErr != nil || !bytes.Equal(payload, corsXML) {
t.Fatalf("%s: peer CORS payload = %q, err = %v", instanceType, payload, decErr)
}
corsSeen = true
case <-time.After(10 * time.Second):
t.Fatalf("%s: healthy peer received no further events (shared=%v cors=%v)", instanceType, sharedSeen, corsSeen)
}
}
if !sharedSeen || !corsSeen {
t.Fatalf("%s: healthy peer events shared=%v cors=%v, want both", instanceType, sharedSeen, corsSeen)
}
// Both hook failures against the unreachable peer stay reported.
if got := strings.Count(st.Err, "->broken:"); got != 2 {
t.Fatalf("%s: bucket error mentions the broken peer %d times, want 2: %q", instanceType, got, st.Err)
}
}})
}
+93 -10
View File
@@ -76,7 +76,7 @@ func (a adminAPIHandlers) PutBucketQuotaConfigHandler(w http.ResponseWriter, r *
return
}
quotaConfig, err := parseBucketQuota(bucket, data)
_, err = parseBucketQuota(bucket, data)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -94,9 +94,6 @@ func (a adminAPIHandlers) PutBucketQuotaConfigHandler(w http.ResponseWriter, r *
Quota: data,
UpdatedAt: updatedAt,
}
if quotaConfig.Size == 0 && quotaConfig.Quota == 0 {
bucketMeta.Quota = nil
}
// Call site replication hook.
replLogIf(ctx, globalSiteReplicationSys.BucketMetaHook(ctx, bucketMeta))
@@ -417,6 +414,7 @@ func (a adminAPIHandlers) ExportBucketMetadataHandler(w http.ResponseWriter, r *
bucketLifecycleConfig,
bucketSSEConfig,
bucketTaggingConfig,
bucketCorsConfig,
bucketQuotaConfigFile,
objectLockConfig,
bucketVersioningConfig,
@@ -437,7 +435,7 @@ func (a adminAPIHandlers) ExportBucketMetadataHandler(w http.ResponseWriter, r *
writeErrorResponse(ctx, w, exportError(ctx, err, cfgFile, bucket), r.URL)
return
}
configData, err := json.Marshal(config)
configData, err := canonicalBucketPolicy(config)
if err != nil {
writeErrorResponse(ctx, w, exportError(ctx, err, cfgFile, bucket), r.URL)
return
@@ -517,6 +515,19 @@ func (a adminAPIHandlers) ExportBucketMetadataHandler(w http.ResponseWriter, r *
return
}
rawDataFn(bytes.NewReader(configData), cfgPath, len(configData))
case bucketCorsConfig:
// Export the stored document verbatim: GetBucketCors returns
// the bytes exactly as they were PUT, so the archive must
// round-trip them unchanged.
configData, _, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil {
if errors.Is(err, errConfigNotFound) {
continue
}
writeErrorResponse(ctx, w, exportError(ctx, err, cfgFile, bucket), r.URL)
return
}
rawDataFn(bytes.NewReader(configData), cfgPath, len(configData))
case objectLockConfig:
config, _, err := globalBucketMetadataSys.GetObjectLockConfig(bucket)
if err != nil {
@@ -616,6 +627,13 @@ func applyImportedBucketMetadata(dst *BucketMetadata, src BucketMetadata, fields
case bucketQuotaConfigFile:
dst.QuotaConfigJSON = bytes.Clone(src.QuotaConfigJSON)
dst.QuotaConfigUpdatedAt = src.QuotaConfigUpdatedAt
case bucketCorsConfig:
// The import stamps its fields before creating any missing bucket,
// and a CORS event stamped before bucket creation is discarded as
// belonging to an older incarnation, so the imported document takes
// the same monotonic timestamp a local PutBucketCors would assign.
dst.CorsConfigUpdatedAt = localCORSUpdatedAt(*dst, src.CorsConfigUpdatedAt)
dst.CorsConfigXML = bytes.Clone(src.CorsConfigXML)
case objectLockConfig:
dst.ObjectLockConfigXML = bytes.Clone(src.ObjectLockConfigXML)
dst.ObjectLockConfigUpdatedAt = src.ObjectLockConfigUpdatedAt
@@ -645,6 +663,8 @@ func (i *importMetaReport) SetStatus(bucket, fname string, err error) {
st.Tagging = madmin.MetaStatus{IsSet: true, Err: errMsg}
case bucketQuotaConfigFile:
st.Quota = madmin.MetaStatus{IsSet: true, Err: errMsg}
case bucketCorsConfig:
st.Cors = madmin.MetaStatus{IsSet: true, Err: errMsg}
case objectLockConfig:
st.ObjectLock = madmin.MetaStatus{IsSet: true, Err: errMsg}
case bucketVersioningConfig:
@@ -899,7 +919,7 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
continue
}
configData, err := json.Marshal(bucketPolicy)
configData, err := canonicalBucketPolicy(bucketPolicy)
if err != nil {
rpt.SetStatus(bucket, fileName, err)
continue
@@ -1015,6 +1035,32 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
bucketMap[bucket].QuotaConfigUpdatedAt = updatedAt
markImported(bucket, fileName)
rpt.SetStatus(bucket, fileName, nil)
case bucketCorsConfig:
if sz > maxBucketCorsSize {
rpt.SetStatus(bucket, fileName, errors.New(ErrEntityTooLarge.String()))
continue
}
// Read one byte past the declared size: stopping exactly at sz
// leaves archive/zip short of EOF, so it never verifies the entry
// checksum and a corrupt entry carrying well formed XML would be
// stored as a valid document. The extra byte also lets the reader
// reject an entry longer than it declares.
corsData, err := io.ReadAll(io.LimitReader(reader, sz+1))
if err != nil {
rpt.SetStatus(bucket, fileName, err)
continue
}
if err = validateCORSReplicationPayload(corsData); err != nil {
rpt.SetStatus(bucket, fileName, fmt.Errorf("%s (%s)", errorCodes[ErrMalformedXML].Description, err))
continue
}
bucketMap[bucket].CorsConfigXML = corsData
bucketMap[bucket].CorsConfigUpdatedAt = updatedAt
markImported(bucket, fileName)
rpt.SetStatus(bucket, fileName, nil)
}
}
@@ -1032,18 +1078,39 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
continue
}
var merged BucketMetadata
var commitAt time.Time
err := func() error {
lockCtx, unlock, err := lockBucketMetadata(ctx, objectAPI, bucket)
if err != nil {
return err
}
defer unlock()
merged, err = loadBucketMetadataParse(lockCtx, objectAPI, bucket, true)
merged, err = loadBucketMetadataParse(lockCtx, objectAPI, bucket, false)
if err != nil {
return err
}
if err := ensureBucketMetadataCreated(lockCtx, objectAPI, &merged); err != nil {
return err
}
commitAt = UTCNow()
for _, file := range replicatedBucketConfigs {
if _, ok := fields[file]; ok {
commitAt = localBucketConfigUpdatedAt(merged, file, commitAt)
}
}
applyImportedBucketMetadata(&merged, *meta, fields)
return globalBucketMetadataSys.saveMetadata(lockCtx, objectAPI, merged)
for _, file := range replicatedBucketConfigs {
if _, ok := fields[file]; !ok {
continue
}
data, at := replicatedBucketConfig(&merged, file)
payload, _, err := bucketConfigPayload(bucket, file, *data, len(merged.ObjectLockConfigXML) != 0)
if err != nil {
return err
}
*data, *at = payload, commitAt
}
return globalBucketMetadataSys.saveMetadata(lockCtx, objectAPI, &merged)
}()
if err != nil {
rpt.SetStatus(bucket, "", err)
@@ -1051,7 +1118,7 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
}
*meta = merged
globalNotificationSys.LoadBucketMetadata(bgContext(ctx), bucket)
hook := madmin.SRBucketMeta{Bucket: bucket, UpdatedAt: updatedAt}
hook := madmin.SRBucketMeta{Bucket: bucket, UpdatedAt: commitAt}
var hookNeeded bool
if _, ok := fields[bucketQuotaConfigFile]; ok {
hook.Quota = meta.QuotaConfigJSON
@@ -1059,7 +1126,7 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
}
if _, ok := fields[bucketPolicyConfig]; ok {
hook.Policy = meta.PolicyConfigJSON
hookNeeded = true
hookNeeded = hookNeeded || len(hook.Policy) != 0
}
if _, ok := fields[bucketVersioningConfig]; ok {
hook.Versioning = enc(meta.VersioningConfigXML)
@@ -1080,6 +1147,22 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
if hookNeeded {
err = globalSiteReplicationSys.BucketMetaHook(ctx, hook)
}
if _, ok := fields[bucketPolicyConfig]; ok && len(meta.PolicyConfigJSON) == 0 {
// An omitted bulk Policy cannot express deletion.
err = errors.Join(err, globalSiteReplicationSys.BucketMetaHook(ctx, madmin.SRBucketMeta{
Type: madmin.SRBucketMetaTypePolicy, Bucket: bucket, UpdatedAt: commitAt,
}))
}
if _, ok := fields[bucketCorsConfig]; ok {
// CORS carries its own timestamp, so it replicates through the
// dedicated event rather than the shared bucket metadata hook. It
// is announced even when the shared hook failed: the document is
// already committed locally, and a peer that is unreachable for
// one config must not withhold CORS from the reachable ones.
if corsEvent, live := newBucketCORSReplicationEvent(bucket, *meta); live {
err = errors.Join(err, globalSiteReplicationSys.BucketMetaHook(ctx, corsEvent))
}
}
if err != nil {
rpt.SetStatus(bucket, "", err)
continue
+5 -1
View File
@@ -502,11 +502,15 @@ func (a adminAPIHandlers) AddUser(w http.ResponseWriter, r *http.Request) {
}
checkDenyOnly := accessKey == cred.AccessKey
action := policy.Action(policy.CreateUserAdminAction)
if checkDenyOnly {
action = policy.ChangeMyPasswordAdminAction
}
if !globalIAMSys.IsAllowed(policy.Args{
AccountName: cred.AccessKey,
Groups: cred.Groups,
Action: policy.CreateUserAdminAction,
Action: action,
ConditionValues: getConditionValues(r, "", cred),
IsOwner: owner,
Claims: cred.Claims,
+102
View File
@@ -243,6 +243,7 @@ func TestIAMInternalIDPServerSuite(t *testing.T) {
suite.SetUpSuite(c)
suite.TestUserCreate(c)
suite.TestUserPasswordActionAuthorization(c)
suite.TestUserStatusActionAuthorization(c)
suite.TestGroupStatusActionAuthorization(c)
suite.TestUserPolicyEscalationBug(c)
@@ -356,6 +357,106 @@ func (s *TestSuiteIAM) TestUserCreate(c *check) {
}
}
func (s *TestSuiteIAM) TestUserPasswordActionAuthorization(c *check) {
for _, tt := range []struct {
name string
statements string
self bool
other bool
}{
{"readonly", "", true, false},
{"consolereadonly", "", true, false},
{"password grant", `{"Effect":"Allow","Action":"admin:ChangeMyPassword"}`, true, false},
{"legacy CreateUser deny", `{"Effect":"Deny","Action":"admin:CreateUser","Resource":"arn:aws:s3:::*"}`, true, false},
{"password deny", `{"Effect":"Deny","Action":"admin:ChangeMyPassword"}`, false, false},
{"user admin", `{"Effect":"Allow","Action":"admin:CreateUser"}`, true, true},
{"user admin with password deny", `{"Effect":"Allow","Action":"admin:CreateUser"},{"Effect":"Deny","Action":"admin:ChangeMyPassword"}`, false, true},
{"password deny overrides grant", `{"Effect":"Allow","Action":"admin:ChangeMyPassword"},{"Effect":"Deny","Action":"admin:ChangeMyPassword"}`, false, false},
{"wildcard deny", `{"Effect":"Deny","Action":"admin:*"}`, false, false},
} {
c.Run(tt.name, func(t *testing.T) {
c := &check{t, s.serverType}
ctx, cancel := context.WithTimeout(context.Background(), testDefaultTimeout)
defer cancel()
var users []string
policyName := tt.name
defer func() {
for _, user := range users {
if err := s.adm.RemoveUser(ctx, user); err != nil {
c.Errorf("remove test user: %v", err)
}
}
if tt.statements != "" {
if err := s.adm.RemoveCannedPolicy(ctx, policyName); err != nil {
c.Errorf("remove test policy: %v", err)
}
}
}()
createUser := func() (string, string) {
accessKey, secretKey := mustGenerateCredentials(c)
if err := s.adm.SetUser(ctx, accessKey, secretKey, madmin.AccountEnabled); err != nil {
c.Fatalf("create test user: %v", err)
}
users = append(users, accessKey)
return accessKey, secretKey
}
client := func(accessKey, secretKey string) *madmin.AdminClient {
adm, err := madmin.New(s.endpoint, accessKey, secretKey, s.secure)
if err != nil {
c.Fatal(err)
}
adm.SetCustomTransport(s.TestSuiteCommon.client.Transport)
return adm
}
if tt.statements != "" {
policyName = getRandomBucketName()
doc := []byte(`{"Version":"2012-10-17","Statement":[` + tt.statements + `]}`)
if err := s.adm.AddCannedPolicy(ctx, policyName, doc); err != nil {
c.Fatalf("save test policy: %v", err)
}
}
accessKey, secretKey := createUser()
if _, err := s.adm.AttachPolicy(ctx, madmin.PolicyAssociationReq{
User: accessKey, Policies: []string{policyName},
}); err != nil {
c.Fatalf("attach test policy: %v", err)
}
adm := client(accessKey, secretKey)
_, newSecretKey := mustGenerateCredentials(c)
err := adm.SetUser(ctx, accessKey, newSecretKey, madmin.AccountEnabled)
if tt.self {
if err != nil {
c.Fatalf("change own password: %v", err)
}
if _, err = adm.AccountInfo(ctx, madmin.AccountOpts{}); err == nil {
c.Fatal("old password still authenticates")
}
adm = client(accessKey, newSecretKey)
} else if err == nil || madmin.ToErrorResponse(err).Code != "AccessDenied" {
c.Fatalf("self password change: expected AccessDenied, got %v", err)
}
if _, err := adm.AccountInfo(ctx, madmin.AccountOpts{}); err != nil {
c.Fatalf("current password no longer authenticates: %v", err)
}
target, _ := createUser()
newUser, newUserSecret := mustGenerateCredentials(c)
for _, key := range []string{target, newUser} {
err := adm.SetUser(ctx, key, newUserSecret, madmin.AccountEnabled)
if tt.other {
if err != nil {
c.Fatalf("create or update another user: %v", err)
}
if key == newUser {
users = append(users, newUser)
}
} else if err == nil || madmin.ToErrorResponse(err).Code != "AccessDenied" {
c.Fatalf("create or update another user: expected AccessDenied, got %v", err)
}
}
})
}
}
func (s *TestSuiteIAM) TestUserStatusActionAuthorization(c *check) {
ctx, cancel := context.WithTimeout(context.Background(), testDefaultTimeout)
defer cancel()
@@ -946,6 +1047,7 @@ func (s *TestSuiteIAM) TestCannedPolicies(c *check) {
defaultPolicies := []string{
"readwrite",
"readonly",
"consolereadonly",
"writeonly",
"diagnostics",
"consoleAdmin",
+1
View File
@@ -388,6 +388,7 @@ func registerAdminRouter(router *mux.Router, enableConfigOps bool) {
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/join").HandlerFunc(adminMiddleware(adminAPI.SRPeerJoin))
adminRouter.Methods(http.MethodPut).Path(adminVersion+"/site-replication/peer/bucket-ops").HandlerFunc(adminMiddleware(adminAPI.SRPeerBucketOps)).Queries("bucket", "{bucket:.*}").Queries("operation", "{operation:.*}")
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/iam-item").HandlerFunc(adminMiddleware(adminAPI.SRPeerReplicateIAMItem))
adminRouter.Methods(http.MethodGet, http.MethodPut).Path(adminVersion + "/site-replication/peer/iam-revisions").HandlerFunc(adminMiddleware(adminAPI.SRPeerIAMRevisions))
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/bucket-meta").HandlerFunc(adminMiddleware(adminAPI.SRPeerReplicateBucketItem))
adminRouter.Methods(http.MethodGet).Path(adminVersion + "/site-replication/peer/idp-settings").HandlerFunc(adminMiddleware(adminAPI.SRPeerGetIDPSettings))
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/edit").HandlerFunc(adminMiddleware(adminAPI.SiteReplicationEdit))
+5 -1
View File
@@ -1523,10 +1523,14 @@ var errorCodes = errorCodeMap{
Description: "Your Host header is malformed.",
HTTPStatusCode: http.StatusBadRequest,
},
// The stored object cannot be served: a server-side data condition, not a
// successful partial read. Upstream maps it to http.StatusPartialContent
// (since ca6b4773e, 2017), which lets SDKs accept the XML error document
// as object content; SILO deliberately diverges and returns 500.
ErrObjectTampered: {
Code: "XMinioObjectTampered",
Description: errObjectTampered.Error(),
HTTPStatusCode: http.StatusPartialContent,
HTTPStatusCode: http.StatusInternalServerError,
},
ErrSiteReplicationInvalidRequest: {
+4 -1
View File
@@ -212,7 +212,10 @@ func setObjectHeaders(ctx context.Context, w http.ResponseWriter, objInfo Object
}
if rs == nil && opts.PartNumber > 0 {
rs = partNumberToRangeSpec(objInfo, opts.PartNumber)
rs, err = partNumberToRangeSpec(objInfo, opts.PartNumber)
if err != nil {
return err
}
}
// For providing ranged content
+4 -11
View File
@@ -632,18 +632,11 @@ func isReqAuthenticated(ctx context.Context, r *http.Request, region string, sty
return ErrInvalidDigest
}
// Extract either 'X-Amz-Content-Sha256' header or 'X-Amz-Content-Sha256' query parameter (if V4 presigned)
// Do not verify 'X-Amz-Content-Sha256' if skipSHA256.
// Honor the selected header/query checksum, including the header fallback
// for a presigned request. STS separately hashes its body for its signature.
var contentSHA256 []byte
if skipSHA256 := skipContentSha256Cksum(r); !skipSHA256 && isRequestPresignedSignatureV4(r) {
if sha256Sum, ok := r.Form[xhttp.AmzContentSha256]; ok && len(sha256Sum) > 0 {
contentSHA256, err = hex.DecodeString(sha256Sum[0])
if err != nil {
return ErrContentSHA256Mismatch
}
}
} else if _, ok := r.Header[xhttp.AmzContentSha256]; !skipSHA256 && ok {
contentSHA256, err = hex.DecodeString(r.Header.Get(xhttp.AmzContentSha256))
if !skipContentSha256Cksum(r) {
contentSHA256, err = hex.DecodeString(getContentSha256Cksum(r, serviceS3))
if err != nil || len(contentSHA256) == 0 {
return ErrContentSHA256Mismatch
}
-19
View File
@@ -467,25 +467,6 @@ func testAdversarialHealCorsPropagatesNewerEqualValueTimestamp(obj ObjectLayer,
}
}
func TestAdversarialBucketMetadataComparisonIsBase64CaseSensitive(t *testing.T) {
upper := "QQ=="
lower := "qQ=="
upperBytes, err := base64.StdEncoding.Strict().DecodeString(upper)
if err != nil {
t.Fatal(err)
}
lowerBytes, err := base64.StdEncoding.Strict().DecodeString(lower)
if err != nil {
t.Fatal(err)
}
if string(upperBytes) == string(lowerBytes) {
t.Fatal("test inputs unexpectedly decode to the same bytes")
}
if isBucketMetadataEqual(&upper, &lower) {
t.Fatal("different decoded payloads were treated as equal")
}
}
func TestSiteReplicationStatusDetectsCorsTimestampMismatch(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
+111
View File
@@ -0,0 +1,111 @@
// Copyright 2026 PGSTY contributors.
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"errors"
"fmt"
"net/http"
"sync"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
func TestDeleteBucketMetadataLockCancellation(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testDeleteBucketMetadataLockCancellation})
}
func testDeleteBucketMetadataLockCancellation(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
_, unlock, err := lockBucketMetadata(t.Context(), obj, bucket)
if err != nil {
t.Fatal(err)
}
release := sync.OnceFunc(unlock)
defer release()
// Observe DeleteBucket's ACTUAL metadata.lock attempt. Set the hook after
// our own acquisition above so it only trips on the delete.
delAtLock := make(chan struct{})
var once sync.Once
hook := func(b string) {
if b == bucket {
once.Do(func() { close(delAtLock) })
}
}
lockBucketMetadataAcquireHook.Store(&hook)
defer lockBucketMetadataAcquireHook.Store(nil)
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
done := make(chan error, 1)
go func() { done <- obj.DeleteBucket(ctx, bucket, DeleteBucketOptions{Force: true, NoLock: true}) }()
select {
case <-delAtLock:
// Fixed tree: the delete reached metadata.lock and is blocking on the
// lock we hold. Cancel it and confirm it fails WITHOUT deleting, while we
// still hold the lock (release stays deferred until after the checks).
cancel()
if err := <-done; err == nil {
t.Errorf("%s: canceled deletion succeeded", instanceType)
}
if _, err := obj.GetBucketInfo(t.Context(), bucket, BucketOptions{}); err != nil {
t.Errorf("%s: bucket disappeared while metadata.lock was held: %v", instanceType, err)
}
if _, err := readBucketMetadata(t.Context(), obj, bucket); err != nil {
t.Errorf("%s: canceled deletion removed metadata: %v", instanceType, err)
}
case err := <-done:
// Broken tree: the delete finished without ever taking metadata.lock,
// i.e. it did not serialize the destructive operation behind the lock.
t.Errorf("%s: delete bypassed metadata.lock (err=%v)", instanceType, err)
}
}
func TestQueuedMetadataUpdateAfterDelete(t *testing.T) {
defer DetectTestLeak(t)()
for _, expiry := range []bool{false, true} {
t.Run(fmt.Sprintf("expiry=%v", expiry), func(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
previous := newObjectLayerFn()
barrier := &lcMergeBarrier{ObjectLayer: obj, bucket: bucket, mAtLock: make(chan struct{}), mProceed: make(chan struct{})}
setObjectLayer(barrier)
defer setObjectLayer(previous)
release := sync.OnceFunc(func() { close(barrier.mProceed) })
defer release()
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
ctx = context.WithValue(ctx, lcMergeWriterKey{}, "M")
done := make(chan error, 1)
go func() {
if expiry {
done <- globalBucketMetadataSys.UpdateExpiryLCConfig(ctx, bucket, nil, UTCNow())
return
}
_, err := globalBucketMetadataSys.Update(ctx, bucket, bucketTaggingConfig, []byte(`<Tagging><TagSet/></Tagging>`))
done <- err
}()
select {
case <-barrier.mAtLock:
case <-ctx.Done():
t.Fatal("writer did not reach metadata.lock")
}
if err := obj.DeleteBucket(t.Context(), bucket, DeleteBucketOptions{Force: true}); err != nil {
t.Fatal(err)
}
release()
if err := <-done; !isErrBucketNotFound(err) {
t.Errorf("%s: queued update should reject a deleted bucket, got %v", instanceType, err)
}
if _, err := readBucketMetadata(t.Context(), obj, bucket); !errors.Is(err, errConfigNotFound) && !isErrBucketNotFound(err) {
t.Errorf("%s: queued update recreated metadata: %v", instanceType, err)
}
}})
})
}
}
+278
View File
@@ -0,0 +1,278 @@
// Copyright (c) 2015-2026 MinIO, Inc.
// Copyright (c) 2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
package cmd
import (
"context"
"encoding/base64"
"errors"
"net/http"
"sync"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
// ---------------------------------------------------------------------------
// Target 1 (issue #105): a higher-level lifecycle XML merge racing another
// bucket-metadata transition outside BucketMetadataSys.Delete.
//
// Before the fix, PeerBucketLCConfigHandler / healBucketILMExpiry read the
// current lifecycle document without metadata.lock, merged the replicated expiry
// rules with the local transition rules, and then persisted the merged blob with
// BucketMetadataSys.Update. Update re-read the record under metadata.lock but
// overwrote LifecycleConfigXML wholesale with the pre-computed blob, so any
// lifecycle transition change committed between the merge read and the merge
// write was silently lost on disk. BucketMetadataSys.UpdateExpiryLCConfig now
// performs the read, merge, and save under a single metadata.lock. This test
// drives PeerBucketLCConfigHandler and asserts the concurrent change survives.
// ---------------------------------------------------------------------------
type lcMergeWriterKey struct{}
// lcMergeBarrier pauses the merge writer (context value "M") exactly when it
// tries to take metadata.lock for its persisting Update. By that point the
// merge has already read the stale lifecycle document, so the test can commit a
// concurrent transition change before releasing the merge write.
type lcMergeBarrier struct {
ObjectLayer
bucket string
mAtLock chan struct{}
mProceed chan struct{}
mOnce sync.Once
}
func (o *lcMergeBarrier) metadataLock() string {
return pathJoin(bucketMetaPrefix, o.bucket, "metadata.lock")
}
func (o *lcMergeBarrier) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket != minioMetaBucket || len(objects) != 1 || objects[0] != o.metadataLock() {
return lock
}
return metadataObservedRWLocker{RWLocker: lock, onLock: func(ctx context.Context) {
if ctx.Value(lcMergeWriterKey{}) != "M" {
return
}
o.mOnce.Do(func() { close(o.mAtLock) })
select {
case <-o.mProceed:
case <-ctx.Done():
}
}}
}
func TestLifecycleExpiryMergeRaceLosesConcurrentTransition(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testLifecycleExpiryMergeRaceLosesConcurrentTransition,
})
}
func testLifecycleExpiryMergeRaceLosesConcurrentTransition(obj ObjectLayer, instanceType, bucket string,
_ http.Handler, _ auth.Credentials, t *testing.T,
) {
// Revision N: a single transition-only rule.
baseXML := []byte(`<LifecycleConfiguration><Rule><ID>keep</ID><Filter><Prefix>data/</Prefix></Filter><Status>Enabled</Status><Transition><Days>30</Days><StorageClass>WARM</StorageClass></Transition></Rule></LifecycleConfiguration>`)
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketLifecycleConfig, baseXML); err != nil {
t.Fatalf("%s: seed lifecycle: %v", instanceType, err)
}
// A replicated expiry-only rule arriving from a peer site.
expXML := `<LifecycleConfiguration><Rule><ID>expire</ID><Filter><Prefix>tmp/</Prefix></Filter><Status>Enabled</Status><Expiration><Days>7</Days></Expiration></Rule></LifecycleConfiguration>`
expLCConfig := base64.StdEncoding.EncodeToString([]byte(expXML))
previousObjectAPI := newObjectLayerFn()
barrier := &lcMergeBarrier{
ObjectLayer: obj,
bucket: bucket,
mAtLock: make(chan struct{}),
mProceed: make(chan struct{}),
}
setObjectLayer(barrier)
defer setObjectLayer(previousObjectAPI)
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
mCtx := context.WithValue(ctx, lcMergeWriterKey{}, "M")
mDone := make(chan error, 1)
go func() {
mDone <- globalSiteReplicationSys.PeerBucketLCConfigHandler(mCtx, bucket, &expLCConfig, UTCNow())
}()
// Wait until the merge writer has read revision N and is about to persist.
select {
case <-barrier.mAtLock:
case err := <-mDone:
t.Fatalf("%s: merge writer finished before persisting: %v", instanceType, err)
case <-ctx.Done():
t.Fatalf("%s: merge writer never reached metadata.lock: %v", instanceType, ctx.Err())
}
// Concurrent local lifecycle transition change commits revision N+1.
concurrentXML := []byte(`<LifecycleConfiguration><Rule><ID>keep</ID><Filter><Prefix>data/</Prefix></Filter><Status>Enabled</Status><Transition><Days>10</Days><StorageClass>COLD</StorageClass></Transition></Rule></LifecycleConfiguration>`)
if _, err := globalBucketMetadataSys.Update(ctx, bucket, bucketLifecycleConfig, concurrentXML); err != nil {
t.Fatalf("%s: concurrent transition update: %v", instanceType, err)
}
// Release the merge write so it lands after the concurrent commit.
close(barrier.mProceed)
if err := <-mDone; err != nil {
t.Fatalf("%s: merge writer failed: %v", instanceType, err)
}
// The persisted lifecycle must contain both the replicated expiry rule and
// the concurrent transition change.
cfg, _, err := globalBucketMetadataSys.GetLifecycleConfig(bucket)
if err != nil {
t.Fatalf("%s: read merged lifecycle: %v", instanceType, err)
}
var keep, expire bool
for i := range cfg.Rules {
switch cfg.Rules[i].ID {
case "keep":
keep = true
if cfg.Rules[i].Transition.Days != 10 || cfg.Rules[i].Transition.StorageClass != "COLD" {
t.Fatalf("%s: lifecycle merge overwrote the concurrent transition change: got Days=%d StorageClass=%q, want Days=10 StorageClass=COLD",
instanceType, cfg.Rules[i].Transition.Days, cfg.Rules[i].Transition.StorageClass)
}
case "expire":
expire = true
}
}
if !expire {
t.Fatalf("%s: merged lifecycle dropped the replicated expiry rule: %+v", instanceType, cfg.Rules)
}
if !keep {
t.Fatalf("%s: merged lifecycle dropped the transition rule entirely: %+v", instanceType, cfg.Rules)
}
}
// ---------------------------------------------------------------------------
// Target 2 (issue #105): DeleteBucket racing an in-flight metadata writer must
// not resurrect a ghost .metadata.bin record.
//
// Before the fix, erasureServerPools.DeleteBucket took only <bucket>.lck and
// purged the metadata prefix while config writers (updateAndParse) took only
// metadata.lock, so a writer already MID-SAVE (holding metadata.lock, past
// saveMetadata's existence recheck) could persist .metadata.bin after the purge.
// DeleteBucket now takes metadata.lock before deleting, so it waits for that
// writer and then purges whatever the writer wrote.
//
// This test isolates the DeleteBucket-lock fix specifically: the writer holds
// metadata.lock and is paused at the .metadata.bin PutObject, so the earlier
// saveMetadata existence recheck cannot save it — only serializing the delete
// behind the writer can. Removing just DeleteBucket's metadata.lock (keeping the
// recheck) therefore makes this test fail. The delete's ACTUAL metadata.lock
// attempt is observed with lockBucketMetadataAcquireHook (its lock is taken
// through the erasureServerPools receiver, invisible to the object-layer
// barrier), so the handshake is deterministic with no timing assumption.
// ---------------------------------------------------------------------------
func TestDeleteBucketResurrectsGhostMetadata(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testDeleteBucketResurrectsGhostMetadata,
})
}
func testDeleteBucketResurrectsGhostMetadata(obj ObjectLayer, instanceType, bucket string,
_ http.Handler, _ auth.Credentials, t *testing.T,
) {
previousObjectAPI := newObjectLayerFn()
// Writer A holds metadata.lock and pauses at the .metadata.bin PutObject,
// i.e. already past saveMetadata's existence recheck and mid-save.
barrier := &metadataRMWBarrierObjectLayer{
ObjectLayer: obj,
bucket: bucket,
aReady: make(chan struct{}),
aRelease: make(chan struct{}),
bLockAttempt: make(chan struct{}),
}
setObjectLayer(barrier)
defer setObjectLayer(previousObjectAPI)
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
aCtx := context.WithValue(ctx, metadataRMWWriterKey{}, "A")
policyJSON := []byte(`{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":"*","Action":"s3:GetObject","Resource":"arn:aws:s3:::` + bucket + `/*"}]}`)
aReleased := sync.OnceFunc(func() { close(barrier.aRelease) })
defer aReleased()
aDone := make(chan error, 1)
go func() {
_, err := globalBucketMetadataSys.Update(aCtx, bucket, bucketPolicyConfig, policyJSON)
aDone <- err
}()
select {
case <-barrier.aReady:
case err := <-aDone:
t.Fatalf("%s: writer A finished before persisting: %v", instanceType, err)
case <-ctx.Done():
t.Fatalf("%s: writer A never reached metadata save: %v", instanceType, ctx.Err())
}
// A now holds metadata.lock mid-save. Observe DeleteBucket's ACTUAL
// metadata.lock attempt via the acquire hook, set only now so A's earlier
// acquisition does not trip it.
delAtLock := make(chan struct{})
var once sync.Once
hook := func(b string) {
if b == bucket {
once.Do(func() { close(delAtLock) })
}
}
lockBucketMetadataAcquireHook.Store(&hook)
defer lockBucketMetadataAcquireHook.Store(nil)
delDone := make(chan error, 1)
go func() {
delDone <- obj.DeleteBucket(ctx, bucket, DeleteBucketOptions{Force: true})
}()
select {
case <-delAtLock:
// Fixed tree: DeleteBucket reached metadata.lock and blocks on A. Release
// A so it finishes its save and unlocks; the delete then acquires the
// lock and purges the record A wrote.
aReleased()
if err := <-aDone; err != nil {
t.Fatalf("%s: writer A save failed while holding metadata.lock: %v", instanceType, err)
}
if err := <-delDone; err != nil {
t.Fatalf("%s: delete bucket: %v", instanceType, err)
}
case err := <-delDone:
// Broken tree: DeleteBucket purged without taking metadata.lock. Release
// A so its mid-save PutObject recreates .metadata.bin (the ghost).
if err != nil {
t.Fatalf("%s: delete bucket: %v", instanceType, err)
}
aReleased()
if err := <-aDone; err != nil {
t.Fatalf("%s: writer A save failed: %v", instanceType, err)
}
}
// After DeleteBucket, no .metadata.bin record may remain on disk.
if meta, err := readBucketMetadata(ctx, obj, bucket); err == nil {
t.Fatalf("%s: ghost .metadata.bin resurrected after DeleteBucket: name=%q created=%s policyLen=%d",
instanceType, meta.Name, meta.Created, len(meta.PolicyConfigJSON))
} else if !errors.Is(err, errConfigNotFound) && !isErrBucketNotFound(err) && !errors.Is(err, errVolumeNotFound) {
t.Fatalf("%s: unexpected error reading deleted bucket metadata: %v", instanceType, err)
}
}
+54
View File
@@ -0,0 +1,54 @@
// Copyright 2026 PGSTY contributors.
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"net/http"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/grid"
)
func TestPeerMetadataReloadWithEqualMaximumTimestamp(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testPeerMetadataReloadWithEqualMaximumTimestamp})
}
func testPeerMetadataReloadWithEqualMaximumTimestamp(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
disk, err := loadBucketMetadata(t.Context(), obj, bucket)
if err != nil {
t.Fatal(err)
}
disk.TaggingConfigXML = []byte(`<Tagging><TagSet><Tag><Key>revision</Key><Value>new</Value></Tag></TagSet></Tagging>`)
disk.TaggingConfigUpdatedAt = UTCNow()
// An unrelated configuration has the greatest timestamp in both records.
disk.PolicyConfigUpdatedAt = disk.TaggingConfigUpdatedAt.Add(time.Hour)
if err := disk.Save(t.Context(), obj); err != nil {
t.Fatal(err)
}
resident := disk
resident.TaggingConfigXML = []byte(`<Tagging><TagSet><Tag><Key>revision</Key><Value>old</Value></Tag></TagSet></Tagging>`)
resident.TaggingConfigUpdatedAt = disk.TaggingConfigUpdatedAt.Add(-time.Minute)
if err := resident.parseAllConfigs(t.Context(), obj); err != nil {
t.Fatal(err)
}
if !resident.lastUpdate().Equal(disk.lastUpdate()) {
t.Fatal("fixture must have equal maximum timestamps")
}
globalBucketMetadataSys.Set(bucket, resident)
args := grid.MSS{peerRESTBucket: bucket}
if _, err := (&peerRESTServer{}).LoadBucketMetadataHandler(&args); err != nil {
t.Fatal(err)
}
tagging, _, err := globalBucketMetadataSys.GetTaggingConfig(bucket)
if err != nil {
t.Fatal(err)
}
if got := tagging.ToMap()["revision"]; got != "new" {
t.Errorf("%s: peer reload retained tag %q despite a newer tagging configuration", instanceType, got)
}
}
+330
View File
@@ -0,0 +1,330 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful,
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"encoding/json"
"errors"
"sort"
"time"
"github.com/minio/minio-go/v7/pkg/tags"
bucketsse "github.com/minio/minio/internal/bucket/encryption"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
"github.com/minio/minio/internal/bucket/versioning"
"github.com/pgsty/silo-pkg/v3/policy"
)
// Only these fields share the site-replication source-time ordering contract.
// Bulk apply/import process Object Lock before Versioning, whose effective
// document depends on it. Periodic heal retains its existing type order.
var replicatedBucketConfigs = [...]string{
objectLockConfig, bucketVersioningConfig, bucketPolicyConfig,
bucketTaggingConfig, bucketSSEConfig, bucketQuotaConfigFile,
}
func replicatedBucketConfig(meta *BucketMetadata, file string) (*[]byte, *time.Time) {
switch file {
case bucketPolicyConfig:
return &meta.PolicyConfigJSON, &meta.PolicyConfigUpdatedAt
case bucketTaggingConfig:
return &meta.TaggingConfigXML, &meta.TaggingConfigUpdatedAt
case bucketSSEConfig:
return &meta.EncryptionConfigXML, &meta.EncryptionConfigUpdatedAt
case bucketQuotaConfigFile:
return &meta.QuotaConfigJSON, &meta.QuotaConfigUpdatedAt
case bucketVersioningConfig:
return &meta.VersioningConfigXML, &meta.VersioningConfigUpdatedAt
case objectLockConfig:
return &meta.ObjectLockConfigXML, &meta.ObjectLockConfigUpdatedAt
}
return nil, nil
}
// Callers that only need to know whether a file is under the contract must not
// probe replicatedBucketConfig with a throwaway BucketMetadata.
func isReplicatedBucketConfig(file string) bool {
for _, replicated := range replicatedBucketConfigs {
if replicated == file {
return true
}
}
return false
}
func bucketConfigUpdateOnly(file string) bool {
return file == bucketVersioningConfig || file == objectLockConfig
}
// Reuse the persistence rule before comparison, so an accepted Versioning
// event and the document Save actually writes have the same comparison key.
func effectiveBucketVersioning(data []byte, lockEnabled bool) []byte {
if lockEnabled {
config, err := versioning.ParseConfig(bytes.NewReader(data))
if err != nil || !config.Enabled() || config.PrefixesExcluded() {
return enabledBucketVersioningConfig
}
}
return data
}
// A parsed policy still contains map-backed sets with nondeterministic Marshal
// order. Sort every set array recursively, including statements and conditions.
// RawMessage keeps integer values intact; decoding through float64 would not.
func canonicalBucketPolicyJSON(data json.RawMessage) (json.RawMessage, error) {
data = bytes.TrimSpace(data)
if len(data) == 0 {
return nil, nil
}
switch data[0] {
case '{':
var obj map[string]json.RawMessage
if err := json.Unmarshal(data, &obj); err != nil {
return nil, err
}
for key, value := range obj {
var err error
obj[key], err = canonicalBucketPolicyJSON(value)
if err != nil {
return nil, err
}
}
return json.Marshal(obj)
case '[':
var arr []json.RawMessage
if err := json.Unmarshal(data, &arr); err != nil {
return nil, err
}
for i := range arr {
var err error
arr[i], err = canonicalBucketPolicyJSON(arr[i])
if err != nil {
return nil, err
}
}
sort.Slice(arr, func(i, j int) bool { return bytes.Compare(arr[i], arr[j]) < 0 })
return json.Marshal(arr)
default:
var compact bytes.Buffer
if err := json.Compact(&compact, data); err != nil {
return nil, err
}
return compact.Bytes(), nil
}
}
// Encode the validated policy fields explicitly: BPStatement's required
// Action/Resource tags otherwise try to marshal empty sets for the supported
// NotAction/NotResource alternatives. This stays within the existing schema.
func canonicalBucketPolicy(cfg *policy.BucketPolicy) ([]byte, error) {
if cfg.IsEmpty() {
return nil, nil
}
doc := map[string]any{"Version": cfg.Version}
if cfg.ID != "" {
doc["ID"] = cfg.ID
}
statements := make([]map[string]any, 0, len(cfg.Statements))
for _, st := range cfg.Statements {
statement := map[string]any{"Effect": st.Effect, "Principal": st.Principal}
if st.SID != "" {
statement["Sid"] = st.SID
}
if len(st.Actions) != 0 {
statement["Action"] = st.Actions
}
if len(st.NotActions) != 0 {
statement["NotAction"] = st.NotActions
}
if len(st.Resources) != 0 {
statement["Resource"] = st.Resources
}
if len(st.NotResources) != 0 {
statement["NotResource"] = st.NotResources
}
if len(st.Conditions) != 0 {
statement["Condition"] = st.Conditions
}
statements = append(statements, statement)
}
doc["Statement"] = statements
data, err := json.Marshal(doc)
if err != nil {
return nil, err
}
return canonicalBucketPolicyJSON(data)
}
// Validate with the same parsers as Save. Policy's established empty-policy
// semantics are deletion; a parsed zero quota is still a live document.
func bucketConfigPayload(bucket, file string, data []byte, lockEnabled bool) ([]byte, []byte, error) {
if len(data) == 0 {
return nil, nil, nil
}
var err error
key := data
switch file {
case bucketPolicyConfig:
var cfg *policy.BucketPolicy
cfg, err = policy.ParseBucketPolicyConfig(bytes.NewReader(data), bucket)
if err == nil {
if cfg.IsEmpty() {
return nil, nil, nil
}
key, err = canonicalBucketPolicy(cfg)
}
case bucketQuotaConfigFile:
cfg, parseErr := parseBucketQuota(bucket, data)
err = parseErr
if err == nil {
key, err = json.Marshal(cfg)
}
case bucketTaggingConfig:
_, err = tags.ParseBucketXML(bytes.NewReader(data))
case bucketSSEConfig:
_, err = bucketsse.ParseBucketSSEConfig(bytes.NewReader(data))
case objectLockConfig:
_, err = objectlock.ParseObjectLockConfig(bytes.NewReader(data))
case bucketVersioningConfig:
data = effectiveBucketVersioning(data, lockEnabled)
key = data
_, err = versioning.ParseConfig(bytes.NewReader(data))
}
return data, key, err
}
type bucketConfigState struct {
data, key []byte
at time.Time
real, valid bool
}
func newBucketConfigState(bucket, file string, data []byte, at, created time.Time, lockEnabled bool) (bucketConfigState, error) {
data, key, err := bucketConfigPayload(bucket, file, data, lockEnabled)
if err != nil {
return bucketConfigState{}, err
}
if at.IsZero() {
at = created
}
valid := !created.IsZero() && !at.Before(created)
modified := valid && at.After(created)
if bucketConfigUpdateOnly(file) && len(data) == 0 {
modified = false
}
return bucketConfigState{data: data, key: key, at: at, real: modified, valid: valid}, nil
}
func (s bucketConfigState) candidate() bool {
return s.valid && (s.real || len(s.data) != 0)
}
func compareBucketConfigStates(a, b bucketConfigState) int {
if a.valid != b.valid {
if a.valid {
return 1
}
return -1
}
if a.real != b.real {
if a.real {
return 1
}
return -1
}
if a.real {
if n := a.at.Compare(b.at); n != 0 {
return n
}
// At equal source time a real deletion wins, preventing resurrection.
if (len(a.data) == 0) != (len(b.data) == 0) {
if len(a.data) == 0 {
return 1
}
return -1
}
}
return bytes.Compare(a.key, b.key)
}
func localBucketConfigUpdatedAt(meta BucketMetadata, file string, now time.Time) time.Time {
_, at := replicatedBucketConfig(&meta, file)
for _, lower := range []time.Time{meta.Created, *at} {
if !now.After(lower) {
now = lower.Add(time.Nanosecond)
}
}
return now.UTC()
}
func ensureBucketMetadataCreated(ctx context.Context, obj ObjectLayer, meta *BucketMetadata) error {
if !meta.Created.IsZero() {
return nil
}
info, err := obj.GetBucketInfo(ctx, meta.Name, BucketOptions{NoMetadata: true})
if err != nil {
return err
}
if info.Created.IsZero() {
return errors.New("bucket metadata creation time is unknown")
}
meta.Created = info.Created.UTC()
return nil
}
// applyBucketConfig runs under metadata.lock, on freshly loaded metadata. It
// changes only the selected field; the caller persists once after all checks.
func applyBucketConfig(meta *BucketMetadata, file string, data []byte, at time.Time) (bool, error) {
if bucketConfigUpdateOnly(file) && len(data) == 0 {
return false, nil
}
current, currentAt := replicatedBucketConfig(meta, file)
lockEnabled := len(meta.ObjectLockConfigXML) != 0
incoming, err := newBucketConfigState(meta.Name, file, data, at, meta.Created, lockEnabled)
if err != nil {
return false, err
}
if !incoming.candidate() {
return false, nil
}
local, err := newBucketConfigState(meta.Name, file, *current, *currentAt, meta.Created, lockEnabled)
if err != nil {
return false, err
}
if compareBucketConfigStates(incoming, local) <= 0 {
return false, nil
}
*current, *currentAt = bytes.Clone(incoming.data), incoming.at.UTC()
return true, nil
}
func rebaseBucketConfigDefaults(meta *BucketMetadata, oldCreated time.Time) {
if meta.Created.Equal(oldCreated) {
return
}
// These six fields alone use Created to distinguish a baseline from a
// tombstone. Preserve actual source times when adopting an existing bucket.
for _, file := range replicatedBucketConfigs {
_, at := replicatedBucketConfig(meta, file)
if at.IsZero() || at.Equal(oldCreated) {
*at = meta.Created
}
}
}
+160 -51
View File
@@ -24,6 +24,7 @@ import (
"fmt"
"math/rand"
"sync"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
@@ -125,19 +126,42 @@ func (sys *BucketMetadataSys) Set(bucket string, meta BucketMetadata) {
}
}
func (sys *BucketMetadataSys) updateAndParse(ctx context.Context, bucket string, configFile string, configData []byte, parse, lifecycleDelete bool) (updatedAt time.Time, err error) {
// bucketMetadataUpdate returns the committed snapshot to the caller. meta and
// updatedAt hold the saved state only when changed is true. Local writes always
// change state, because localBucketConfigUpdatedAt is strictly greater than the
// current field time, so their handlers can broadcast meta without rechecking.
type bucketMetadataUpdate struct {
meta BucketMetadata
updatedAt time.Time
changed bool
}
func (sys *BucketMetadataSys) updateAndParse(ctx context.Context, bucket, configFile string, configData []byte, parse, lifecycleDelete bool) (time.Time, error) {
result, err := sys.updateAndParseMetadata(ctx, bucket, configFile, configData, parse, lifecycleDelete, nil)
return result.updatedAt, err
}
func (sys *BucketMetadataSys) updateAndParseMetadata(ctx context.Context, bucket string, configFile string, configData []byte, parse, lifecycleDelete bool, sourceTime *time.Time) (result bucketMetadataUpdate, err error) {
objAPI := newObjectLayerFn()
if objAPI == nil {
return updatedAt, errServerNotInitialized
return result, errServerNotInitialized
}
if isMinioMetaBucketName(bucket) {
return updatedAt, errInvalidArgument
return result, errInvalidArgument
}
// Load deletions without parsed caches (notably quota), and compare the
// six replicated fields against the raw document under the same lock.
if isReplicatedBucketConfig(configFile) {
parse = false
if bucketConfigUpdateOnly(configFile) && len(configData) == 0 {
return result, nil
}
}
notifyCtx := ctx
ctx, unlock, err := lockBucketMetadata(ctx, objAPI, bucket)
if err != nil {
return updatedAt, err
return result, err
}
err = func() error {
@@ -157,55 +181,73 @@ func (sys *BucketMetadataSys) updateAndParse(ctx context.Context, bucket string,
return err
}
}
updatedAt = UTCNow()
switch configFile {
case bucketPolicyConfig:
meta.PolicyConfigJSON = configData
meta.PolicyConfigUpdatedAt = updatedAt
case bucketNotificationConfig:
meta.NotificationConfigXML = configData
meta.NotificationConfigUpdatedAt = updatedAt
case bucketLifecycleConfig:
meta.LifecycleConfigXML = configData
meta.LifecycleConfigUpdatedAt = updatedAt
case bucketSSEConfig:
meta.EncryptionConfigXML = configData
meta.EncryptionConfigUpdatedAt = updatedAt
case bucketTaggingConfig:
meta.TaggingConfigXML = configData
meta.TaggingConfigUpdatedAt = updatedAt
case bucketQuotaConfigFile:
meta.QuotaConfigJSON = configData
meta.QuotaConfigUpdatedAt = updatedAt
case objectLockConfig:
meta.ObjectLockConfigXML = configData
meta.ObjectLockConfigUpdatedAt = updatedAt
case bucketVersioningConfig:
meta.VersioningConfigXML = configData
meta.VersioningConfigUpdatedAt = updatedAt
case bucketReplicationConfig:
meta.ReplicationConfigXML = configData
meta.ReplicationConfigUpdatedAt = updatedAt
case bucketTargetsFile:
meta.BucketTargetsConfigJSON, meta.BucketTargetsConfigMetaJSON, err = encryptBucketMetadata(ctx, meta.Name, configData, kms.Context{
bucket: meta.Name,
bucketTargetsFile: bucketTargetsFile,
})
if err != nil {
return fmt.Errorf("Error encrypting bucket target metadata %w", err)
updatedAt := UTCNow()
if isReplicatedBucketConfig(configFile) {
if err := ensureBucketMetadataCreated(ctx, objAPI, &meta); err != nil {
var at time.Time
if sourceTime != nil {
at = *sourceTime
}
logBucketConfigReplication(ctx, bucket, configFile, "indeterminate", at, meta.Created, err.Error())
return err
}
if sourceTime == nil || sourceTime.IsZero() {
updatedAt = localBucketConfigUpdatedAt(meta, configFile, updatedAt)
if sourceTime != nil {
logBucketConfigReplication(ctx, bucket, configFile, "legacy-zero", *sourceTime, meta.Created, "assigned local source time")
}
} else {
updatedAt = sourceTime.UTC()
}
if updatedAt.Before(meta.Created) {
logBucketConfigReplication(ctx, bucket, configFile, "before-created", updatedAt, meta.Created, "peer event")
return nil
}
changed, err := applyBucketConfig(&meta, configFile, configData, updatedAt)
if err != nil {
return err
}
if !changed {
return nil
}
} else {
switch configFile {
case bucketNotificationConfig:
meta.NotificationConfigXML = configData
meta.NotificationConfigUpdatedAt = updatedAt
case bucketLifecycleConfig:
meta.LifecycleConfigXML = configData
meta.LifecycleConfigUpdatedAt = updatedAt
case bucketReplicationConfig:
meta.ReplicationConfigXML = configData
meta.ReplicationConfigUpdatedAt = updatedAt
case bucketTargetsFile:
meta.BucketTargetsConfigJSON, meta.BucketTargetsConfigMetaJSON, err = encryptBucketMetadata(ctx, meta.Name, configData, kms.Context{
bucket: meta.Name,
bucketTargetsFile: bucketTargetsFile,
})
if err != nil {
return fmt.Errorf("Error encrypting bucket target metadata %w", err)
}
meta.BucketTargetsConfigUpdatedAt = updatedAt
meta.BucketTargetsConfigMetaUpdatedAt = updatedAt
default:
return fmt.Errorf("Unknown bucket %s metadata update requested %s", bucket, configFile)
}
meta.BucketTargetsConfigUpdatedAt = updatedAt
meta.BucketTargetsConfigMetaUpdatedAt = updatedAt
default:
return fmt.Errorf("Unknown bucket %s metadata update requested %s", bucket, configFile)
}
return sys.saveMetadata(ctx, objAPI, meta)
if err := sys.saveMetadata(ctx, objAPI, &meta); err != nil {
return err
}
result = bucketMetadataUpdate{meta: meta, updatedAt: updatedAt, changed: true}
return nil
}()
if err != nil {
return updatedAt, err
return result, err
}
globalNotificationSys.LoadBucketMetadata(bgContext(notifyCtx), bucket) // Do not use caller context here
return updatedAt, nil
if result.changed {
globalNotificationSys.LoadBucketMetadata(bgContext(notifyCtx), bucket)
}
return result, nil
}
func (sys *BucketMetadataSys) save(ctx context.Context, meta BucketMetadata) error {
@@ -218,7 +260,7 @@ func (sys *BucketMetadataSys) save(ctx context.Context, meta BucketMetadata) err
return errInvalidArgument
}
if err := sys.saveMetadata(ctx, objAPI, meta); err != nil {
if err := sys.saveMetadata(ctx, objAPI, &meta); err != nil {
return err
}
@@ -228,11 +270,16 @@ func (sys *BucketMetadataSys) save(ctx context.Context, meta BucketMetadata) err
// saveMetadata persists and publishes metadata locally. Callers performing a
// read-modify-write must hold metadata.lock and release it before peer fan-out.
func (sys *BucketMetadataSys) saveMetadata(ctx context.Context, objAPI ObjectLayer, meta BucketMetadata) error {
func (sys *BucketMetadataSys) saveMetadata(ctx context.Context, objAPI ObjectLayer, meta *BucketMetadata) error {
// A writer may have queued for metadata.lock before DeleteBucket completed.
// Recheck the physical bucket under that lock, before recreating metadata.
if _, err := objAPI.GetBucketInfo(ctx, meta.Name, BucketOptions{NoMetadata: true}); err != nil {
return err
}
if err := meta.Save(ctx, objAPI); err != nil {
return err
}
sys.Set(meta.Name, meta)
sys.Set(meta.Name, *meta)
return nil
}
@@ -240,8 +287,19 @@ func lockBucketMetadata(ctx context.Context, objectAPI ObjectLayer, bucket strin
return lockBucketMetadataWithTimeout(ctx, objectAPI, bucket, globalOperationTimeout)
}
// lockBucketMetadataAcquireHook, when set, is invoked at the start of every
// metadata.lock acquisition, immediately before the blocking Lock() call. It is
// nil in production (a single atomic load, no behavior change) and exists only
// so tests can deterministically observe a caller reaching the metadata lock —
// notably DeleteBucket, whose lock is taken through its erasureServerPools
// receiver and is therefore invisible to an injected object layer.
var lockBucketMetadataAcquireHook atomic.Pointer[func(bucket string)]
func lockBucketMetadataWithTimeout(ctx context.Context, objectAPI ObjectLayer, bucket string, timeout *dynamicTimeout) (context.Context, func(), error) {
lock := objectAPI.NewNSLock(minioMetaBucket, pathJoin(bucketMetaPrefix, bucket, "metadata.lock"))
if hook := lockBucketMetadataAcquireHook.Load(); hook != nil {
(*hook)(bucket)
}
lkctx, err := lock.GetLock(ctx, timeout)
if err != nil {
return nil, nil, err
@@ -292,6 +350,57 @@ func (sys *BucketMetadataSys) Update(ctx context.Context, bucket string, configF
return sys.updateAndParse(ctx, bucket, configFile, configData, true, false)
}
// UpdateExpiryLCConfig merges a replicated ILM expiry configuration with the
// bucket's current lifecycle document and persists the merged result while
// holding metadata.lock across the read, merge, and save. The site-replication
// expiry heal and peer-apply paths must use this instead of computing the merge
// from an unlocked GetConfigFromDisk read and then writing it with Update: that
// two-step sequence drops any lifecycle transition change committed in between
// (issue #105). Lock order stays <bucket>.lck -> metadata.lock -> .metadata.bin;
// the merge and save run under metadata.lock and the peer fan-out runs after it
// is released.
func (sys *BucketMetadataSys) UpdateExpiryLCConfig(ctx context.Context, bucket string, expLCConfig *string, updatedAt time.Time) error {
objAPI := newObjectLayerFn()
if objAPI == nil {
return errServerNotInitialized
}
if isMinioMetaBucketName(bucket) {
return errInvalidArgument
}
notifyCtx := ctx
ctx, unlock, err := lockBucketMetadata(ctx, objAPI, bucket)
if err != nil {
return err
}
err = func() error {
defer unlock()
meta, err := loadBucketMetadataParse(ctx, objAPI, bucket, true)
if err != nil {
if !globalIsErasure && !globalIsDistErasure && errors.Is(err, errVolumeNotFound) {
// Only single drive mode needs this fallback.
meta = newBucketMetadata(bucket)
} else {
return err
}
}
configData, err := mergeExpiryWithLCConfig(bucket, meta, expLCConfig, updatedAt)
if err != nil {
return err
}
meta.LifecycleConfigXML = configData
meta.LifecycleConfigUpdatedAt = UTCNow()
return sys.saveMetadata(ctx, objAPI, &meta)
}()
if err != nil {
return err
}
globalNotificationSys.LoadBucketMetadata(bgContext(notifyCtx), bucket) // Do not use caller context here
return nil
}
// Get metadata for a bucket.
// If no metadata exists errConfigNotFound is returned and a new metadata is returned.
// Only a shallow copy is returned, so referenced data should not be modified,
+4 -9
View File
@@ -303,6 +303,9 @@ func loadBucketMetadataParseUnderLock(ctx context.Context, objectAPI ObjectLayer
return newBucketMetadata(bucket), fmt.Errorf("%w: %v", errBucketMetadataMigrationLockUnavailable, err)
}
defer unlock()
if _, err := objectAPI.GetBucketInfo(ctx, bucket, BucketOptions{NoMetadata: true}); err != nil {
return newBucketMetadata(bucket), err
}
return loadBucketMetadataParse(ctx, objectAPI, bucket, parse)
}
@@ -385,15 +388,7 @@ func (b *BucketMetadata) parseAllConfigs(ctx context.Context, objectAPI ObjectLa
} else {
b.objectLockConfig = nil
}
if b.objectLockConfig != nil {
// Object Lock requires every object to be versioned. Whatever the lock
// document contains, a suspended or prefix-excluded versioning document
// is replaced by plain Enabled versioning; Save persists the result.
config, versioningErr := versioning.ParseConfig(bytes.NewReader(b.VersioningConfigXML))
if versioningErr != nil || !config.Enabled() || config.PrefixesExcluded() {
b.VersioningConfigXML = enabledBucketVersioningConfig
}
}
b.VersioningConfigXML = effectiveBucketVersioning(b.VersioningConfigXML, b.objectLockConfig != nil)
if len(b.VersioningConfigXML) != 0 {
b.versioningConfig, err = versioning.ParseConfig(bytes.NewReader(b.VersioningConfigXML))
+25
View File
@@ -39,3 +39,28 @@ func TestBucketMetadataCorsRoundTrip(t *testing.T) {
t.Fatalf("CorsConfigUpdatedAt not preserved")
}
}
// A persisted retired extension must not prevent the whole bucket's metadata
// from loading, including unrelated versioning and ordinary lifecycle rules.
func TestBucketMetadataRetiredAccessTiering(t *testing.T) {
meta := newBucketMetadata("retired-access")
meta.LifecycleConfigXML = []byte(`<LifecycleConfiguration><AccessTierQuota>500GiB</AccessTierQuota><Rule><ID>access</ID><Status>Enabled</Status><Filter><Prefix>logs/</Prefix></Filter><AccessTransition><Window>10m</Window><PromoteAfterAccesses>10</PromoteAfterAccesses></AccessTransition></Rule><Rule><ID>ordinary</ID><Status>Enabled</Status><Filter><Prefix>expired/</Prefix></Filter><Expiration><Days>30</Days></Expiration></Rule></LifecycleConfiguration>`)
meta.VersioningConfigXML = []byte(`<VersioningConfiguration xmlns="http://s3.amazonaws.com/doc/2006-03-01/"><Status>Enabled</Status></VersioningConfiguration>`)
data, err := meta.MarshalMsg(nil)
if err != nil {
t.Fatal(err)
}
got := newBucketMetadata(meta.Name)
if _, err := got.UnmarshalMsg(data); err != nil {
t.Fatal(err)
}
if err := got.parseAllConfigs(t.Context(), nil); err != nil {
t.Fatalf("bucket metadata failed to load: %v", err)
}
if got.lifecycleConfig == nil || got.lifecycleConfig.HasActiveRules("logs/") || !got.lifecycleConfig.HasActiveRules("expired/") {
t.Fatal("unexpected lifecycle behavior")
}
if got.versioningConfig == nil || !got.versioningConfig.Enabled() {
t.Fatal("unrelated versioning lost")
}
}
+125 -10
View File
@@ -25,6 +25,7 @@ import (
"strings"
"time"
"github.com/minio/minio/internal/amztime"
"github.com/minio/minio/internal/auth"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
xhttp "github.com/minio/minio/internal/http"
@@ -333,11 +334,10 @@ func checkPutObjectLockAllowed(ctx context.Context, rq *http.Request, bucket, ob
return mode, retainDate, legalHold, ErrObjectLocked
}
if !legalHoldRequested && retentionCfg.LockEnabled {
// inherit retention from bucket configuration
return retentionCfg.Mode, objectlock.RetentionDate{Time: t.Add(retentionCfg.Validity)}, legalHold, ErrNone
}
return "", objectlock.RetentionDate{}, legalHold, ErrNone
// Inherit retention from the bucket configuration. A legal-hold header
// on the same request, ON or OFF, is independent of retention and must
// not suppress the default (#165).
return retentionCfg.Mode, objectlock.RetentionDate{Time: t.Add(retentionCfg.Validity)}, legalHold, ErrNone
}
return mode, retainDate, legalHold, ErrNone
}
@@ -386,22 +386,137 @@ func (s objectLockState) legalHoldIsOlderThan(src time.Time) bool {
// restoreRetention and restoreLegalHold put the stored state back into
// metadata that was rebuilt from a request whose update was not applied.
func (s objectLockState) restoreRetention(metadata map[string]string) {
// The stored timestamp orders the next update and must survive even when
// the stored value is empty, which is how a removal is recorded.
if s.retentionTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = s.retentionTimestamp
}
if s.mode == "" {
return
}
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = s.mode
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = s.retainUntil
if s.retentionTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = s.retentionTimestamp
}
}
func (s objectLockState) restoreLegalHold(metadata map[string]string) {
if s.legalHoldTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = s.legalHoldTimestamp
}
if s.legalHold == "" {
return
}
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = s.legalHold
if s.legalHoldTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = s.legalHoldTimestamp
}
// replicaStoredLock reads the Object Lock state stored on the addressed version
// so a trusted replica write can order its update against it. A missing object
// or version yields an empty state, which is correct for the first write of a
// version; any other read error is returned so the caller fails the write rather
// than ordering an incoming update against lock state it merely failed to read
// (an older incoming value must not win over a newer stored one just because the
// read timed out).
func replicaStoredLock(ctx context.Context, getObjectInfo GetObjectInfoFn, bucket, object, versionID string) (objectLockState, error) {
oi, err := getObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: versionID})
switch {
case err == nil:
return storedObjectLockState(oi.UserDefined), nil
case isErrObjectNotFound(err) || isErrVersionNotFound(err):
return objectLockState{}, nil
default:
return objectLockState{}, err
}
}
// applyReplicatedObjectLock writes the retention and legal-hold decision into
// metadata for a PUT, CopyObject, or multipart-initiation request. A request
// that is not an actual trusted replica -- a normal user write, or a trusted
// peer that carried the replication marker without REPLICA status -- takes
// ordinary write semantics: a validated value is applied and stamped now, and a
// missing value is left as is. Only an actual replica update is ordered against
// the state already stored on the addressed version, so a stale value cannot
// overwrite a newer one and a full retransmit cannot roll a destination back.
// The stored argument is meaningful only for a replica; callers pass an empty
// state otherwise. Only the two Object Lock keys and their reserved ordering
// timestamps are touched; any encryption-metadata reconciliation stays with the
// caller.
func applyReplicatedObjectLock(metadata map[string]string, stored objectLockState,
replicaTrusted bool,
retentionMode objectlock.RetMode, retentionDate objectlock.RetentionDate,
legalHold objectlock.ObjectLegalHold, srcRetentionTimestamp, srcLegalholdTimestamp time.Time,
) {
switch {
case !replicaTrusted:
// Ordinary write semantics: apply a validated retention and stamp it now;
// a missing value carries no instruction, so leave the metadata as it is.
if retentionMode.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
case !stored.retentionIsOlderThan(srcRetentionTimestamp):
// The stored update is at least as new as this replica's, or the replica
// carries no ordering timestamp: keep what is stored. This is also how a
// stale retransmit is rejected.
stored.restoreRetention(metadata)
default:
// The replica update wins. A removal carries no value but still records
// the source timestamp that orders it.
if retentionMode.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
}
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = srcRetentionTimestamp.UTC().Format(time.RFC3339Nano)
}
// Legal hold has no removal in S3: an explicitly empty header is already
// rejected as an invalid status, so the only value-less shape that gets here
// is an absent one, which conveys no legal-hold change. Only a valid status
// can win.
switch {
case !replicaTrusted:
if legalHold.Status.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
case legalHold.Status.Valid() && stored.legalHoldIsOlderThan(srcLegalholdTimestamp):
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = srcLegalholdTimestamp.UTC().Format(time.RFC3339Nano)
default:
stored.restoreLegalHold(metadata)
}
}
// reconcileStoredObjectLock re-orders the Object Lock already written into
// metadata against the state currently stored on the destination version, both
// compared by their reserved ordering timestamps. It runs inside the object
// layer under the namespace write lock that guards the version replacement,
// after the destination version is read and before the new one is committed, so
// a replica update whose ordering was decided at handler time (or, for multipart,
// at initiation) cannot overwrite a newer lock update that reached the version in
// between. metadata already carries the incoming update with its source
// timestamps; a stored value that is not older than the incoming one is put back,
// which for a stored removal means clearing the incoming value and keeping only
// the removal's timestamp. Only the two lock keys and their reserved timestamps
// move; a non-replica write never sets the flag that invokes this.
func reconcileStoredObjectLock(metadata map[string]string, stored objectLockState) {
incoming := storedObjectLockState(metadata)
incomingRetentionTS, _ := time.Parse(time.RFC3339Nano, incoming.retentionTimestamp)
if !stored.retentionIsOlderThan(incomingRetentionTS) {
// The stored retention is at least as new as the incoming one (or the
// incoming update is unordered): drop the incoming value and put the stored
// state back, which may itself be a removal (value keys absent, timestamp
// present).
delete(metadata, strings.ToLower(xhttp.AmzObjectLockMode))
delete(metadata, strings.ToLower(xhttp.AmzObjectLockRetainUntilDate))
delete(metadata, ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp)
stored.restoreRetention(metadata)
}
incomingLegalHoldTS, _ := time.Parse(time.RFC3339Nano, incoming.legalHoldTimestamp)
if incoming.legalHold == "" || !stored.legalHoldIsOlderThan(incomingLegalHoldTS) {
delete(metadata, strings.ToLower(xhttp.AmzObjectLockLegalHold))
delete(metadata, ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp)
stored.restoreLegalHold(metadata)
}
}
+5 -6
View File
@@ -19,7 +19,6 @@ package cmd
import (
"bytes"
"encoding/json"
"io"
"net/http"
@@ -100,13 +99,13 @@ func (api objectAPIHandlers) PutBucketPolicyHandler(w http.ResponseWriter, r *ht
return
}
configData, err := json.Marshal(bucketPolicy)
configData, err := canonicalBucketPolicy(bucketPolicy)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
updatedAt, err := globalBucketMetadataSys.Update(ctx, bucket, bucketPolicyConfig, configData)
result, err := globalBucketMetadataSys.updateAndParseMetadata(ctx, bucket, bucketPolicyConfig, configData, false, false, nil)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -116,8 +115,8 @@ func (api objectAPIHandlers) PutBucketPolicyHandler(w http.ResponseWriter, r *ht
replLogIf(ctx, globalSiteReplicationSys.BucketMetaHook(ctx, madmin.SRBucketMeta{
Type: madmin.SRBucketMetaTypePolicy,
Bucket: bucket,
Policy: bucketPolicyBytes,
UpdatedAt: updatedAt,
Policy: result.meta.PolicyConfigJSON,
UpdatedAt: result.updatedAt,
}))
// Success.
@@ -200,7 +199,7 @@ func (api objectAPIHandlers) GetBucketPolicyHandler(w http.ResponseWriter, r *ht
return
}
configData, err := json.Marshal(config)
configData, err := canonicalBucketPolicy(config)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
+26 -7
View File
@@ -255,13 +255,32 @@ func getConditionValuesWithTags(r *http.Request, lc string, cred auth.Credential
}
cloneHeader := r.Header.Clone()
signatureAge := cloneHeader.Get("x-amz-signature-age")
cloneHeader.Del("x-amz-signature-age")
// The presigned V4 verifier overwrites this internal scratch header after
// validating the signature. Ignore a value supplied on every other request
// type, where it would otherwise synthesize s3:signatureAge.
if authType == authTypePresigned && signatureAge != "" {
args["signatureAge"] = []string{signatureAge}
// s3:signatureAge is derived from the presigned X-Amz-Date rather than from
// anything the verifier writes back: PutObject and UploadPart authorize
// before they verify the signature, so a post-verification value is not yet
// available on the first evaluation. The date is bound by the signature
// (doesPresignedSignatureMatch rebuilds and compares it), so a forged date
// only changes the authorization outcome of a request that then fails
// verification. A date that does not parse leaves the key absent; the
// verifier rejects the request as ErrMalformedPresignedDate.
if authType == authTypePresigned {
if signedDate, err := time.Parse(iso8601Format, r.Form.Get(xhttp.AmzDate)); err == nil {
args["signatureAge"] = []string{strconv.FormatInt(currTime.Sub(signedDate).Milliseconds(), 10)}
}
}
// s3:x-amz-content-sha256 must name the payload hash the request is actually
// verified and enforced against, and only one such value. Presence of the
// header controls whether the key exists at all (AWS documents that the
// query-string form does not populate it), but the value comes from the same
// selection getContentSha256Cksum makes for verification: the presigned query
// value takes precedence over the header, and a repeated header contributes
// only its first value. Exposing every raw header value instead let a
// request satisfy a policy with a value the verifier never checked.
if _, ok := cloneHeader[xhttp.AmzContentSha256]; ok {
args[xhttp.AmzContentSha256] = []string{getContentSha256Cksum(r, serviceS3)}
cloneHeader.Del(xhttp.AmzContentSha256)
}
userTags := cloneHeader.Get(xhttp.AmzObjectTagging)
+34 -6
View File
@@ -23,8 +23,10 @@ import (
"net/url"
"os"
"slices"
"strconv"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/handlers"
@@ -579,8 +581,20 @@ func TestGetConditionValuesRejectsAbsentInternalKeys(t *testing.T) {
}
}
// s3:signatureAge is derived from the presigned X-Amz-Date, which the signature
// binds. A client header under the former scratch name must never supply it on
// any auth type, and a presign whose date is missing or malformed leaves the
// key absent (the verifier then rejects the request).
func TestGetConditionValuesOnlyAcceptsPresignedSignatureAge(t *testing.T) {
const signatureAgeHeader = "x-amz-signature-age"
signedDate := UTCNow().Add(-90 * time.Second)
presignQuery := func(date string) string {
q := url.Values{xhttp.AmzCredential: {"access/20260803/us-east-1/s3/aws4_request"}}
if date != "" {
q.Set(xhttp.AmzDate, date)
}
return "http://minio.local/bkt/obj?" + q.Encode()
}
for _, tc := range []struct {
name string
@@ -602,19 +616,33 @@ func TestGetConditionValuesOnlyAcceptsPresignedSignatureAge(t *testing.T) {
},
},
{
name: "presigned verifier value",
target: "http://minio.local/bkt/obj?" + url.Values{
xhttp.AmzCredential: {"access/20260803/us-east-1/s3/aws4_request"},
}.Encode(),
name: "presigned client header without date",
target: presignQuery(""),
headers: map[string]string{signatureAgeHeader: "250"},
},
{
name: "presigned malformed date",
target: presignQuery("yesterday"),
},
{
name: "presigned signed date",
target: presignQuery(signedDate.Format(iso8601Format)),
headers: map[string]string{signatureAgeHeader: "250"},
want: true,
},
} {
t.Run(tc.name, func(t *testing.T) {
got := condValuesForRequest(t, tc.target, tc.headers)
_, ok := got["signatureAge"]
v, ok := got["signatureAge"]
if ok != tc.want {
t.Fatalf("signatureAge presence: expected %v, got %v", tc.want, got["signatureAge"])
t.Fatalf("signatureAge presence: expected %v, got %v", tc.want, v)
}
if !tc.want {
return
}
age, err := strconv.ParseInt(strings.Join(v, ""), 10, 64)
if err != nil || age < (90*time.Second).Milliseconds() || age > (2*time.Minute).Milliseconds() {
t.Fatalf("signatureAge = %v, want about 90s derived from X-Amz-Date rather than the client header", v)
}
})
}
+12 -8
View File
@@ -43,6 +43,17 @@ func NewBucketQuotaSys() *BucketQuotaSys {
return &BucketQuotaSys{}
}
// getBucketQuotaSize returns the effective enforced hard-quota size.
func getBucketQuotaSize(quota *madmin.BucketQuota) uint64 {
if quota == nil || quota.Type != madmin.HardQuota {
return 0
}
if quota.Size > 0 {
return quota.Size
}
return quota.Quota
}
var bucketStorageCache = cachevalue.New[DataUsageInfo]()
// Init initialize bucket quota.
@@ -110,14 +121,7 @@ func (sys *BucketQuotaSys) enforceQuotaHard(ctx context.Context, bucket string,
return err
}
var quotaSize uint64
if q != nil && q.Type == madmin.HardQuota {
if q.Size > 0 {
quotaSize = q.Size
} else if q.Quota > 0 {
quotaSize = q.Quota
}
}
quotaSize := getBucketQuotaSize(q)
if quotaSize > 0 {
if uint64(size) >= quotaSize { // check if file size already exceeds the quota
return BucketQuotaExceeded{Bucket: bucket}
+77
View File
@@ -0,0 +1,77 @@
// Copyright (c) 2015-2025 MinIO, Inc.
// Copyright (c) 2025-2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful,
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"testing"
"github.com/minio/madmin-go/v3"
)
func TestGetBucketQuotaSize(t *testing.T) {
tests := []struct {
name string
quota *madmin.BucketQuota
want uint64
}{
{name: "nil"},
{name: "empty", quota: &madmin.BucketQuota{}},
{name: "current size", quota: &madmin.BucketQuota{Type: madmin.HardQuota, Size: 1024}, want: 1024},
{name: "legacy quota", quota: &madmin.BucketQuota{Type: madmin.HardQuota, Quota: 2048}, want: 2048},
{name: "size takes precedence", quota: &madmin.BucketQuota{Type: madmin.HardQuota, Size: 1024, Quota: 2048}, want: 1024},
{name: "missing type", quota: &madmin.BucketQuota{Size: 1024}},
{name: "unsupported type", quota: &madmin.BucketQuota{Type: "fifo", Size: 1024}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := getBucketQuotaSize(tt.quota); got != tt.want {
t.Fatalf("getBucketQuotaSize() = %d, want %d", got, tt.want)
}
})
}
}
func TestIsBktQuotaCfgReplicated(t *testing.T) {
hardQuota := func(size, legacy uint64) *madmin.BucketQuota {
return &madmin.BucketQuota{Type: madmin.HardQuota, Size: size, Quota: legacy}
}
tests := []struct {
name string
quotas []*madmin.BucketQuota
want bool
}{
{name: "none configured", quotas: []*madmin.BucketQuota{nil, nil}, want: true},
{name: "missing from one site", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), nil}},
{name: "matching size", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), hardQuota(1024, 0)}, want: true},
{name: "different size", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), hardQuota(2048, 0)}},
{name: "equivalent representations", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), hardQuota(0, 1024)}, want: true},
{name: "different typeless size", quotas: []*madmin.BucketQuota{{Size: 1024}, {Size: 2048}}},
{name: "different type", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), {Type: "fifo", Size: 1024}}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := isBktQuotaCfgReplicated(len(tt.quotas), tt.quotas); got != tt.want {
t.Fatalf("isBktQuotaCfgReplicated() = %v, want %v", got, tt.want)
}
})
}
}
+7 -4
View File
@@ -418,6 +418,9 @@ func getReplicationState(rinfos replicatedInfos, prevState ReplicationState, vID
for _, rinfo := range rinfos.Targets {
if rinfo.ResyncTimestamp != "" {
if rs.ResetStatusesMap == nil {
rs.ResetStatusesMap = make(map[string]string)
}
rs.ResetStatusesMap[targetResetHeader(rinfo.Arn)] = rinfo.ResyncTimestamp
}
}
@@ -640,10 +643,10 @@ type VersionPurgeStatusType = replication.VersionPurgeStatusType
type replicationResyncer struct {
// map of bucket to their resync status
statusMap map[string]BucketReplicationResyncStatus
workerSize int
resyncCancelCh chan struct{}
workerCh chan struct{}
statusMap map[string]BucketReplicationResyncStatus
workerSize int
cancelResyncs map[resyncOpts]context.CancelCauseFunc
workerCh chan struct{}
sync.RWMutex
}
+565 -176
View File
File diff suppressed because it is too large Load Diff
+788
View File
@@ -18,13 +18,21 @@
package cmd
import (
"context"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"path"
"strings"
"sync"
"testing"
"testing/synctest"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio-go/v7"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
)
@@ -307,3 +315,783 @@ func TestReplicationValidationObjectUsesRulePrefix(t *testing.T) {
})
}
}
// The resync-finalization tests below exercise the real result sink, the
// finish() shutdown ordering, and the sendResyncResult / finalResyncStatus
// helpers, plus (for the persistence cases) markStatus with on-disk
// round-tripping. resyncBucket cannot be driven end to end in a unit test
// because its workers call a live remote target (StatObject), so the helpers it
// uses are exercised directly. The blocking-order assertions run under
// testing/synctest so a removed wait fails deterministically, with no timing
// windows.
func newTestResyncer(bucket, arn string) (*replicationResyncer, resyncOpts) {
s := &replicationResyncer{
statusMap: map[string]BucketReplicationResyncStatus{},
}
brs := newBucketResyncStatus(bucket)
brs.TargetsMap[arn] = TargetReplicationResyncStatus{ResyncStatus: ResyncStarted, ResyncID: "reset-" + bucket}
s.statusMap[bucket] = brs
return s, resyncOpts{bucket: bucket, arn: arn, resyncID: "reset-" + bucket}
}
// TestResyncBucketFinalize round-trips the terminal status through a real
// ObjectLayer: a clean run persists Completed with every result, while a run
// whose parent context was canceled during the drain is downgraded to Failed;
// a user-canceled run persists Canceled. Completed never misrepresents an
// incomplete resync.
func TestResyncBucketFinalize(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
objAPI, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatalf("prepare erasure backend: %v", err)
}
defer removeRoots(fsDirs)
// persistTerminal applies resyncBucket's finalizer logic (finalResyncStatus
// then markStatus, which persists) and reads the status back the way the
// resync status API does.
persistTerminal := func(t *testing.T, s *replicationResyncer, opts resyncOpts, status ResyncStatusType, cause error) TargetReplicationResyncStatus {
t.Helper()
s.markStatus(finalResyncStatus(status, cause), opts, objAPI)
brs, err := loadBucketResyncMetadata(ctx, opts.bucket, objAPI)
if err != nil {
t.Fatalf("load persisted resync metadata: %v", err)
}
return brs.TargetsMap[opts.arn]
}
// 1. Clean completion: every result - including the failed object - is folded
// into the persisted status, which stays Completed.
t.Run("persists complete counts", func(t *testing.T) {
s, opts := newTestResyncer("finalize-counts", "arn1")
results := s.newResyncResults(opts)
results.ch <- TargetReplicationResyncStatus{Object: "ok-1", ReplicatedCount: 1, ReplicatedSize: 100}
results.ch <- TargetReplicationResyncStatus{Object: "ok-2", ReplicatedCount: 1, ReplicatedSize: 200}
results.ch <- TargetReplicationResyncStatus{Object: "bad", FailedCount: 1, FailedSize: 300}
var wg sync.WaitGroup // no producer workers for this case
results.finish(nil, &wg)
st := persistTerminal(t, s, opts, ResyncCompleted, nil)
if st.ResyncStatus != ResyncCompleted {
t.Fatalf("persisted status = %s, want Completed", st.ResyncStatus)
}
if st.ReplicatedCount != 2 || st.ReplicatedSize != 300 || st.FailedCount != 1 || st.FailedSize != 300 {
t.Fatalf("persisted counts = {replicated:%d/%d failed:%d/%d}, want {2/300 1/300}",
st.ReplicatedCount, st.ReplicatedSize, st.FailedCount, st.FailedSize)
}
})
// 2. Parent context canceled during the drain -> Completed downgraded to
// Failed (markStatus persists under its own context, so nothing else stops
// a bare Completed from being recorded).
t.Run("parent cancel during drain downgrades to failed", func(t *testing.T) {
s, opts := newTestResyncer("finalize-parent-cancel", "arn1")
results := s.newResyncResults(opts)
results.ch <- TargetReplicationResyncStatus{Object: "ok-1", ReplicatedCount: 1, ReplicatedSize: 100}
var wg sync.WaitGroup
results.finish(nil, &wg)
cctx, ccancel := context.WithCancel(context.Background())
ccancel()
st := persistTerminal(t, s, opts, ResyncCompleted, context.Cause(cctx))
if st.ResyncStatus != ResyncFailed {
t.Fatalf("persisted status = %s, want Failed (parent canceled during drain)", st.ResyncStatus)
}
})
// 3. A user-canceled worker cannot report a completed resync.
t.Run("user cancel persists canceled", func(t *testing.T) {
s, opts := newTestResyncer("finalize-worker-abort", "arn1")
ctx, cancel := context.WithCancelCause(context.Background())
cancel(errResyncCanceled)
ch := make(chan TargetReplicationResyncStatus)
if s.sendResyncResult(ctx, ch, TargetReplicationResyncStatus{Object: "dropped", ReplicatedCount: 1}) {
t.Fatal("sendResyncResult reported success after cancellation")
}
st := persistTerminal(t, s, opts, ResyncCompleted, context.Cause(ctx))
if st.ResyncStatus != ResyncCanceled {
t.Fatalf("persisted status = %s, want Canceled", st.ResyncStatus)
}
})
}
// TestResyncFinishDrainsResults asserts finish() does not return until the
// consumer has applied the final result (the #136 defect). A gated apply holds
// the last result unapplied; under synctest finish() must stay durably blocked
// until it is released - if rr.wg.Wait() is removed, finish() returns early and
// the test fails deterministically.
func TestResyncFinishDrainsResults(t *testing.T) {
synctest.Test(t, func(t *testing.T) {
s, opts := newTestResyncer("drain", "arn1")
reachedFinal := make(chan struct{})
release := make(chan struct{})
results := startResyncResults(func(r TargetReplicationResyncStatus) {
if r.Object == "final" {
close(reachedFinal)
<-release
}
s.incStats(r, opts)
})
results.ch <- TargetReplicationResyncStatus{Object: "ok-1", ReplicatedCount: 1, ReplicatedSize: 100}
results.ch <- TargetReplicationResyncStatus{Object: "final", FailedCount: 1, FailedSize: 200}
<-reachedFinal // consumer received "final" but is gated before incStats(final)
var wg sync.WaitGroup
finishDone := make(chan struct{})
go func() {
results.finish(nil, &wg)
close(finishDone)
}()
synctest.Wait()
select {
case <-finishDone:
close(release)
synctest.Wait()
t.Fatal("finish() returned before the final result was drained (drain wait missing)")
default:
// finish() is durably blocked in rr.wg.Wait() - correct.
}
close(release)
synctest.Wait()
<-finishDone
st := s.statusMap[opts.bucket].TargetsMap[opts.arn]
if st.ReplicatedCount != 1 || st.FailedCount != 1 || st.FailedSize != 200 {
t.Fatalf("status after finish = {replicated:%d failed:%d/%d}, want {1 1/200}",
st.ReplicatedCount, st.FailedCount, st.FailedSize)
}
})
}
// TestResyncFinishWaitsForInflightWorker asserts finish() stops the producer
// workers before it closes the result channel, so an in-flight worker (as on an
// early-return path) never sends on a closed channel and its result is not lost.
// A gated worker stays in flight past the shutdown request; under synctest
// finish() must stay durably blocked until the worker is released - if
// workerWg.Wait() is removed, finish() returns early and the test fails
// deterministically.
func TestResyncFinishWaitsForInflightWorker(t *testing.T) {
synctest.Test(t, func(t *testing.T) {
s, opts := newTestResyncer("workers", "arn1")
results := startResyncResults(func(r TargetReplicationResyncStatus) { s.incStats(r, opts) })
workers := []chan ReplicateObjectInfo{make(chan ReplicateObjectInfo, 1)}
var wg sync.WaitGroup
gotRoi := make(chan struct{})
release := make(chan struct{})
wg.Add(1)
go func() {
defer wg.Done()
for roi := range workers[0] {
close(gotRoi)
<-release
// Mirror the real worker's send; recover so that if finish()
// wrongly closed the result channel first, the test fails via the
// assertion below instead of crashing on send-on-closed.
func() {
defer func() { _ = recover() }()
results.ch <- TargetReplicationResyncStatus{Object: roi.Name, ReplicatedCount: 1, ReplicatedSize: 500}
}()
}
}()
workers[0] <- ReplicateObjectInfo{Name: "inflight"}
<-gotRoi // worker holds a result in flight, not yet delivered
finishDone := make(chan struct{})
go func() {
results.finish(workers, &wg)
close(finishDone)
}()
synctest.Wait()
select {
case <-finishDone:
close(release)
synctest.Wait()
t.Fatal("finish() closed the result channel before the in-flight worker finished (worker wait missing)")
default:
// finish() is durably blocked in workerWg.Wait() - correct.
}
close(release)
synctest.Wait()
<-finishDone
st := s.statusMap[opts.bucket].TargetsMap[opts.arn]
if st.ReplicatedCount != 1 || st.ReplicatedSize != 500 {
t.Fatalf("status after finish = {replicated:%d/%d}, want {1/500}", st.ReplicatedCount, st.ReplicatedSize)
}
})
}
// TestResyncResultFor asserts the resync worker classifies a target from the
// actual replication outcome, not from whether the target version merely exists.
// The key regression is the "failed update over an existing version" case: a
// quota-rejected update leaves the old version in place, and counting existence
// (the previous behavior) would score it a success. It also checks a genuine
// success, an errored-but-Completed result, a delete failure, a delete-marker
// success (zero bytes), and an ARN that was never attempted.
func TestResyncResultFor(t *testing.T) {
const arn = "arn:minio:replication::id:bucket"
obj := ReplicateObjectInfo{Name: "obj", Bucket: "bucket", Size: 196608}
deleteMarker := ReplicateObjectInfo{Name: "dm", Bucket: "bucket", Size: 0, DeleteMarker: true}
tests := []struct {
name string
roi ReplicateObjectInfo
rinfos replicatedInfos
wantRepl, wantReplSize, wantFail, wantFailSize int64
}{
{
name: "completed update",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed, Size: 196608},
}},
wantRepl: 1, wantReplSize: 196608,
},
{
name: "failed update over existing version",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Failed, Err: fmt.Errorf("quota exceeded"), Size: 196608},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "completed but errored is a failure",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed, Err: fmt.Errorf("boom"), Size: 196608},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "delete failed",
roi: deleteMarker,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Failed},
}},
wantFail: 1, wantFailSize: 0,
},
{
name: "delete marker replicated counts zero bytes",
roi: deleteMarker,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed},
}},
wantRepl: 1, wantReplSize: 0,
},
{
name: "arn not attempted is a failure",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: "arn:minio:replication::id2:bucket", ReplicationStatus: replication.Completed, Size: 196608},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "completed with zero size falls back to object size",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed, Size: 0},
}},
wantRepl: 1, wantReplSize: 196608,
},
{
name: "version purge complete is a success",
roi: ReplicateObjectInfo{Name: "purge", Bucket: "bucket", Size: 196608, VersionPurgeStatus: replication.VersionPurgePending},
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
// a successful purge sets only VersionPurgeStatus; ReplicationStatus stays empty.
{Arn: arn, VersionPurgeStatus: replication.VersionPurgeComplete},
}},
wantRepl: 1, wantReplSize: 196608,
},
{
name: "version purge failed is a failure",
roi: ReplicateObjectInfo{Name: "purge", Bucket: "bucket", Size: 196608, VersionPurgeStatus: replication.VersionPurgePending},
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, VersionPurgeStatus: replication.VersionPurgeFailed, Err: fmt.Errorf("quota exceeded")},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "benign duplicate 412 is a success",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
// the destination answers PreconditionFailed for an exact duplicate;
// replicateAll keeps Completed but retains the error.
{Arn: arn, ReplicationStatus: replication.Completed, Err: minio.ErrorResponse{Code: "PreconditionFailed"}, Size: 196608},
}},
wantRepl: 1, wantReplSize: 196608,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
st := resyncResultFor(tc.rinfos, arn, tc.roi)
if st.Object != tc.roi.Name || st.Bucket != tc.roi.Bucket {
t.Fatalf("object/bucket = %s/%s, want %s/%s", st.Object, st.Bucket, tc.roi.Name, tc.roi.Bucket)
}
if st.ReplicatedCount != tc.wantRepl || st.ReplicatedSize != tc.wantReplSize ||
st.FailedCount != tc.wantFail || st.FailedSize != tc.wantFailSize {
t.Fatalf("resyncResultFor = {replicated:%d/%d failed:%d/%d}, want {%d/%d %d/%d}",
st.ReplicatedCount, st.ReplicatedSize, st.FailedCount, st.FailedSize,
tc.wantRepl, tc.wantReplSize, tc.wantFail, tc.wantFailSize)
}
})
}
}
// TestObjectNeedsResyncForARN asserts the resync dispatch is scoped to the
// target being resynced. The worker pool runs for a single target (opts.arn),
// so an object that only qualifies for a different target must be skipped: with
// A/B rules and a resync of A, an object that needs replication only for B must
// not be admitted to A's worker. Otherwise (after outcome-based classification)
// A would be absent from that object's result and miscounted as an A failure.
func TestObjectNeedsResyncForARN(t *testing.T) {
const (
arnA = "arn:minio:replication::id:bucket"
arnB = "arn:minio:replication::id2:bucket"
)
tests := []struct {
name string
decision ResyncDecision
arn string
want bool
}{
{
name: "target must resync",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{arnA: {Replicate: true}}},
arn: arnA,
want: true,
},
{
name: "object qualifies for B only, resyncing A",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{arnB: {Replicate: true}}},
arn: arnA,
want: false,
},
{
name: "A present but not replicating, B replicating, resyncing A",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{
arnA: {Replicate: false},
arnB: {Replicate: true},
}},
arn: arnA,
want: false,
},
{
name: "object qualifies for both, resyncing A",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{
arnA: {Replicate: true},
arnB: {Replicate: true},
}},
arn: arnA,
want: true,
},
{
name: "no resync decision",
decision: ResyncDecision{},
arn: arnA,
want: false,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
roi := ReplicateObjectInfo{Name: "obj", Bucket: "bucket", ExistingObjResync: tc.decision}
if got := objectNeedsResyncForARN(roi, tc.arn); got != tc.want {
t.Fatalf("objectNeedsResyncForARN(arn=%s) = %v, want %v", tc.arn, got, tc.want)
}
})
}
}
// newMatchingReplicationPair returns a source/target pair that getReplicationAction must
// classify as replicateNone: same ETag, version id, size, modification time and content
// type. Any action other than replicateNone is therefore attributable to the object lock
// entries a caller adds on top.
func newMatchingReplicationPair() (ObjectInfo, minio.ObjectInfo) {
mtime := time.Date(2026, 9, 5, 10, 0, 0, 0, time.UTC)
size := int64(7)
src := ObjectInfo{
Bucket: "bucket",
Name: "object",
ETag: "d41d8cd98f00b204e9800998ecf8427e",
VersionID: "b0ff1d6e-0000-4000-8000-000000000001",
Size: size,
ActualSize: &size,
ModTime: mtime,
ContentType: "application/octet-stream",
UserDefined: map[string]string{"content-type": "application/octet-stream"},
}
tgt := minio.ObjectInfo{
ETag: src.ETag,
VersionID: src.VersionID,
Size: size,
LastModified: mtime,
ContentType: src.ContentType,
Metadata: http.Header{},
}
return src, tgt
}
// TestGetReplicationActionEmptyObjectLockValues covers the comparison of object lock entries
// whose value is empty. Removing retention from a version stores the mode and retain-until-date
// keys with empty values, while the target's HEAD response omits them entirely, so the two must
// compare equal or the version can never be reported as in sync. Cases 3 and 4 are synthetic
// comparison inputs, since a SILO target cannot return empty lock headers; cases 7 and 8 guard
// against over-normalizing.
func TestGetReplicationActionEmptyObjectLockValues(t *testing.T) {
var (
modeKey = strings.ToLower(xhttp.AmzObjectLockMode)
dateKey = strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
until = "2026-10-05T10:00:00.000Z"
)
emptyRetention := map[string]string{modeKey: "", dateKey: ""}
realRetention := map[string]string{modeKey: "GOVERNANCE", dateKey: until}
tests := []struct {
name string
srcMeta map[string]string
tgtHdr map[string]string
want replicationAction
}{
{"1-both-clean-never-had-retention", nil, nil, replicateNone},
{"2-source-present-empty-target-absent", emptyRetention, nil, replicateNone},
{"3-source-absent-target-present-empty", nil, emptyRetention, replicateNone},
{"4-both-present-empty", emptyRetention, emptyRetention, replicateNone},
{"5-both-governance-equal", realRetention, realRetention, replicateNone},
{"6-source-governance-target-absent", realRetention, nil, replicateMetadata},
{"7-source-empty-target-real-retention", emptyRetention, realRetention, replicateMetadata},
{"8-empty-user-metadata-is-not-normalized", map[string]string{"x-amz-meta-foo": ""}, nil, replicateMetadata},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgt := newMatchingReplicationPair()
for k, v := range test.srcMeta {
src.UserDefined[k] = v
}
for k, v := range test.tgtHdr {
tgt.Metadata.Set(k, v)
}
if got := getReplicationAction(src, tgt, replication.HealReplicationType); got != test.want {
t.Fatalf("getReplicationAction() = %q, want %q (source %v, target %v)", got, test.want, src.UserDefined, tgt.Metadata)
}
})
}
}
// TestEmptyRetentionValuesAreOmittedFromObjectResponseHeaders records why the target half of the
// comparison in getReplicationAction can never report an empty object lock entry:
// FilterObjectLockMetadata drops both keys because an empty mode is not a valid retention mode,
// and setObjectHeaders skips them when writing response headers. Neither filter reaches the
// replication wire: the empty entries are still carried by getCopyObjMetadata and sent by the
// metadata CopyObject, which is why the sender's comparison is what has to tolerate them.
// FilterObjectLockMetadata is also applied by CopyObject (cmd/object-handlers.go:1708), where it
// strips the source's lock metadata before the destination re-derives it from the request.
func TestEmptyRetentionValuesAreOmittedFromObjectResponseHeaders(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
meta := map[string]string{
modeKey: "",
dateKey: "",
"content-type": "application/octet-stream",
}
filtered := objectlock.FilterObjectLockMetadata(meta, false, false)
if _, ok := filtered[modeKey]; ok {
t.Errorf("FilterObjectLockMetadata() kept the empty lock mode key: %v", filtered)
}
if _, ok := filtered[dateKey]; ok {
t.Errorf("FilterObjectLockMetadata() kept the empty retain-until-date key: %v", filtered)
}
rec := httptest.NewRecorder()
if err := setObjectHeaders(t.Context(), rec, ObjectInfo{UserDefined: meta, ModTime: time.Now(), Size: 7}, nil, ObjectOptions{}); err != nil {
t.Fatalf("setObjectHeaders() = %v", err)
}
if v, ok := rec.Header()[http.CanonicalHeaderKey(xhttp.AmzObjectLockMode)]; ok {
t.Errorf("setObjectHeaders() emitted an empty lock mode header: %v", v)
}
if v, ok := rec.Header()[http.CanonicalHeaderKey(xhttp.AmzObjectLockRetainUntilDate)]; ok {
t.Errorf("setObjectHeaders() emitted an empty retain-until-date header: %v", v)
}
}
// fakeRetentionGetter answers GetObjectRetention with a fixed result and counts its calls.
type fakeRetentionGetter struct {
mode *minio.RetentionMode
err error
calls int
}
func (f *fakeRetentionGetter) GetObjectRetention(_ context.Context, _, _, _ string) (*minio.RetentionMode, *time.Time, error) {
f.calls++
return f.mode, nil, f.err
}
func TestRetentionRemovedAtSource(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
tsKey := ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp
stamp := "2026-09-06T01:00:00Z"
tests := []struct {
name string
meta map[string]string
want bool
}{
{"no lock keys", map[string]string{"content-type": "text/plain"}, false},
{"empty pair", map[string]string{modeKey: "", dateKey: ""}, true},
{"empty mode only", map[string]string{modeKey: ""}, true},
{"empty date only", map[string]string{dateKey: ""}, true},
{"real retention", map[string]string{modeKey: "GOVERNANCE", dateKey: "2026-10-05T10:00:00.000Z"}, false},
{"canonical case", map[string]string{xhttp.AmzObjectLockMode: ""}, true},
{"empty user metadata", map[string]string{"x-amz-meta-foo": ""}, false},
// Representation (2): a replicated removal persists the ordering timestamp alone,
// with the mode and retain-until-date keys absent (restoreRetention).
{"timestamp only, mode absent", map[string]string{tsKey: stamp}, true},
{"timestamp with empty mode", map[string]string{tsKey: stamp, modeKey: ""}, true},
{"timestamp with real retention is a set, not a removal", map[string]string{tsKey: stamp, modeKey: "GOVERNANCE", dateKey: "2026-10-05T10:00:00.000Z"}, false},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
if got := retentionRemovedAtSource(ObjectInfo{UserDefined: test.meta}); got != test.want {
t.Fatalf("retentionRemovedAtSource() = %v, want %v", got, test.want)
}
})
}
}
// TestTargetRetentionConfirmedAbsent pins the rule that only an explicit answer from the
// destination clears a removed retention. A denied or unreachable destination must read as still
// holding retention, because HEAD hides a real retention from a credential without
// s3:GetObjectRetention exactly as it hides one that does not exist.
func TestTargetRetentionConfirmedAbsent(t *testing.T) {
governance := minio.Governance
var emptyMode minio.RetentionMode
unknownMode := minio.RetentionMode("ARCHIVE")
tests := []struct {
name string
mode *minio.RetentionMode
err error
want bool
}{
{"version holds governance retention", &governance, nil, false},
{"no retention on the version", nil, minio.ErrorResponse{Code: "NoSuchObjectLockConfiguration"}, true},
{
// The destination also answers this when its own read of the bucket's Object Lock
// configuration fails, so it does not establish that Object Lock is disabled.
"invalid request naming a missing object lock configuration",
nil,
minio.ErrorResponse{Code: "InvalidRequest", Message: "Bucket is missing ObjectLockConfiguration"},
false,
},
{
"unrelated invalid request",
nil,
minio.ErrorResponse{Code: "InvalidRequest", Message: "Object is WORM protected and cannot be overwritten"},
false,
},
{"retention read denied", nil, minio.ErrorResponse{Code: "AccessDenied"}, false},
{"destination unreachable", nil, errors.New("dial tcp: connection refused"), false},
{"empty mode returned", &emptyMode, nil, true},
{"unknown non-empty mode returned", &unknownMode, nil, false},
{"nil mode returned", nil, nil, true},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
if got := targetRetentionConfirmedAbsent(t.Context(), tgt, "bucket", "object", "v1"); got != test.want {
t.Fatalf("targetRetentionConfirmedAbsent() = %v, want %v", got, test.want)
}
if tgt.calls != 1 {
t.Fatalf("GetObjectRetention called %d times, want 1", tgt.calls)
}
})
}
}
// TestReplicationActionForTargetRetentionRemoval covers the decision the replication worker makes
// for a version whose retention was removed. The destination's HEAD never reports the empty keys,
// so the comparison alone reads every one of these as in sync; only the confirmation separates a
// destination that really dropped the retention from one that is hiding it.
func TestReplicationActionForTargetRetentionRemoval(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
governance := minio.Governance
tests := []struct {
name string
srcMeta map[string]string
mode *minio.RetentionMode
err error
want replicationAction
wantCalls int
}{
{
name: "removal confirmed by destination",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
err: minio.ErrorResponse{Code: "NoSuchObjectLockConfiguration"},
want: replicateNone,
wantCalls: 1,
},
{
name: "destination still holds the retention hidden from HEAD",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
mode: &governance,
want: replicateMetadata,
wantCalls: 1,
},
{
name: "retention hidden from HEAD by permissions",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
err: minio.ErrorResponse{Code: "AccessDenied"},
want: replicateMetadata,
wantCalls: 1,
},
{
// A destination that names a missing Object Lock configuration answers the same way
// when its own read of that configuration failed, so it confirms nothing.
name: "destination reports no object lock configuration",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
err: minio.ErrorResponse{Code: "InvalidRequest", Message: "Bucket is missing ObjectLockConfiguration"},
want: replicateMetadata,
wantCalls: 1,
},
{
name: "version never had retention is not confirmed",
srcMeta: nil,
want: replicateNone,
wantCalls: 0,
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgtInfo := newMatchingReplicationPair()
for k, v := range test.srcMeta {
src.UserDefined[k] = v
}
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
got := replicationActionForTarget(t.Context(), src, tgtInfo, replication.HealReplicationType, tgt, "bucket", "object")
if got != test.want {
t.Fatalf("replicationActionForTarget() = %q, want %q", got, test.want)
}
if tgt.calls != test.wantCalls {
t.Fatalf("GetObjectRetention called %d times, want %d", tgt.calls, test.wantCalls)
}
})
}
}
// TestReplicationActionForTargetNullVersionResync pins that the confirmation does not reopen the
// null-version exclusion at the head of getReplicationAction. An existing object resync returns
// replicateNone for a null version whose source modification time is later than the target's,
// before comparing anything, and that must stand even when the source carries a removed retention
// and the destination would report retention or refuse to answer.
func TestReplicationActionForTargetNullVersionResync(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
governance := minio.Governance
tests := []struct {
name string
mode *minio.RetentionMode
err error
}{
{"destination holds retention", &governance, nil},
{"retention read denied", nil, minio.ErrorResponse{Code: "AccessDenied"}},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgtInfo := newMatchingReplicationPair()
// A null version whose source modification time is later, and whose content differs,
// so only the exclusion can hold the action at replicateNone.
src.VersionID = nullVersionID
src.ModTime = tgtInfo.LastModified.Add(time.Hour)
src.ETag = "5d41402abc4b2a76b9719d911017c592"
src.UserDefined[modeKey] = ""
src.UserDefined[dateKey] = ""
tgtInfo.VersionID = nullVersionID
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
got := replicationActionForTarget(t.Context(), src, tgtInfo, replication.ExistingObjectReplicationType, tgt, "bucket", "object")
if got != replicateNone {
t.Fatalf("replicationActionForTarget() = %q, want %q", got, replicateNone)
}
if tgt.calls != 0 {
t.Fatalf("GetObjectRetention called %d times, want 0", tgt.calls)
}
})
}
}
// TestReplicationActionForTargetTimestampOnlyRemoval covers representation (2) of a removed
// retention. A removal that arrived by replication persists only the retention ordering timestamp,
// with the mode and retain-until-date keys absent, because restoreRetention writes the timestamp
// alone when the mode is empty (cmd/bucket-object-lock.go). The comparison in getReplicationAction
// reads such a source as in sync with a matching destination, so only the GetObjectRetention
// confirmation separates a destination that dropped the retention from one hiding it behind a
// permission-filtered HEAD. The source is built through the real restoreRetention path so the
// fixture is the metadata a replicated removal actually leaves on disk, not a hand-rolled map.
func TestReplicationActionForTargetTimestampOnlyRemoval(t *testing.T) {
governance := minio.Governance
stamp := time.Date(2026, 9, 6, 1, 0, 0, 0, time.UTC).Format(time.RFC3339Nano)
tests := []struct {
name string
mode *minio.RetentionMode
err error
want replicationAction
wantCalls int
}{
{
// The reference case: a destination that denies the retention read is
// indistinguishable from one still holding it, so the removal is resent.
name: "retention hidden from HEAD by permissions",
err: minio.ErrorResponse{Code: "AccessDenied"},
want: replicateMetadata,
wantCalls: 1,
},
{
name: "destination still holds the retention",
mode: &governance,
want: replicateMetadata,
wantCalls: 1,
},
{
name: "removal confirmed by destination",
err: minio.ErrorResponse{Code: "NoSuchObjectLockConfiguration"},
want: replicateNone,
wantCalls: 1,
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgtInfo := newMatchingReplicationPair()
// Persist the timestamp-only tombstone the same way an applied replica removal does.
objectLockState{retentionTimestamp: stamp}.restoreRetention(src.UserDefined)
if !retentionRemovedAtSource(src) {
t.Fatalf("restoreRetention fixture not recognized as a removal: %v", src.UserDefined)
}
if _, ok := src.UserDefined[strings.ToLower(xhttp.AmzObjectLockMode)]; ok {
t.Fatalf("restoreRetention fixture wrote a mode key, fixture is not timestamp-only: %v", src.UserDefined)
}
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
got := replicationActionForTarget(t.Context(), src, tgtInfo, replication.HealReplicationType, tgt, "bucket", "object")
if got != test.want {
t.Fatalf("replicationActionForTarget() = %q, want %q", got, test.want)
}
if tgt.calls != test.wantCalls {
t.Fatalf("GetObjectRetention called %d times, want %d", tgt.calls, test.wantCalls)
}
})
}
}
+6 -2
View File
@@ -319,7 +319,9 @@ func (r *ReplicationStats) getNodeQueueStats(bucket string) (qs ReplQNodeStats)
qs.QStats = r.qCache.getBucketStats(bucket)
qs.TgtXferStats = make(map[string]map[RMetricName]XferStats)
qs.MRFStats = ReplicationMRFStats{
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
TotalDroppedCount: atomic.LoadUint64(&r.mrfStats.TotalDroppedCount),
TotalDroppedBytes: atomic.LoadUint64(&r.mrfStats.TotalDroppedBytes),
}
r.RLock()
@@ -410,7 +412,9 @@ func (r *ReplicationStats) getNodeQueueStatsSummary() (qs ReplQNodeStats) {
qs.XferStats = make(map[RMetricName]XferStats)
qs.QStats = r.qCache.getSiteStats()
qs.MRFStats = ReplicationMRFStats{
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
TotalDroppedCount: atomic.LoadUint64(&r.mrfStats.TotalDroppedCount),
TotalDroppedBytes: atomic.LoadUint64(&r.mrfStats.TotalDroppedBytes),
}
r.RLock()
defer r.RUnlock()
+3 -3
View File
@@ -97,7 +97,7 @@ func (api objectAPIHandlers) PutBucketVersioningHandler(w http.ResponseWriter, r
return
}
updatedAt, err := globalBucketMetadataSys.Update(ctx, bucket, bucketVersioningConfig, configData)
result, err := globalBucketMetadataSys.updateAndParseMetadata(ctx, bucket, bucketVersioningConfig, configData, false, false, nil)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -107,12 +107,12 @@ func (api objectAPIHandlers) PutBucketVersioningHandler(w http.ResponseWriter, r
//
// We encode the xml bytes as base64 to ensure there are no encoding
// errors.
cfgStr := base64.StdEncoding.EncodeToString(configData)
cfgStr := base64.StdEncoding.EncodeToString(result.meta.VersioningConfigXML)
replLogIf(ctx, globalSiteReplicationSys.BucketMetaHook(ctx, madmin.SRBucketMeta{
Type: madmin.SRBucketMetaTypeVersionConfig,
Bucket: bucket,
Versioning: &cfgStr,
UpdatedAt: updatedAt,
UpdatedAt: result.updatedAt,
}))
writeSuccessResponseHeadersOnly(w)
+54 -7
View File
@@ -120,15 +120,40 @@ func init() {
const consolePrefix = "CONSOLE_"
// consoleMinIOServerEnv derives the CONSOLE_MINIO_SERVER value the embedded
// Console uses to reach the S3/STS API, and whether TLS verification of that
// endpoint must be skipped. With no explicit endpoint configured the Console
// reaches the API over the loopback address, whose TLS certificate is not
// expected to carry a 127.0.0.1 SAN; because Console verifies outbound TLS by
// default, the loopback origin has to be exempted or embedded login (local and
// LDAP alike) fails at the STS handshake. The exemption is endpoint-scoped in
// Console, so every other HTTPS peer stays verified. An explicitly configured
// endpoint is always reached under its own verified name and is never exempted.
func consoleMinIOServerEnv(endpoint string, isTLS bool, port string) (server string, skipVerify bool) {
if endpoint != "" {
return endpoint, false
}
return fmt.Sprintf("%s://127.0.0.1:%s", getURLScheme(isTLS), port), isTLS
}
func minioConfigToConsoleFeatures() {
os.Setenv("CONSOLE_PBKDF_SALT", globalDeploymentID())
os.Setenv("CONSOLE_PBKDF_PASSPHRASE", globalDeploymentID())
if globalMinioEndpoint != "" {
os.Setenv("CONSOLE_MINIO_SERVER", globalMinioEndpoint)
consoleServer, skipVerify := consoleMinIOServerEnv(globalMinioEndpoint, globalIsTLS, globalMinioPort)
os.Setenv("CONSOLE_MINIO_SERVER", consoleServer)
if skipVerify {
// The embedded Console reaches the loopback S3/STS endpoint above, whose
// certificate is not expected to carry a 127.0.0.1 SAN. Console verifies
// outbound TLS by default (silo-console v2.3.x), so opt into the
// endpoint-scoped compatibility switch to preserve the documented loopback
// bypass; every other HTTPS peer (IdP, Prometheus, webhooks, ...) stays
// verified. initConsoleServer unsets CONSOLE_* before calling this, so the
// switch cannot be supplied by the operator on the embedded path.
os.Setenv("CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY", "on")
} else {
// Explicitly set 127.0.0.1 so Console will automatically bypass TLS verification to the local S3 API.
// This will save users from providing a certificate with IP or FQDN SAN that points to the local host.
os.Setenv("CONSOLE_MINIO_SERVER", fmt.Sprintf("%s://127.0.0.1:%s", getURLScheme(globalIsTLS), globalMinioPort))
// An explicitly configured endpoint is reached under its own verified name;
// never let a loopback exemption apply to it.
os.Unsetenv("CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY")
}
if value := env.Get(config.EnvMinIOLogQueryURL, ""); value != "" {
os.Setenv("CONSOLE_LOG_QUERY_URL", value)
@@ -232,11 +257,31 @@ func buildOpenIDConsoleConfig() consoleoauth2.OpenIDPCfg {
return m
}
func initConsoleServer() (*consoleapi.Server, error) {
// unset all console_ environment variables.
// resetConsoleEnvironment preserves the embedded Console's supported resource
// settings verbatim. Server derives all other Console settings itself.
func resetConsoleEnvironment() {
for _, cenv := range env.List(consolePrefix) {
switch cenv {
case consoleapi.ConsoleWSMaxConnections,
consoleapi.ConsoleWSMaxConnectionsPerClient,
consoleapi.ConsoleWSMaxAnonymousConnections,
consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient:
continue
}
os.Unsetenv(cenv)
}
}
func initConsoleServer() (*consoleapi.Server, error) {
resetConsoleEnvironment()
// Validate explicitly: ConfigureAPI logs errors, but embedded Console logs
// are normally silenced. Return configuration failures to Server startup.
if err := consoleapi.ConfigureEmbeddedSourceIPTrust(); err != nil {
return nil, err
}
if err := consoleapi.ConfigureWebSocketLimits(); err != nil {
return nil, err
}
// enable all console environment variables
minioConfigToConsoleFeatures()
@@ -400,6 +445,7 @@ func buildServerCtxt(ctx *cli.Context, ctxt *serverCtxt) (err error) {
ctxt.SendBufSize = ctx.Int("send-buf-size")
ctxt.RecvBufSize = ctx.Int("recv-buf-size")
ctxt.IdleTimeout = ctx.Duration("idle-timeout")
ctxt.ReadHeaderTimeout = ctx.Duration("read-header-timeout")
ctxt.UserTimeout = ctx.Duration("conn-user-timeout")
if conf := ctx.String("config"); len(conf) > 0 {
@@ -855,6 +901,7 @@ func serverHandleEnvVars() {
}
globalEnableSyncBoot = env.Get("MINIO_SYNC_BOOT", config.EnableOff) == config.EnableOn
globalSiteReplicationMetadataTombstones = env.Get("MINIO_SITE_REPLICATION_METADATA_TOMBSTONES", config.EnableOff) == config.EnableOn
}
func loadRootCredentials() auth.Credentials {
+136
View File
@@ -26,6 +26,7 @@ import (
"strings"
"testing"
consoleapi "github.com/minio/console/api"
"github.com/minio/minio/internal/config"
)
@@ -372,3 +373,138 @@ func TestConfigEnvFileNamedTargetDiscovery(t *testing.T) {
t.Fatalf("named target %q not discovered from %s: %v", "my-hook", key, targets)
}
}
// TestConsoleMinIOServerEnv locks in the loopback TLS exemption that keeps
// embedded Console login working (issue #108) while ensuring an explicitly
// configured endpoint is never silently exempted from TLS verification.
func TestConsoleMinIOServerEnv(t *testing.T) {
tests := []struct {
name string
endpoint string
isTLS bool
port string
wantServer string
wantSkipVerify bool
}{
{
name: "loopback TLS is exempted so embedded login works",
isTLS: true,
port: "9000",
wantServer: "https://127.0.0.1:9000",
wantSkipVerify: true,
},
{
name: "loopback plain HTTP needs no exemption",
isTLS: false,
port: "9000",
wantServer: "http://127.0.0.1:9000",
},
{
name: "explicit https endpoint stays verified",
endpoint: "https://silo.example:9000",
isTLS: true,
port: "9000",
wantServer: "https://silo.example:9000",
},
{
name: "explicit http endpoint stays verified",
endpoint: "http://silo.example:9000",
isTLS: false,
port: "9000",
wantServer: "http://silo.example:9000",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
server, skipVerify := consoleMinIOServerEnv(tt.endpoint, tt.isTLS, tt.port)
if server != tt.wantServer {
t.Fatalf("server = %q, want %q", server, tt.wantServer)
}
if skipVerify != tt.wantSkipVerify {
t.Fatalf("skipVerify = %v, want %v", skipVerify, tt.wantSkipVerify)
}
})
}
}
// The startup path clears process environment, so preserve all existing Console
// variables, including ones unrelated to this test, before exercising it.
func preserveConsoleEnvironment(t *testing.T) {
t.Helper()
for _, entry := range os.Environ() {
if strings.HasPrefix(entry, consolePrefix) {
name, value, _ := strings.Cut(entry, "=")
t.Setenv(name, value)
}
}
}
func TestResetConsoleEnvironment(t *testing.T) {
preserveConsoleEnvironment(t)
settings := map[string]string{
consoleapi.ConsoleWSMaxConnections: "2048",
consoleapi.ConsoleWSMaxConnectionsPerClient: "512",
consoleapi.ConsoleWSMaxAnonymousConnections: "128",
consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient: "16",
}
for key, value := range settings {
t.Setenv(key, value)
}
decoys := []string{"CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY", "CONSOLE_MINIO_SERVER", "CONSOLE_PBKDF_SALT", "CONSOLE_TRUSTED_PROXIES", "CONSOLE_WS_MAX_UNKNOWN"}
for _, key := range decoys {
t.Setenv(key, "operator-value")
}
resetConsoleEnvironment()
for key, want := range settings {
if got, present := os.LookupEnv(key); !present || got != want {
t.Errorf("%s = %q, present = %v; want %q", key, got, present, want)
}
}
for _, key := range decoys {
if _, present := os.LookupEnv(key); present {
t.Errorf("unsupported override %s survived", key)
}
}
for _, raw := range []string{"", " 16 ", "env://missing-limit"} {
t.Setenv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient, raw)
resetConsoleEnvironment()
if got, present := os.LookupEnv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient); !present || got != raw {
t.Fatalf("raw value %q was changed to %q (present = %v)", raw, got, present)
}
}
os.Unsetenv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient)
resetConsoleEnvironment()
if _, present := os.LookupEnv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient); present {
t.Fatal("unset setting became present")
}
}
func TestInitConsoleServerConfigurationErrors(t *testing.T) {
for _, tt := range []struct {
name, proxy, limit, want string
}{
{"proxy error precedes limit error", "proxy.internal", "bad", "MINIO_API_TRUSTED_PROXIES"},
{"blank limit", "", "", "CONSOLE_WS_MAX_ANONYMOUS_CONNECTIONS_PER_CLIENT"},
{"non-integer limit", "", "bad", "CONSOLE_WS_MAX_ANONYMOUS_CONNECTIONS_PER_CLIENT"},
{"out-of-range limit", "", "0", "CONSOLE_WS_MAX_ANONYMOUS_CONNECTIONS_PER_CLIENT"},
{"inconsistent limits", "", "256", "must be less than"},
} {
t.Run(tt.name, func(t *testing.T) {
// Restore the process-wide library configuration after environment cleanup.
t.Cleanup(func() {
_ = consoleapi.ConfigureEmbeddedSourceIPTrust()
_ = consoleapi.ConfigureWebSocketLimits()
})
preserveConsoleEnvironment(t)
t.Setenv(consoleapi.EnvMinIOTrustedProxies, tt.proxy)
t.Setenv(consoleapi.ConsoleWSMaxConnections, "1024")
t.Setenv(consoleapi.ConsoleWSMaxConnectionsPerClient, "256")
t.Setenv(consoleapi.ConsoleWSMaxAnonymousConnections, "64")
t.Setenv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient, tt.limit)
server, err := initConsoleServer()
if err == nil || !strings.Contains(err.Error(), tt.want) || server != nil {
t.Fatalf("initConsoleServer() = %v, %v; want nil server and %q error", server, err, tt.want)
}
})
}
}
+738
View File
@@ -0,0 +1,738 @@
// Copyright (c) 2015-2026 MinIO, Inc.
// Copyright (c) 2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"archive/tar"
"bytes"
"crypto/md5"
"encoding/base64"
"encoding/xml"
"io"
"net/http"
"net/http/httptest"
"strconv"
"testing"
"time"
"github.com/klauspost/compress/s2"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// disableCompression turns the global compression config off and returns a
// restore func. It is the replication destination that applies no transform of
// its own; setCopyChecksumCompression covers the enabled cases.
func disableCompression() func() {
globalCompressConfigMu.Lock()
previous := globalCompressConfig
globalCompressConfig.Enabled = false
globalCompressConfigMu.Unlock()
return func() {
globalCompressConfigMu.Lock()
globalCompressConfig = previous
globalCompressConfigMu.Unlock()
}
}
// ssecTestHeaders builds the customer key headers for a key made of the given
// repeated byte.
func ssecTestHeaders(b byte) map[string]string {
key := bytes.Repeat([]byte{b}, 32)
keyMD5 := md5.Sum(key)
return map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
}
// assertStoredSSECUncompressed requires the stored object to be SSE-C sealed
// and to carry no compression marker, so that "not compressed" is never
// reported for an object that is not encrypted either.
func assertStoredSSECUncompressed(t *testing.T, obj ObjectLayer, bucketName, object string) ObjectInfo {
t.Helper()
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if _, sealed := info.UserDefined[crypto.MetaSealedKeySSEC]; !sealed {
t.Fatalf("%s is not SSE-C sealed, the fixture proves nothing (userDefined=%v)", object, info.UserDefined)
}
if marker, compressed := info.UserDefined[ReservedMetadataPrefix+"compression"]; compressed {
t.Errorf("%s was stored as a compressed SSE-C object (compression=%q); such an object cannot be replicated",
object, marker)
}
return info
}
// assertSSECPlaintext GETs an SSE-C object with its customer key and requires
// the body to equal the plaintext.
func assertSSECPlaintext(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
bucketName, object string, sseHeaders map[string]string, want []byte,
) {
t.Helper()
req, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("GET %s status %d: %s", object, rec.Code, rec.Body.String())
}
if bytes.Equal(rec.Body.Bytes(), want) {
return
}
body := rec.Body.Bytes()
head := body
if len(head) > 16 {
head = head[:16]
}
t.Errorf("GET %s returned %d bytes, want the %d byte plaintext; first bytes % x",
object, len(body), len(want), head)
// Name the failure mode: a body that s2-decodes to the plaintext is the raw
// S2 stream of a compressed source shipped without its compression marker.
if decoded, derr := io.ReadAll(s2.NewReader(bytes.NewReader(body))); derr == nil && bytes.Equal(decoded, want) {
t.Errorf("the returned body is the raw S2 stream: s2-decoding it yields the %d byte plaintext", len(want))
}
}
// TestAPISSECCompressionReplicaStaysReadable replicates an SSE-C object written
// with compression enabled and allow_encryption=on, and requires the replica to
// read back as the source plaintext.
//
// The source object is read the way the replication worker reads it
// (ReplicationRequest, hence NoDecryption), its wire headers come from the
// production option builder putReplicationOpts, and the replica is written the
// way a destination that applies no transform of its own stores it: compression
// off and no default encryption.
//
// Before the SSE-C compression exclusion the source was stored as
// encrypt(s2(plaintext)) while putReplicationOpts dropped
// X-Minio-Internal-compression, so the replica decrypted to an S2 stream and a
// correct-key GET returned HTTP 200 with the wrong body.
func TestAPISSECCompressionReplicaStaysReadable(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionReplicaStaysReadable,
})
}
func testAPISSECCompressionReplicaStaysReadable(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
replicator := newObjectAttributesAuthzUser(t, instanceType, bucketName, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
sseHeaders := ssecTestHeaders(0x42)
// Highly compressible and comfortably above minCompressibleSize (4096).
data := bytes.Repeat([]byte("silo compressed ssec replication payload "), 8192)
t.Run(instanceType+"/single-put", func(t *testing.T) {
object := "replication/ssec-single.txt"
// --- Source side: compression ON with allow_encryption ON. ---
restore := setCopyChecksumCompression(true)
srcReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
srcRec := httptest.NewRecorder()
apiRouter.ServeHTTP(srcRec, srcReq)
if srcRec.Code != http.StatusOK {
restore()
t.Fatalf("source PUT status %d: %s", srcRec.Code, srcRec.Body.String())
}
sourceInfo := assertStoredSSECUncompressed(t, obj, bucketName, object)
t.Logf("source: stored size=%d compression=%q plaintext=%d", sourceInfo.Size,
sourceInfo.UserDefined[ReservedMetadataPrefix+"compression"], len(data))
// The replication worker's read: raw stored bytes, no decryption.
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{ReplicationRequest: true})
if err != nil {
restore()
t.Fatal(err)
}
sourceInfo = gr.ObjInfo
raw, err := io.ReadAll(gr)
gr.Close()
if err != nil {
restore()
t.Fatal(err)
}
replicationOpts, isMP, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
restore()
t.Fatalf("putReplicationOpts rejected the source: %v", err)
}
if isMP {
restore()
t.Fatal("single PUT source classified as multipart")
}
headers := map[string]string{}
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
headers[name] = values[0]
}
}
restore()
// --- Destination side: NO compression, NO default encryption. ---
restoreDst := disableCompression()
defer restoreDst()
replReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(raw)), bytes.NewReader(raw), replicator.AccessKey, replicator.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
replRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replRec, replReq)
if replRec.Code != http.StatusOK {
t.Fatalf("replica PUT status %d: %s", replRec.Code, replRec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("replica: stored size=%d compression=%q actual-size=%q", info.Size,
info.UserDefined[ReservedMetadataPrefix+"compression"],
info.UserDefined[ReservedMetadataPrefix+"actual-size"])
assertSSECPlaintext(t, apiRouter, credentials, bucketName, object, sseHeaders, data)
})
t.Run(instanceType+"/multipart", func(t *testing.T) {
object := "replication/ssec-mpu.txt"
restore := setCopyChecksumCompression(true)
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
newRec := httptest.NewRecorder()
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
restore()
t.Fatalf("source NewMultipart status %d: %s", newRec.Code, newRec.Body.String())
}
var sourceInit InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &sourceInit, int64(newRec.Body.Len())); err != nil {
restore()
t.Fatal(err)
}
partReq, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, sourceInit.UploadID, "1"),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
partRec := httptest.NewRecorder()
apiRouter.ServeHTTP(partRec, partReq)
if partRec.Code != http.StatusOK {
restore()
t.Fatalf("source PutPart status %d: %s", partRec.Code, partRec.Body.String())
}
completeBody, err := xml.Marshal(CompleteMultipartUpload{Parts: []CompletePart{
{PartNumber: 1, ETag: canonicalizeETag(partRec.Header()[xhttp.ETag][0])},
}})
if err != nil {
restore()
t.Fatal(err)
}
completeReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, sourceInit.UploadID),
int64(len(completeBody)), bytes.NewReader(completeBody), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
completeRec := httptest.NewRecorder()
apiRouter.ServeHTTP(completeRec, completeReq)
if completeRec.Code != http.StatusOK {
restore()
t.Fatalf("source Complete status %d: %s", completeRec.Code, completeRec.Body.String())
}
assertStoredSSECUncompressed(t, obj, bucketName, object)
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{ReplicationRequest: true})
if err != nil {
restore()
t.Fatal(err)
}
sourceInfo := gr.ObjInfo
rawPart, err := io.ReadAll(gr)
gr.Close()
if err != nil {
restore()
t.Fatal(err)
}
actualSize, err := sourceInfo.GetActualSize()
if err != nil {
restore()
t.Fatal(err)
}
t.Logf("source mpu: stored size=%d actual-size=%d rawRead=%d plaintext=%d compression=%q",
sourceInfo.Size, actualSize, len(rawPart), len(data),
sourceInfo.UserDefined[ReservedMetadataPrefix+"compression"])
replicationOpts, isMP, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
restore()
t.Fatalf("putReplicationOpts rejected the source: %v", err)
}
if !isMP {
restore()
t.Fatal("SSE-C multipart source not recognized as multipart")
}
replicationOpts.Internal.SourceMTime = time.Time{}
headers := map[string]string{}
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
headers[name] = values[0]
}
}
restore()
// --- Destination: no compression, no default encryption. ---
restoreDst := disableCompression()
defer restoreDst()
replNewReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, replicator.AccessKey, replicator.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
replNewRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replNewRec, replNewReq)
if replNewRec.Code != http.StatusOK {
t.Fatalf("replica NewMultipart status %d: %s", replNewRec.Code, replNewRec.Body.String())
}
var replicaInit InitiateMultipartUploadResponse
if err = xmlDecoder(replNewRec.Body, &replicaInit, int64(replNewRec.Body.Len())); err != nil {
t.Fatal(err)
}
replPartReq, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, replicaInit.UploadID, "1"),
int64(len(rawPart)), bytes.NewReader(rawPart), replicator.AccessKey, replicator.SecretKey,
map[string]string{xhttp.MinIOSourceReplicationRequest: "true"})
if err != nil {
t.Fatal(err)
}
replPartRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replPartRec, replPartReq)
if replPartRec.Code != http.StatusOK {
t.Fatalf("replica PutPart status %d: %s", replPartRec.Code, replPartRec.Body.String())
}
replCompleteBody, err := xml.Marshal(CompleteMultipartUpload{Parts: []CompletePart{
{PartNumber: 1, ETag: canonicalizeETag(replPartRec.Header()[xhttp.ETag][0])},
}})
if err != nil {
t.Fatal(err)
}
replCompleteReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, replicaInit.UploadID),
int64(len(replCompleteBody)), bytes.NewReader(replCompleteBody), replicator.AccessKey, replicator.SecretKey,
map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.MinIOSourceMTime: sourceInfo.ModTime.Format(time.RFC3339Nano),
xhttp.MinIOSourceETag: sourceInfo.ETag,
xhttp.MinIOReplicationActualObjectSize: strconv.FormatInt(actualSize, 10),
})
if err != nil {
t.Fatal(err)
}
replCompleteRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replCompleteRec, replCompleteReq)
if replCompleteRec.Code != http.StatusOK {
t.Fatalf("replica Complete status %d: %s", replCompleteRec.Code, replCompleteRec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
reportedActual, aerr := info.GetActualSize()
t.Logf("replica mpu: stored size=%d compression=%q actual-size=%q GetActualSize=%d(err=%v)",
info.Size, info.UserDefined[ReservedMetadataPrefix+"compression"],
info.UserDefined[ReservedMetadataPrefix+"actual-size"], reportedActual, aerr)
assertSSECPlaintext(t, apiRouter, credentials, bucketName, object, sseHeaders, data)
})
}
// TestAPISSECCompressionProducerMatrix pins the scope of the exclusion across
// the PutObject and NewMultipartUpload producers: SSE-C is never compressed,
// while plaintext, SSE-S3 and SSE-KMS keep following allow_encryption.
func TestAPISSECCompressionProducerMatrix(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionProducerMatrix,
})
}
func testAPISSECCompressionProducerMatrix(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
ssecHeaders := ssecTestHeaders(0x5a)
sseS3Headers := map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}
sseKMSHeaders := map[string]string{
xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS,
xhttp.AmzServerSideEncryptionKmsID: "compressed-ssec-producer-matrix",
}
previousKMS := GlobalKMS
GlobalKMS = kms.NewStub("compressed-ssec-producer-matrix")
defer func() { GlobalKMS = previousKMS }()
big := bytes.Repeat([]byte("silo producer matrix payload "), 8192)
small := bytes.Repeat([]byte("s"), 1024) // below minCompressibleSize
for _, tc := range []struct {
name string
allowEncrypted bool
headers map[string]string
body []byte
wantCompressed bool
}{
// SSE-C is excluded from compression in both configurations, because the
// replication wire cannot carry the compression state.
{"ssec+allow_encryption-on+large", true, ssecHeaders, big, false},
{"ssec+allow_encryption-on+small", true, ssecHeaders, small, false},
{"ssec+allow_encryption-off+large", false, ssecHeaders, big, false},
// Plaintext still compresses in both configurations.
{"plain+allow_encryption-on+large", true, nil, big, true},
{"plain+allow_encryption-off+large", false, nil, big, true},
// SSE-S3 and SSE-KMS are the reason allow_encryption exists: the server
// owns the key, so the source decompresses before replicating.
{"sse-s3+allow_encryption-on+large", true, sseS3Headers, big, true},
{"sse-s3+allow_encryption-off+large", false, sseS3Headers, big, false},
{"sse-kms+allow_encryption-on+large", true, sseKMSHeaders, big, true},
{"sse-kms+allow_encryption-off+large", false, sseKMSHeaders, big, false},
} {
t.Run(instanceType+"/put/"+tc.name, func(t *testing.T) {
restore := setCopyChecksumCompression(tc.allowEncrypted)
defer restore()
object := "producer/" + tc.name + ".txt"
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(tc.body)), bytes.NewReader(tc.body), credentials.AccessKey, credentials.SecretKey, tc.headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("PUT status %d, want %d: %s", rec.Code, http.StatusOK, rec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
_, compressed := info.UserDefined[ReservedMetadataPrefix+"compression"]
if compressed != tc.wantCompressed {
t.Errorf("compressed=%v, want %v (userDefined=%v)", compressed, tc.wantCompressed, info.UserDefined)
}
})
}
// NewMultipartUpload has no size gate, so the exclusion turns on the SSE-C
// and allow_encryption combination alone.
for _, tc := range []struct {
name string
allowEncrypted bool
headers map[string]string
wantCompressed bool
}{
{"ssec+allow_encryption-on", true, ssecHeaders, false},
{"ssec+allow_encryption-off", false, ssecHeaders, false},
{"plain+allow_encryption-on", true, nil, true},
} {
t.Run(instanceType+"/mpu/"+tc.name, func(t *testing.T) {
restore := setCopyChecksumCompression(tc.allowEncrypted)
defer restore()
object := "producer/mpu-" + tc.name + ".txt"
req, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, tc.headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("NewMultipartUpload status %d, want %d: %s", rec.Code, http.StatusOK, rec.Body.String())
}
var init InitiateMultipartUploadResponse
if err = xmlDecoder(rec.Body, &init, int64(rec.Body.Len())); err != nil {
t.Fatal(err)
}
mi, err := obj.GetMultipartInfo(t.Context(), bucketName, object, init.UploadID, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
_, compressed := mi.UserDefined[ReservedMetadataPrefix+"compression"]
if compressed != tc.wantCompressed {
t.Errorf("upload compressed=%v, want %v", compressed, tc.wantCompressed)
}
})
}
}
// TestAPISSECCompressionSkippedOnCopyObject covers the third producer:
// CopyObjectHandler decides compression before it encrypts, so a copy with a
// destination customer key and allow_encryption=on used to store a compressed
// SSE-C object from a plaintext source. It also covers the reverse direction,
// where only copy-source customer headers are present and compression must
// still apply.
func TestAPISSECCompressionSkippedOnCopyObject(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionSkippedOnCopyObject,
endpoints: []string{"CopyObject", "PutObject", "GetObject"},
})
}
func testAPISSECCompressionSkippedOnCopyObject(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
ssecHeaders := ssecTestHeaders(0x7c)
data := bytes.Repeat([]byte("copy object compressed ssec payload "), 8192)
restore := setCopyChecksumCompression(true)
defer restore()
// An unencrypted source, stored compressed because compression is on. Only
// the copy adds encryption, so only the copy can change the decision.
src := "copysrc/plain.txt"
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, src),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("source PUT status %d: %s", rec.Code, rec.Body.String())
}
dst := "copydst/ssec.txt"
copyHeaders := map[string]string{"X-Amz-Copy-Source": SlashSeparator + bucketName + SlashSeparator + src}
for k, v := range ssecHeaders {
copyHeaders[k] = v
}
copyReq, err := newTestSignedRequestV4(http.MethodPut, getCopyObjectURL("", bucketName, dst),
0, nil, credentials.AccessKey, credentials.SecretKey, copyHeaders)
if err != nil {
t.Fatal(err)
}
copyRec := httptest.NewRecorder()
apiRouter.ServeHTTP(copyRec, copyReq)
if copyRec.Code != http.StatusOK {
t.Fatalf("CopyObject status %d: %s", copyRec.Code, copyRec.Body.String())
}
assertStoredSSECUncompressed(t, obj, bucketName, dst)
assertSSECPlaintext(t, apiRouter, credentials, bucketName, dst, ssecHeaders, data)
// The plaintext source is untouched by the copy and stays compressed.
srcInfo, err := obj.GetObjectInfo(t.Context(), bucketName, src, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if _, compressed := srcInfo.UserDefined[ReservedMetadataPrefix+"compression"]; !compressed {
t.Errorf("the plaintext copy source lost compression (userDefined=%v)", srcInfo.UserDefined)
}
// The reverse direction: a copy-source customer key is not a destination
// key, so copying the SSE-C object on to a plaintext destination still
// compresses. crypto.SSEC.IsRequested ignores the copy-source headers.
plain := "copydst/decrypted.txt"
decryptHeaders := map[string]string{
"X-Amz-Copy-Source": SlashSeparator + bucketName + SlashSeparator + dst,
xhttp.AmzServerSideEncryptionCopyCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCopyCustomerKey: ssecHeaders[xhttp.AmzServerSideEncryptionCustomerKey],
xhttp.AmzServerSideEncryptionCopyCustomerKeyMD5: ssecHeaders[xhttp.AmzServerSideEncryptionCustomerKeyMD5],
}
decryptReq, err := newTestSignedRequestV4(http.MethodPut, getCopyObjectURL("", bucketName, plain),
0, nil, credentials.AccessKey, credentials.SecretKey, decryptHeaders)
if err != nil {
t.Fatal(err)
}
decryptRec := httptest.NewRecorder()
apiRouter.ServeHTTP(decryptRec, decryptReq)
if decryptRec.Code != http.StatusOK {
t.Fatalf("CopyObject to a plaintext destination status %d: %s", decryptRec.Code, decryptRec.Body.String())
}
plainInfo, err := obj.GetObjectInfo(t.Context(), bucketName, plain, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if _, sealed := plainInfo.UserDefined[crypto.MetaSealedKeySSEC]; sealed {
t.Fatalf("%s is still SSE-C sealed, the fixture proves nothing", plain)
}
if _, compressed := plainInfo.UserDefined[ReservedMetadataPrefix+"compression"]; !compressed {
t.Errorf("a copy carrying only copy-source SSE-C headers was not compressed (userDefined=%v)", plainInfo.UserDefined)
}
}
// TestAPISSECCompressionSkippedOnSnowballExtract covers the fourth producer:
// PutObjectExtractHandler decides compression per entry before it encrypts, so
// a tar extract carrying customer key headers used to store every entry as a
// compressed SSE-C object.
func TestAPISSECCompressionSkippedOnSnowballExtract(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionSkippedOnSnowballExtract,
})
}
func testAPISSECCompressionSkippedOnSnowballExtract(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
entry := "extracted/entry.txt"
payload := bytes.Repeat([]byte("snowball compressed ssec entry "), 4096)
var body bytes.Buffer
tw := tar.NewWriter(&body)
if err := tw.WriteHeader(&tar.Header{Name: entry, Mode: 0o600, Size: int64(len(payload))}); err != nil {
t.Fatal(err)
}
if _, err := tw.Write(payload); err != nil {
t.Fatal(err)
}
if err := tw.Close(); err != nil {
t.Fatal(err)
}
restore := setCopyChecksumCompression(true)
defer restore()
ssecHeaders := ssecTestHeaders(0x2d)
headers := map[string]string{xhttp.AmzSnowballExtract: "true"}
for k, v := range ssecHeaders {
headers[k] = v
}
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, "snowball.tar"),
int64(body.Len()), bytes.NewReader(body.Bytes()), credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("snowball extract status %d: %s", rec.Code, rec.Body.String())
}
assertStoredSSECUncompressed(t, obj, bucketName, entry)
assertSSECPlaintext(t, apiRouter, credentials, bucketName, entry, ssecHeaders, payload)
}
// TestSSECBatchReplicationCannotRead is the control for the corruption path:
// batch replication reads without ReplicationRequest, so NoDecryption is never
// set and a non-empty SSE-C source fails at read time. Batch replication
// therefore cannot reach the replica shape in
// TestAPISSECCompressionReplicaStaysReadable; it cannot replicate a non-empty
// SSE-C object at all, compressed or not. A zero-byte object takes the reader
// shortcut, whose key check passes without a customer key.
func TestSSECBatchReplicationCannotRead(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testSSECBatchReplicationCannotRead,
})
}
func testSSECBatchReplicationCannotRead(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
ssecHeaders := ssecTestHeaders(0x6b)
data := bytes.Repeat([]byte("batch ssec payload "), 8192)
object := "batch/ssec-plain.txt"
restore := disableCompression()
defer restore()
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, ssecHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("source PUT status %d: %s", rec.Code, rec.Body.String())
}
// The read shape used by BatchJobReplicateV1.ReplicateToTarget and
// writeAsArchive: no ReplicationRequest, so NoDecryption is never set.
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{}, ObjectOptions{})
if err == nil {
gr.Close()
t.Fatal("batch-shaped read of an SSE-C object unexpectedly succeeded")
}
t.Logf("batch-shaped read of an SSE-C object fails as expected: %v", err)
// Control: the replication worker's read shape succeeds and yields ciphertext.
gr2, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{ReplicationRequest: true})
if err != nil {
t.Fatalf("replication-shaped read failed: %v", err)
}
raw, err := io.ReadAll(gr2)
gr2.Close()
if err != nil {
t.Fatal(err)
}
if bytes.Equal(raw, data) {
t.Fatal("replication-shaped read returned plaintext")
}
t.Logf("replication-shaped read returns %d bytes of ciphertext (plaintext %d)", len(raw), len(data))
}
+89
View File
@@ -0,0 +1,89 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"encoding/json"
"os"
"reflect"
"testing"
)
// Read a real pre-removal scanner cache with nonzero hot-tier bytes. A v8
// payload with a relabeled header would not exercise the ignored hts field.
func TestDataUsageCacheReadV9(t *testing.T) {
raw, err := os.ReadFile("testdata/data-usage-v9/data-usage-v9.bin")
if err != nil {
t.Fatal(err)
}
if len(raw) == 0 || raw[0] != 9 {
t.Fatal("fixture is not v9")
}
expected, err := os.ReadFile("testdata/data-usage-v9/data-usage-v9.json")
if err != nil {
t.Fatal(err)
}
var want, got dataUsageCache
if err := json.Unmarshal(expected, &want); err != nil {
t.Fatal(err)
}
if err := got.deserialize(bytes.NewReader(raw)); err != nil {
t.Fatal(err)
}
if !got.Info.LastUpdate.Equal(want.Info.LastUpdate) {
t.Fatal("cache timestamp changed")
}
// msgp restores local time while JSON preserves the UTC representation.
want.Info.LastUpdate = got.Info.LastUpdate
if !reflect.DeepEqual(got, want) {
t.Fatalf("v9 ordinary cache fields changed:\ngot: %#v\nwant: %#v", got, want)
}
flat := got.flatten(*got.root())
if flat.Size != 74962 || flat.Objects != 3 || flat.Versions != 5 || flat.DeleteMarkers != 2 {
t.Fatalf("statistics changed: %+v", flat)
}
if flat.AllTierStats == nil || flat.AllTierStats.Tiers["COLD"] != (tierStats{TotalSize: 65536, NumVersions: 1, NumObjects: 1}) {
t.Fatalf("remote tier statistics changed: %+v", flat.AllTierStats)
}
buckets := []BucketInfo{{Name: "v9-bucket"}}
if gotInfo, wantInfo := got.dui(dataUsageRoot, buckets), want.dui(dataUsageRoot, buckets); !reflect.DeepEqual(gotInfo, wantInfo) {
t.Fatalf("bucket usage aggregation changed: %+v != %+v", gotInfo, wantInfo)
}
var buf bytes.Buffer
if err := got.serializeTo(&buf); err != nil {
t.Fatal(err)
}
if buf.Bytes()[0] != 8 {
t.Fatalf("wrote cache version %d, want 8", buf.Bytes()[0])
}
var roundtrip dataUsageCache
if err := roundtrip.deserialize(&buf); err != nil {
t.Fatal(err)
}
if !reflect.DeepEqual(got, roundtrip) {
t.Fatal("v8 round trip lost ordinary fields")
}
if err := roundtrip.deserialize(bytes.NewReader(raw[:len(raw)/2])); err == nil {
t.Fatal("truncated v9 cache accepted")
}
raw[0] = 10
if err := roundtrip.deserialize(bytes.NewReader(raw)); err == nil {
t.Fatal("unknown cache version accepted")
}
}
+2 -1
View File
@@ -981,6 +981,7 @@ func (d *dataUsageCache) save(ctx context.Context, store objectIO, name string)
// and write new data with the new version.
const (
dataUsageCacheVerCurrent = 8
dataUsageCacheVerV9 = 9 // Retired access-tier cache; only adds the ignored "hts" entry key.
dataUsageCacheVerV7 = 7
dataUsageCacheVerV6 = 6
dataUsageCacheVerV5 = 5
@@ -1182,7 +1183,7 @@ func (d *dataUsageCache) deserialize(r io.Reader) error {
}
return nil
case dataUsageCacheVerCurrent:
case dataUsageCacheVerCurrent, dataUsageCacheVerV9:
// Zstd compressed.
dec, err := zstd.NewReader(r, zstd.WithDecoderConcurrency(2))
if err != nil {
+5
View File
@@ -1068,6 +1068,11 @@ func (er erasureObjects) HealObject(ctx context.Context, bucket, object, version
newReqInfo = logger.NewReqInfo("", "", globalDeploymentID(), "", "Heal", bucket, object)
}
healCtx := logger.SetReqInfo(GlobalContext, newReqInfo)
if opts.NoLock {
// The caller owns the namespace lock. Stop if that lock's context is
// canceled instead of continuing metadata writes after losing it.
healCtx = logger.SetReqInfo(ctx, newReqInfo)
}
// Healing directories handle it separately.
if HasSuffix(object, SlashSeparator) {
+790
View File
@@ -0,0 +1,790 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"crypto/md5"
"encoding/base64"
"encoding/xml"
"fmt"
"io"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
"github.com/minio/minio/internal/hash"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/sio"
)
// TestAPISSECReplicaPartNumberReads replicates a three-part SSE-C multipart
// object through the trusted-replication write path and compares what
// GET ?partNumber=N returns before and after the replica overwrite.
func TestAPISSECReplicaPartNumberReads(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECReplicaPartNumberReads,
})
}
func testAPISSECReplicaPartNumberReads(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
replicator := newObjectAttributesAuthzUser(t, instanceType, bucketName, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
key := bytes.Repeat([]byte{0x42}, 32)
keyMD5 := md5.Sum(key)
sseHeaders := map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
const mib = 1024 * 1024
partLens := []int{5 * mib, 5 * mib, 1 * mib}
plaintext := make([]byte, 0, 11*mib)
partData := make([][]byte, len(partLens))
for i, n := range partLens {
b := make([]byte, n)
for j := range b {
// Distinct, position-dependent bytes so an off-by-N shift is visible.
b[j] = byte(i*7 + j%251)
}
partData[i] = b
plaintext = append(plaintext, b...)
}
object := "ssec-mp-3part"
// ---- 1. Build the source: a real three-part SSE-C multipart object. ----
newRec := httptest.NewRecorder()
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
t.Fatalf("source NewMultipart status %d: %s", newRec.Code, newRec.Body.String())
}
var srcInit InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &srcInit, int64(newRec.Body.Len())); err != nil {
t.Fatal(err)
}
srcParts := make([]CompletePart, len(partLens))
for i, b := range partData {
pn := strconv.Itoa(i + 1)
partReq, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, srcInit.UploadID, pn),
int64(len(b)), bytes.NewReader(b), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, partReq)
if rec.Code != http.StatusOK {
t.Fatalf("source PutPart %s status %d: %s", pn, rec.Code, rec.Body.String())
}
srcParts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
srcCompleteBody, err := xml.Marshal(CompleteMultipartUpload{Parts: srcParts})
if err != nil {
t.Fatal(err)
}
completeReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, srcInit.UploadID), int64(len(srcCompleteBody)),
bytes.NewReader(srcCompleteBody), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
completeRec := httptest.NewRecorder()
apiRouter.ServeHTTP(completeRec, completeReq)
if completeRec.Code != http.StatusOK {
t.Fatalf("source Complete status %d: %s", completeRec.Code, completeRec.Body.String())
}
// ---- 2. Record the source's per-part metadata and per-part GET answers. ----
srcOI, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] SOURCE parts:", instanceType)
for _, p := range srcOI.Parts {
t.Logf(" part %d Size=%d ActualSize=%d", p.Number, p.Size, p.ActualSize)
}
srcActual, err := srcOI.GetActualSize()
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] SOURCE object Size=%d GetActualSize=%d actual-size-meta=%q",
instanceType, srcOI.Size, srcActual, srcOI.UserDefined[ReservedMetadataPrefix+"actual-size"])
type getResult struct {
status int
clen string
crange string
body []byte
}
doGet := func(query string, extra map[string]string) getResult {
hdrs := map[string]string{}
for k, v := range sseHeaders {
hdrs[k] = v
}
for k, v := range extra {
hdrs[k] = v
}
u := getGetObjectURL("", bucketName, object) + query
req, err := newTestSignedRequestV4(http.MethodGet, u, 0, nil, credentials.AccessKey, credentials.SecretKey, hdrs)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return getResult{
status: rec.Code,
clen: rec.Header().Get(xhttp.ContentLength),
crange: rec.Header().Get(xhttp.ContentRange),
body: append([]byte(nil), rec.Body.Bytes()...),
}
}
doHead := func(query string) getResult {
u := getGetObjectURL("", bucketName, object) + query
req, err := newTestSignedRequestV4(http.MethodHead, u, 0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return getResult{
status: rec.Code,
clen: rec.Header().Get(xhttp.ContentLength),
crange: rec.Header().Get(xhttp.ContentRange),
}
}
queries := []string{"?partNumber=1", "?partNumber=2", "?partNumber=3"}
srcGets := make([]getResult, len(queries))
for i, q := range queries {
srcGets[i] = doGet(q, nil)
t.Logf("[%s] SOURCE GET %s -> status=%d Content-Length=%s Content-Range=%s len(body)=%d",
instanceType, q, srcGets[i].status, srcGets[i].clen, srcGets[i].crange, len(srcGets[i].body))
}
// Range GET crossing the part1/part2 boundary.
boundaryRange := fmt.Sprintf("bytes=%d-%d", 5*mib-16, 5*mib+15)
srcRange := doGet("", map[string]string{"Range": boundaryRange})
t.Logf("[%s] SOURCE GET Range %s -> status=%d Content-Length=%s len(body)=%d",
instanceType, boundaryRange, srcRange.status, srcRange.clen, len(srcRange.body))
// Sanity: the source must return exactly the part bytes.
for i := range partData {
if !bytes.Equal(srcGets[i].body, partData[i]) {
t.Fatalf("[%s] SOURCE partNumber=%d returned wrong bytes (len %d want %d)",
instanceType, i+1, len(srcGets[i].body), len(partData[i]))
}
}
// ---- 3. Read the raw ciphertext the replication worker would ship. ----
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{}, ObjectOptions{ReplicationRequest: true})
if err != nil {
t.Fatal(err)
}
sourceInfo := gr.ObjInfo
rawAll, err := io.ReadAll(gr)
gr.Close()
if err != nil {
t.Fatal(err)
}
if bytes.Equal(rawAll, plaintext) {
t.Fatal("source replication read did not return encrypted bytes")
}
rawParts := make([][]byte, len(sourceInfo.Parts))
off := int64(0)
for i, p := range sourceInfo.Parts {
rawParts[i] = rawAll[off : off+p.Size]
off += p.Size
}
if off != int64(len(rawAll)) {
t.Fatalf("raw ciphertext length %d != sum of part sizes %d", len(rawAll), off)
}
// ---- 4. Replicate onto the same key through the trusted write path. ----
replicationOpts, isMP, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
t.Fatal(err)
}
if !isMP {
t.Fatal("SSE-C multipart source was not recognized as multipart")
}
replicationOpts.Internal.SourceMTime = time.Time{}
replicationHeaders := make(map[string]string)
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
replicationHeaders[name] = values[0]
}
}
replNewReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, replicator.AccessKey, replicator.SecretKey, replicationHeaders)
if err != nil {
t.Fatal(err)
}
replNewRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replNewRec, replNewReq)
if replNewRec.Code != http.StatusOK {
t.Fatalf("replica NewMultipart status %d: %s", replNewRec.Code, replNewRec.Body.String())
}
var replInit InitiateMultipartUploadResponse
if err = xmlDecoder(replNewRec.Body, &replInit, int64(replNewRec.Body.Len())); err != nil {
t.Fatal(err)
}
replParts := make([]CompletePart, len(rawParts))
for i, raw := range rawParts {
pn := strconv.Itoa(i + 1)
req, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, replInit.UploadID, pn),
int64(len(raw)), bytes.NewReader(raw), replicator.AccessKey, replicator.SecretKey,
map[string]string{xhttp.MinIOSourceReplicationRequest: "true"})
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("replica PutPart %s status %d: %s", pn, rec.Code, rec.Body.String())
}
replParts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
replCompleteBody, err := xml.Marshal(CompleteMultipartUpload{Parts: replParts})
if err != nil {
t.Fatal(err)
}
srcActualSize, err := sourceInfo.GetActualSize()
if err != nil {
t.Fatal(err)
}
replCompleteHeaders := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.MinIOSourceMTime: sourceInfo.ModTime.Format(time.RFC3339Nano),
xhttp.MinIOSourceETag: sourceInfo.ETag,
xhttp.MinIOReplicationActualObjectSize: strconv.FormatInt(srcActualSize, 10),
}
replCompleteReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, replInit.UploadID), int64(len(replCompleteBody)),
bytes.NewReader(replCompleteBody), replicator.AccessKey, replicator.SecretKey, replCompleteHeaders)
if err != nil {
t.Fatal(err)
}
replCompleteRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replCompleteRec, replCompleteReq)
if replCompleteRec.Code != http.StatusOK {
t.Fatalf("replica Complete status %d: %s", replCompleteRec.Code, replCompleteRec.Body.String())
}
// ---- 5. Same reads against the replica. ----
repOI, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] REPLICA parts:", instanceType)
for _, p := range repOI.Parts {
t.Logf(" part %d Size=%d ActualSize=%d", p.Number, p.Size, p.ActualSize)
}
repActual, err := repOI.GetActualSize()
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] REPLICA object Size=%d GetActualSize=%d actual-size-meta=%q",
instanceType, repOI.Size, repActual, repOI.UserDefined[ReservedMetadataPrefix+"actual-size"])
// Whole-object GET must still be byte-identical.
whole := doGet("", nil)
if whole.status != http.StatusOK || !bytes.Equal(whole.body, plaintext) {
t.Errorf("[%s] REPLICA whole-object GET: status=%d len=%d want %d, equal=%v",
instanceType, whole.status, len(whole.body), len(plaintext), bytes.Equal(whole.body, plaintext))
} else {
t.Logf("[%s] REPLICA whole-object GET: OK, %d bytes identical", instanceType, len(whole.body))
}
// Fresh replica parts must record the plaintext lengths, not the
// ciphertext lengths the sender shipped.
if len(repOI.Parts) != len(partLens) {
t.Fatalf("[%s] REPLICA has %d parts, want %d", instanceType, len(repOI.Parts), len(partLens))
}
for i, p := range repOI.Parts {
if p.ActualSize != int64(partLens[i]) {
t.Errorf("[%s] REPLICA part %d ActualSize=%d, want the uploaded length %d",
instanceType, p.Number, p.ActualSize, partLens[i])
}
}
// Expected framing from independent prefix sums of the uploaded lengths.
total := len(plaintext)
start := 0
for i, q := range queries {
wantLen := partLens[i]
wantRange := fmt.Sprintf("bytes %d-%d/%d", start, start+wantLen-1, total)
start += wantLen
got := doGet(q, nil)
want := srcGets[i]
if got.status != want.status || got.status != http.StatusPartialContent {
t.Errorf("[%s] REPLICA GET %s status=%d, source=%d, want 206", instanceType, q, got.status, want.status)
}
if got.clen != strconv.Itoa(wantLen) || got.crange != wantRange {
t.Errorf("[%s] REPLICA GET %s Content-Length=%s Content-Range=%s, want %d and %q",
instanceType, q, got.clen, got.crange, wantLen, wantRange)
}
if want.clen != strconv.Itoa(wantLen) || want.crange != wantRange {
t.Errorf("[%s] SOURCE GET %s Content-Length=%s Content-Range=%s, want %d and %q",
instanceType, q, want.clen, want.crange, wantLen, wantRange)
}
if len(got.body) != wantLen {
t.Errorf("[%s] REPLICA GET %s body is %d bytes, want %d", instanceType, q, len(got.body), wantLen)
}
if !bytes.Equal(got.body, want.body) {
firstDiff := -1
for k := 0; k < len(got.body) && k < len(want.body); k++ {
if got.body[k] != want.body[k] {
firstDiff = k
break
}
}
t.Errorf("[%s] REPLICA partNumber=%d returned DIFFERENT bytes than the source: got %d bytes (Content-Range %q), want %d bytes (Content-Range %q), first differing byte at %d",
instanceType, i+1, len(got.body), got.crange, len(want.body), want.crange, firstDiff)
}
head := doHead(q)
if head.status != http.StatusPartialContent || head.clen != strconv.Itoa(wantLen) || head.crange != wantRange {
t.Errorf("[%s] REPLICA HEAD %s status=%d Content-Length=%s Content-Range=%s, want 206, %d and %q",
instanceType, q, head.status, head.clen, head.crange, wantLen, wantRange)
}
}
repRange := doGet("", map[string]string{"Range": boundaryRange})
sameRange := bytes.Equal(repRange.body, srcRange.body)
t.Logf("[%s] REPLICA GET Range %s -> status=%d Content-Length=%s len(body)=%d | source len=%d | bytes-equal=%v",
instanceType, boundaryRange, repRange.status, repRange.clen, len(repRange.body), len(srcRange.body), sameRange)
if !sameRange {
t.Errorf("[%s] REPLICA boundary Range GET returned different bytes", instanceType)
}
}
// TestSSECReplicaPartActualSizeDataMovement reproduces what a decommission or
// rebalance does to a replica whose parts already carry the ciphertext length in
// ActualSize: it replays the object through the object layer exactly the way
// decommissionObject does (cmd/erasure-server-pool-decom.go:605-667), passing the
// stale ActualSize to PutObjectPart and completing without ReplicationRequest, so
// CompleteMultipartUpload recomputes the object-level actual-size from the sum of
// part ActualSizes (cmd/erasure-multipart.go:1365,1440).
func TestSSECReplicaPartActualSizeDataMovement(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testSSECReplicaPartActualSizeDataMovement,
})
}
func testSSECReplicaPartActualSizeDataMovement(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
key := bytes.Repeat([]byte{0x37}, 32)
keyMD5 := md5.Sum(key)
sseHeaders := map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
const mib = 1024 * 1024
partLens := []int{5 * mib, 5 * mib, 1 * mib}
plaintext := make([]byte, 0, 11*mib)
partData := make([][]byte, len(partLens))
for i, n := range partLens {
b := make([]byte, n)
for j := range b {
b[j] = byte(i*13 + j%241)
}
partData[i] = b
plaintext = append(plaintext, b...)
}
object := "ssec-mp-datamovement"
newRec := httptest.NewRecorder()
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
t.Fatalf("NewMultipart status %d: %s", newRec.Code, newRec.Body.String())
}
var init InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &init, int64(newRec.Body.Len())); err != nil {
t.Fatal(err)
}
srcParts := make([]CompletePart, len(partLens))
for i, b := range partData {
req, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, init.UploadID, strconv.Itoa(i+1)),
int64(len(b)), bytes.NewReader(b), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("PutPart %d status %d: %s", i+1, rec.Code, rec.Body.String())
}
srcParts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
body, err := xml.Marshal(CompleteMultipartUpload{Parts: srcParts})
if err != nil {
t.Fatal(err)
}
cReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, init.UploadID), int64(len(body)),
bytes.NewReader(body), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
cRec := httptest.NewRecorder()
apiRouter.ServeHTTP(cRec, cReq)
if cRec.Code != http.StatusOK {
t.Fatalf("Complete status %d: %s", cRec.Code, cRec.Body.String())
}
// Replay it the way decommissionObject does, but hand PutObjectPart the STALE
// ActualSize an already-written SSE-C replica carries: the ciphertext length.
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{NoDecryption: true, NoLock: true, NoAuditLog: true})
if err != nil {
t.Fatal(err)
}
oi := gr.ObjInfo
res, err := obj.NewMultipartUpload(t.Context(), bucketName, object, ObjectOptions{
UserDefined: oi.UserDefined,
DataMovement: true,
NoAuditLog: true,
})
if err != nil {
gr.Close()
t.Fatal(err)
}
moved := make([]CompletePart, len(oi.Parts))
for i, part := range oi.Parts {
staleActual := part.Size // what a bad replica records
hr, herr := hash.NewReader(t.Context(), io.LimitReader(gr, part.Size), part.Size, "", "", staleActual)
if herr != nil {
gr.Close()
t.Fatal(herr)
}
pi, perr := obj.PutObjectPart(t.Context(), bucketName, object, res.UploadID, part.Number,
NewPutObjReader(hr), ObjectOptions{
PreserveETag: part.ETag,
IndexCB: func() []byte { return part.Index },
NoAuditLog: true,
})
if perr != nil {
gr.Close()
t.Fatalf("data-movement PutObjectPart part %d: %v", part.Number, perr)
}
moved[i] = CompletePart{ETag: pi.ETag, PartNumber: pi.PartNumber}
}
gr.Close()
// decommissionObject/rebalanceObject complete WITHOUT ReplicationRequest, so
// the object-level actual-size is recomputed from the part ActualSizes.
if _, err = obj.CompleteMultipartUpload(t.Context(), bucketName, object, res.UploadID, moved,
ObjectOptions{DataMovement: true, MTime: oi.ModTime, NoAuditLog: true}); err != nil {
t.Fatalf("data-movement CompleteMultipartUpload: %v", err)
}
after, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] AFTER DATA MOVEMENT parts:", instanceType)
for _, p := range after.Parts {
t.Logf(" part %d Size=%d ActualSize=%d", p.Number, p.Size, p.ActualSize)
}
gotActual, err := after.GetActualSize()
if err != nil {
t.Fatalf("GetActualSize after data movement: %v", err)
}
t.Logf("[%s] AFTER DATA MOVEMENT object Size=%d GetActualSize=%d actual-size-meta=%q",
instanceType, after.Size, gotActual, after.UserDefined[ReservedMetadataPrefix+"actual-size"])
wantActual := int64(len(plaintext))
if gotActual != wantActual {
t.Errorf("[%s] object-level actual size after data movement = %d, want %d",
instanceType, gotActual, wantActual)
}
for i, p := range after.Parts {
if p.ActualSize != int64(partLens[i]) {
t.Errorf("[%s] part %d ActualSize after data movement = %d, want %d",
instanceType, p.Number, p.ActualSize, partLens[i])
}
}
getReq, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
getRec := httptest.NewRecorder()
apiRouter.ServeHTTP(getRec, getReq)
if getRec.Code != http.StatusOK || !bytes.Equal(getRec.Body.Bytes(), plaintext) {
t.Errorf("[%s] whole-object GET after data movement: status=%d len=%d want %d equal=%v",
instanceType, getRec.Code, getRec.Body.Len(), len(plaintext), bytes.Equal(getRec.Body.Bytes(), plaintext))
}
// The advertised length must be the body length: a poisoned object-level
// actual-size shows up here as a Content-Length larger than the body.
if clen := getRec.Header().Get(xhttp.ContentLength); clen != strconv.Itoa(len(plaintext)) || clen != strconv.Itoa(getRec.Body.Len()) {
t.Errorf("[%s] whole-object GET after data movement advertises Content-Length=%s for a %d-byte body (plaintext %d)",
instanceType, clen, getRec.Body.Len(), len(plaintext))
}
}
// TestPartNumberToRangeSpecEncryptedParts pins the read-side repair: for an
// encrypted, uncompressed object the part range is derived from the stored
// ciphertext length, so a replica whose parts still record the ciphertext
// length in ActualSize reads correctly, while plaintext and compressed objects
// keep using ActualSize, and a part whose length cannot be a valid encrypted
// stream is reported as tampered by both callers. See pgsty/silo#119.
func TestPartNumberToRangeSpecEncryptedParts(t *testing.T) {
const mib = 1024 * 1024
plain := []int64{5 * mib, 5 * mib, 1024}
cipher := make([]int64, len(plain))
for i, n := range plain {
c, err := sio.EncryptedSize(uint64(n))
if err != nil {
t.Fatal(err)
}
cipher[i] = int64(c)
}
sum := func(v []int64) (s int64) {
for _, n := range v {
s += n
}
return s
}
encMeta := map[string]string{
crypto.MetaSealedKeySSEC: "sealed-key",
crypto.MetaIV: "iv",
crypto.MetaAlgorithm: crypto.InsecureSealAlgorithm,
}
compressedEncMeta := map[string]string{ReservedMetadataPrefix + "compression": compressionAlgorithmV2}
for k, v := range encMeta {
compressedEncMeta[k] = v
}
mkParts := func(sizes, actual []int64) []ObjectPartInfo {
parts := make([]ObjectPartInfo, len(sizes))
for i := range sizes {
parts[i] = ObjectPartInfo{Number: i + 1, Size: sizes[i], ActualSize: actual[i]}
}
return parts
}
// A compressed part's ActualSize is the uploaded length before compression,
// which bears no relation to the ciphertext length: use lengths whose
// decrypted size differs from ActualSize so that dropping the compression
// exclusion is detectable.
uploaded := []int64{2 * plain[0], 2 * plain[1], 2 * plain[2]}
wantRangeOf := func(lens []int64, pn int) (start, end int64) {
for i := 0; i < pn-1; i++ {
start += lens[i]
}
return start, start + lens[pn-1] - 1
}
for _, tc := range []struct {
name string
oi ObjectInfo
lens []int64
}{
{"encrypted, stale ciphertext ActualSize", ObjectInfo{Size: sum(cipher), UserDefined: encMeta, Parts: mkParts(cipher, cipher)}, plain},
{"encrypted, correct ActualSize", ObjectInfo{Size: sum(cipher), UserDefined: encMeta, Parts: mkParts(cipher, plain)}, plain},
{"plaintext", ObjectInfo{Size: sum(plain), UserDefined: map[string]string{}, Parts: mkParts(plain, plain)}, plain},
{"compressed and encrypted keeps ActualSize", ObjectInfo{Size: sum(cipher), UserDefined: compressedEncMeta, Parts: mkParts(cipher, uploaded)}, uploaded},
} {
for pn := 1; pn <= len(plain); pn++ {
rs, err := partNumberToRangeSpec(tc.oi, pn)
if err != nil {
t.Fatalf("%s: partNumber=%d: %v", tc.name, pn, err)
}
start, end := wantRangeOf(tc.lens, pn)
if rs == nil || rs.Start != start || rs.End != end {
t.Errorf("%s: partNumber=%d range %+v, want %d-%d", tc.name, pn, rs, start, end)
}
}
}
// A 31-byte part cannot be a sio stream: both callers report it as tampered.
bad := ObjectInfo{
Size: 31 + cipher[1] + cipher[2],
UserDefined: encMeta,
Parts: mkParts([]int64{31, cipher[1], cipher[2]}, []int64{31, plain[1], plain[2]}),
}
for pn := 1; pn <= 2; pn++ {
if _, err := partNumberToRangeSpec(bad, pn); err != errObjectTampered {
t.Errorf("malformed part: partNumber=%d err=%v, want errObjectTampered", pn, err)
}
}
if _, _, _, err := NewGetObjectReader(nil, bad, ObjectOptions{PartNumber: 1}, http.Header{}); err != errObjectTampered {
t.Errorf("NewGetObjectReader on a malformed part: err=%v, want errObjectTampered", err)
}
if err := setObjectHeaders(t.Context(), httptest.NewRecorder(), bad, nil, ObjectOptions{PartNumber: 1}); err != errObjectTampered {
t.Errorf("setObjectHeaders on a malformed part: err=%v, want errObjectTampered", err)
}
// A zero-length trailing part is a valid (empty) stream and stays accepted.
zero := ObjectInfo{Size: cipher[0], UserDefined: encMeta, Parts: mkParts([]int64{cipher[0], 0}, []int64{cipher[0], 0})}
rs, err := partNumberToRangeSpec(zero, 2)
if err != nil || rs == nil || rs.Start != plain[0] {
t.Errorf("zero-length trailing part: range %+v err=%v, want start %d", rs, err, plain[0])
}
}
// TestAPISSECReplicaMalformedPartIsRejected asserts that a trusted SSE-C replica
// part whose ciphertext length cannot be a valid encrypted stream is rejected
// as tampered before the part is committed. See pgsty/silo#119.
func TestAPISSECReplicaMalformedPartIsRejected(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPISSECReplicaMalformedPartIsRejected})
}
func testAPISSECReplicaMalformedPartIsRejected(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
replicator := newObjectAttributesAuthzUser(t, instanceType, bucketName, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
key := bytes.Repeat([]byte{0x45}, 32)
keyMD5 := md5.Sum(key)
sseHeaders := map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
object := "ssec-replica-malformed"
data := bytes.Repeat([]byte("malformed-part-"), 1024)
putReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object), int64(len(data)),
bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
putRec := httptest.NewRecorder()
apiRouter.ServeHTTP(putRec, putReq)
if putRec.Code != http.StatusOK {
t.Fatalf("%s: source PUT %d: %s", instanceType, putRec.Code, putRec.Body.String())
}
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{}, ObjectOptions{ReplicationRequest: true})
if err != nil {
t.Fatal(err)
}
sourceInfo := gr.ObjInfo
raw, err := io.ReadAll(gr)
gr.Close()
if err != nil {
t.Fatal(err)
}
replicationOpts, _, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
t.Fatal(err)
}
replicationOpts.Internal.SourceMTime = time.Time{}
replicationHeaders := make(map[string]string)
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
replicationHeaders[name] = values[0]
}
}
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, replicator.AccessKey, replicator.SecretKey, replicationHeaders)
if err != nil {
t.Fatal(err)
}
newRec := httptest.NewRecorder()
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
t.Fatalf("%s: replica NewMultipart %d: %s", instanceType, newRec.Code, newRec.Body.String())
}
var init InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &init, int64(newRec.Body.Len())); err != nil {
t.Fatal(err)
}
partHeaders := map[string]string{xhttp.MinIOSourceReplicationRequest: "true"}
// 31 bytes cannot be a sio stream (the package header plus its
// authentication tag occupy 32 bytes).
badReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectPartURL("", bucketName, object, init.UploadID, "1"),
31, bytes.NewReader(raw[:31]), replicator.AccessKey, replicator.SecretKey, partHeaders)
if err != nil {
t.Fatal(err)
}
badRec := httptest.NewRecorder()
apiRouter.ServeHTTP(badRec, badReq)
if badRec.Code == http.StatusOK || !strings.Contains(badRec.Body.String(), "XMinioObjectTampered") {
t.Fatalf("%s: malformed replica part answered %d: %s", instanceType, badRec.Code, badRec.Body.String())
}
lpi, err := obj.ListObjectParts(t.Context(), bucketName, object, init.UploadID, 0, 10, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if len(lpi.Parts) != 0 {
t.Fatalf("%s: malformed replica part was committed: %+v", instanceType, lpi.Parts)
}
// Control: the real ciphertext still uploads on the same upload.
goodReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectPartURL("", bucketName, object, init.UploadID, "1"),
int64(len(raw)), bytes.NewReader(raw), replicator.AccessKey, replicator.SecretKey, partHeaders)
if err != nil {
t.Fatal(err)
}
goodRec := httptest.NewRecorder()
apiRouter.ServeHTTP(goodRec, goodReq)
if goodRec.Code != http.StatusOK {
t.Fatalf("%s: valid replica part answered %d: %s", instanceType, goodRec.Code, goodRec.Body.String())
}
lpi, err = obj.ListObjectParts(t.Context(), bucketName, object, init.UploadID, 0, 10, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if len(lpi.Parts) != 1 || lpi.Parts[0].ActualSize != int64(len(data)) {
t.Fatalf("%s: valid replica part recorded %+v, want one part with ActualSize %d", instanceType, lpi.Parts, len(data))
}
}
@@ -157,6 +157,13 @@ func copyPartWithoutChecksumHTTP(t *testing.T, apiRouter http.Handler, creds aut
if sourceRange != "" {
req.Header.Set(xhttp.AmzCopySourceRange, sourceRange)
}
// Re-sign so the copy-source x-amz-* headers are covered by the signature,
// as real S3 clients send them; the verifier rejects unsigned x-amz-*.
if creds.AccessKey != "" && creds.SecretKey != "" {
if err := signRequestV4(req, creds.AccessKey, creds.SecretKey); err != nil {
t.Fatalf("failed to re-sign UploadPartCopy request: %v", err)
}
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
+72 -26
View File
@@ -710,20 +710,26 @@ func (er erasureObjects) PutObjectPart(ctx context.Context, bucket, object, uplo
}
actualSize := data.ActualSize()
if actualSize < 0 {
_, encrypted := crypto.IsEncrypted(fi.Metadata)
compressed := fi.IsCompressed()
switch {
case compressed:
// ... nothing changes for compressed stream.
// if actualSize is -1 we have no known way to
// determine what is the actualSize.
case encrypted:
decSize, err := sio.DecryptedSize(uint64(n))
if err == nil {
actualSize = int64(decSize)
}
default:
_, encrypted := crypto.IsEncrypted(fi.Metadata)
compressed := fi.IsCompressed()
switch {
case compressed:
// ... nothing changes for compressed stream.
// if actualSize is -1 we have no known way to
// determine what is the actualSize.
case encrypted:
// The uploaded length of an encrypted part is always derivable from the
// bytes just written, and the caller's value cannot be trusted: trusted
// SSE-C replication and the data movement paths hand over the ciphertext
// length. Derive it with the arithmetic the read path applies to
// part.Size, so the stored value matches how the part is read back.
decSize, err := sio.DecryptedSize(uint64(n))
if err != nil {
return pi, toObjectErr(errObjectTampered, bucket, object, uploadID)
}
actualSize = int64(decSize)
default:
if actualSize < 0 {
actualSize = n
}
}
@@ -1108,7 +1114,7 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
auditObjectErasureSet(ctx, "CompleteMultipartUpload", object, &er)
}
if opts.CheckPrecondFn != nil {
if opts.CheckPrecondFn != nil || opts.ReplicaLockReconcile {
if !opts.NoLock {
ns := er.NewNSLock(bucket, object)
lkctx, err := ns.GetLock(ctx, globalOperationTimeout)
@@ -1120,18 +1126,24 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
opts.NoLock = true
}
obj, err := er.getObjectInfo(ctx, bucket, object, opts)
if err == nil && opts.CheckPrecondFn(obj) {
return ObjectInfo{}, PreConditionFailed{}
}
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
return ObjectInfo{}, err
}
// The Object Lock reconcile below needs the version being committed, read
// after checkUploadIDExists, so only the precondition read happens here;
// both run under this same write lock, held until the version is renamed
// into place.
if opts.CheckPrecondFn != nil {
obj, err := er.getObjectInfo(ctx, bucket, object, opts)
if err == nil && opts.CheckPrecondFn(obj) {
return ObjectInfo{}, PreConditionFailed{}
}
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
return ObjectInfo{}, err
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return ObjectInfo{}, err
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return ObjectInfo{}, err
}
}
}
@@ -1143,6 +1155,40 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
return oi, toObjectErr(err, bucket, object, uploadID)
}
// Reconcile against the upload's persisted version, under the object lock.
// Multi-pool callers supply a resolver spanning all pools, including those
// draining their contents; the upload itself stays in its original pool.
if opts.ReplicaLockReconcile {
// A persisted upload records the null version as an empty VersionID; look
// it up as the null version so the reconcile reads the addressed version's
// stored lock, not the latest version's.
lookupVersionID := fi.VersionID
if lookupVersionID == "" {
lookupVersionID = nullVersionID
}
getObjectInfo := er.getObjectInfo
if opts.replicaObjectInfo != nil {
getObjectInfo = opts.replicaObjectInfo
}
curr, gerr := getObjectInfo(ctx, bucket, object, ObjectOptions{
VersionID: lookupVersionID,
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoLock: true,
})
switch {
case gerr == nil:
reconcileStoredObjectLock(fi.Metadata, storedObjectLockState(curr.UserDefined))
reconcileStoredObjectTags(fi.Metadata, curr.UserTags, curr.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
case isErrVersionNotFound(gerr) || isErrObjectNotFound(gerr):
// No existing version to order against: keep the upload's own accepted
// lock, including a pre-upgrade upload that persisted values without
// their ordering timestamps.
default:
return oi, toObjectErr(gerr, bucket, object)
}
}
uploadIDPath := er.getUploadIDDir(bucket, object, uploadID)
onlineDisks := er.getDisks()
writeQuorum := fi.WriteQuorum(er.defaultWQuorum())
@@ -0,0 +1,295 @@
// Copyright (c) 2015-2025 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"net/http"
"testing"
)
// TestDeleteObjectConditional verifies that a conditional DeleteObject
// (If-Match, wired through opts.CheckPrecondFn) is evaluated atomically at the
// object layer: a non-matching ETag must fail with PreConditionFailed and leave
// the object intact, a matching ETag must delete it, and an If-Match against a
// missing object must return a not-found error rather than silently succeeding.
func TestDeleteObjectConditional(t *testing.T) {
ctx := context.Background()
obj, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatal(err)
}
defer obj.Shutdown(context.Background())
defer removeRoots(fsDirs)
bucket := "test-bucket"
object := "test-object"
if err = obj.MakeBucket(ctx, bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
if _, err = obj.PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewReader([]byte("test-value")),
int64(len("test-value")), "", ""), ObjectOptions{}); err != nil {
t.Fatal(err)
}
objInfo, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
existingETag := objInfo.ETag
// If-Match with a wrong ETag must fail and preserve the object.
t.Run("wrong-etag-precondition-failed", func(t *testing.T) {
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, "wrong-etag")
},
}
if _, err := obj.DeleteObject(ctx, bucket, object, opts); !isErrPreconditionFailed(err) {
t.Errorf("expected PreConditionFailed, got: %v", err)
}
if _, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{}); err != nil {
t.Errorf("object must still exist after a failed conditional delete, got: %v", err)
}
})
// If-Match against a missing object must return a not-found error.
t.Run("missing-object-not-found", func(t *testing.T) {
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, existingETag)
},
}
_, err := obj.DeleteObject(ctx, bucket, "does-not-exist", opts)
if !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
t.Errorf("expected ObjectNotFound/VersionNotFound, got: %v", err)
}
})
// If-Match with the correct ETag must delete the object (run last).
t.Run("correct-etag-succeeds", func(t *testing.T) {
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, existingETag)
},
}
if _, err := obj.DeleteObject(ctx, bucket, object, opts); err != nil {
t.Errorf("expected a successful delete with matching ETag, got: %v", err)
}
if _, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{}); !isErrObjectNotFound(err) {
t.Errorf("object must be removed after a matching conditional delete, got: %v", err)
}
})
}
// TestDeleteObjectConditionalWithReadQuorumFailure verifies that a conditional
// (If-Match) DeleteObject does NOT proceed when the object's current state
// cannot be read due to read-quorum loss: without a verified ETag the delete
// must fail rather than remove the object blindly.
func TestDeleteObjectConditionalWithReadQuorumFailure(t *testing.T) {
ctx := context.Background()
obj, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatal(err)
}
defer obj.Shutdown(context.Background())
defer removeRoots(fsDirs)
z := obj.(*erasureServerPools)
xl := z.serverPools[0].sets[0]
bucket := "test-bucket"
object := "test-object"
if err = obj.MakeBucket(ctx, bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
if _, err = obj.PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewReader([]byte("test-value")),
int64(len("test-value")), "", ""), ObjectOptions{}); err != nil {
t.Fatal(err)
}
objInfo, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
existingETag := objInfo.ETag
// Simulate read-quorum loss by taking 8 of 16 disks offline (EC 8+8).
erasureDisks := xl.getDisks()
z.serverPools[0].erasureDisksMu.Lock()
xl.getDisks = func() []StorageAPI {
for i := range erasureDisks[:8] {
erasureDisks[i] = nil
}
return erasureDisks
}
z.serverPools[0].erasureDisksMu.Unlock()
// Even with the correct ETag we must not delete: the current state (hence the
// ETag) cannot be verified under read-quorum loss.
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, existingETag)
},
}
if _, err := obj.DeleteObject(ctx, bucket, object, opts); err == nil {
t.Error("expected an error for a conditional delete under read-quorum loss, got nil (object may have been deleted without ETag verification)")
}
}
// TestDeleteObjectConditionalVersioned verifies conditional DeleteObject on a
// versioned bucket, where the precondition is evaluated at the server-pool layer
// against the version that will actually be removed:
// - If-Match "*" when the latest version is a delete marker must fail (412),
// because there is no live object to match.
// - An explicit versionId If-Match is evaluated against the addressed version,
// not the latest one (match deletes it, mismatch is refused).
// - An If-Match against a missing version returns VersionNotFound, which the
// handler maps to NoSuchVersion.
func TestDeleteObjectConditionalVersioned(t *testing.T) {
ctx := context.Background()
obj, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatal(err)
}
defer obj.Shutdown(context.Background())
defer removeRoots(fsDirs)
bucket := "test-bucket"
if err = obj.MakeBucket(ctx, bucket, MakeBucketOptions{VersioningEnabled: true}); err != nil {
t.Fatal(err)
}
versioned := globalBucketVersioningSys.PrefixEnabled(bucket, "any")
if !versioned {
t.Fatalf("expected versioning to be enabled on %q", bucket)
}
put := func(object, content string) ObjectInfo {
oi, perr := obj.PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewReader([]byte(content)), int64(len(content)), "", ""),
ObjectOptions{Versioned: versioned})
if perr != nil {
t.Fatalf("put %q: %v", object, perr)
}
return oi
}
ifMatch := func(value string) CheckPreconditionFn {
return func(oi ObjectInfo) bool {
return deleteIfMatchPreconditionFailed(http.Header{}, value, oi)
}
}
// If-Match "*" against a delete-marker-latest must fail with 412.
t.Run("wildcard-on-delete-marker-latest", func(t *testing.T) {
object := "dm-object"
put(object, "v1")
// Create a delete marker (unconditional), making the latest a delete marker.
if _, derr := obj.DeleteObject(ctx, bucket, object, ObjectOptions{Versioned: versioned}); derr != nil {
t.Fatalf("create delete marker: %v", derr)
}
opts := ObjectOptions{Versioned: versioned, HasIfMatch: true, CheckPrecondFn: ifMatch("*")}
if _, derr := obj.DeleteObject(ctx, bucket, object, opts); !isErrPreconditionFailed(derr) {
t.Errorf("expected PreConditionFailed for If-Match:* on a delete-marker-latest, got: %v", derr)
}
})
// Explicit versionId is evaluated against the addressed (older) version.
t.Run("explicit-version-selection", func(t *testing.T) {
object := "ver-object"
v1 := put(object, "first")
v2 := put(object, "second-longer") // v2 is now the latest with a different ETag
if v1.ETag == v2.ETag {
t.Fatalf("test setup: versions must have distinct ETags")
}
// Mismatch: delete v2 with v1's ETag must be refused, v2 preserved.
mismatch := ObjectOptions{Versioned: versioned, VersionID: v2.VersionID, HasIfMatch: true, CheckPrecondFn: ifMatch(v1.ETag)}
if _, derr := obj.DeleteObject(ctx, bucket, object, mismatch); !isErrPreconditionFailed(derr) {
t.Errorf("expected PreConditionFailed deleting v2 with v1 ETag, got: %v", derr)
}
if _, gerr := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: v2.VersionID}); gerr != nil {
t.Errorf("v2 must still exist after a refused conditional delete, got: %v", gerr)
}
// Match: delete v1 with v1's ETag must succeed even though v1 is not latest.
match := ObjectOptions{Versioned: versioned, VersionID: v1.VersionID, HasIfMatch: true, CheckPrecondFn: ifMatch(v1.ETag)}
if _, derr := obj.DeleteObject(ctx, bucket, object, match); derr != nil {
t.Errorf("expected the addressed version to be deleted, got: %v", derr)
}
if _, gerr := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: v1.VersionID}); !isErrVersionNotFound(gerr) {
t.Errorf("v1 must be gone after a matching conditional delete, got: %v", gerr)
}
if _, gerr := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: v2.VersionID}); gerr != nil {
t.Errorf("v2 must remain after deleting v1, got: %v", gerr)
}
})
// If-Match against a missing version on an EXISTING key returns VersionNotFound.
t.Run("missing-version", func(t *testing.T) {
object := "missing-version-object"
put(object, "only")
opts := ObjectOptions{Versioned: versioned, VersionID: mustGetUUID(), HasIfMatch: true, CheckPrecondFn: ifMatch("anything")}
if _, derr := obj.DeleteObject(ctx, bucket, object, opts); !isErrVersionNotFound(derr) {
t.Errorf("expected VersionNotFound for If-Match on a missing version, got: %v", derr)
}
})
// If-Match against a missing version on an ABSENT key must also return
// VersionNotFound (NoSuchVersion), not NoSuchKey: the request addresses a
// specific version, which does not exist regardless of the key.
t.Run("missing-version-absent-key", func(t *testing.T) {
opts := ObjectOptions{Versioned: versioned, VersionID: mustGetUUID(), HasIfMatch: true, CheckPrecondFn: ifMatch("anything")}
if _, derr := obj.DeleteObject(ctx, bucket, "never-existed", opts); !isErrVersionNotFound(derr) {
t.Errorf("expected VersionNotFound for If-Match on a version of an absent key, got: %v", derr)
}
})
// If-Match "*" addressing a delete-marker VERSION by id must fail with 412,
// not 405: a delete marker has no entity-tag to match. getObjectInfo returns
// the marker alongside MethodNotAllowed; the precondition runs on the marker.
t.Run("wildcard-on-explicit-delete-marker-version", func(t *testing.T) {
object := "explicit-dm-object"
put(object, "live")
dm, derr := obj.DeleteObject(ctx, bucket, object, ObjectOptions{Versioned: versioned})
if derr != nil {
t.Fatalf("create delete marker: %v", derr)
}
if !dm.DeleteMarker || dm.VersionID == "" {
t.Fatalf("expected a delete-marker version, got DeleteMarker=%v VersionID=%q", dm.DeleteMarker, dm.VersionID)
}
opts := ObjectOptions{Versioned: versioned, VersionID: dm.VersionID, HasIfMatch: true, CheckPrecondFn: ifMatch("*")}
if _, derr := obj.DeleteObject(ctx, bucket, object, opts); !isErrPreconditionFailed(derr) {
t.Errorf("expected PreConditionFailed for If-Match:* on an addressed delete-marker version, got: %v", derr)
}
})
}
+32 -8
View File
@@ -133,6 +133,11 @@ func (er erasureObjects) CopyObject(ctx context.Context, srcBucket, srcObject, d
return fi.ToObjectInfo(srcBucket, srcObject, srcOpts.Versioned || srcOpts.VersionSuspended), toObjectErr(errMethodNotAllowed, srcBucket, srcObject)
}
if dstOpts.ReplicaLockReconcile {
reconcileStoredObjectLock(srcInfo.UserDefined, storedObjectLockState(fi.Metadata))
reconcileStoredObjectTags(srcInfo.UserDefined, fi.Metadata[xhttp.AmzObjectTagging], fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp])
}
filterOnlineDisksInplace(fi, metaArr, onlineDisks)
versionID := srcInfo.VersionID
@@ -1268,7 +1273,7 @@ func (er erasureObjects) putObject(ctx context.Context, bucket string, object st
data := r.Reader
if opts.CheckPrecondFn != nil {
if opts.CheckPrecondFn != nil || opts.ReplicaLockReconcile {
if !opts.NoLock {
ns := er.NewNSLock(bucket, object)
lkctx, err := ns.GetLock(ctx, globalOperationTimeout)
@@ -1280,18 +1285,33 @@ func (er erasureObjects) putObject(ctx context.Context, bucket string, object st
opts.NoLock = true
}
obj, err := er.getObjectInfo(ctx, bucket, object, opts)
if err == nil && opts.CheckPrecondFn(obj) {
return objInfo, PreConditionFailed{}
getObjectInfo := er.getObjectInfo
if opts.replicaObjectInfo != nil {
getObjectInfo = opts.replicaObjectInfo
}
obj, err := getObjectInfo(ctx, bucket, object, opts)
// A destination read that fails for a reason other than not-found must not
// be taken as a passed precondition or as absent lock state.
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
return objInfo, err
}
if opts.CheckPrecondFn != nil {
if err == nil && opts.CheckPrecondFn(obj) {
return objInfo, PreConditionFailed{}
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return objInfo, err
}
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return objInfo, err
// Re-read under the write lock, using the pools-layer resolver when
// present. A missing version keeps the incoming accepted state; an
// existing version contributes independently ordered lock and tags.
if opts.ReplicaLockReconcile && err == nil {
reconcileStoredObjectLock(opts.UserDefined, storedObjectLockState(obj.UserDefined))
reconcileStoredObjectTags(opts.UserDefined, obj.UserTags, obj.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
}
}
@@ -2307,7 +2327,11 @@ func (er erasureObjects) PutObjectTags(ctx context.Context, bucket, object strin
fi.Metadata[xhttp.AmzObjectTagging] = tags
fi.ReplicationState = opts.PutReplicationState()
stamp := monotonicTaggingTimestamp(opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp])
maps.Copy(fi.Metadata, opts.UserDefined)
if stamp != "" {
fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = stamp
}
if err = er.updateObjectMeta(ctx, bucket, object, fi, onlineDisks); err != nil {
return ObjectInfo{}, toObjectErr(err, bucket, object)
+392
View File
@@ -0,0 +1,392 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"context"
"maps"
"sort"
"strings"
"time"
madmin "github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/bucket/lifecycle"
xhttp "github.com/minio/minio/internal/http"
)
// objectPoolInfos reads the addressed version in every pool, including pools
// draining their contents. The caller must hold the pools-layer object lock.
// An unreadable pool may contain a newer version or metadata; it is not absence.
func (z *erasureServerPools) objectPoolInfos(ctx context.Context, bucket, object string, opts ObjectOptions) ([]PoolObjInfo, error) {
opts.NoLock = true
opts.CheckPrecondFn = nil
var copies []PoolObjInfo
for i, pool := range z.serverPools {
oi, err := pool.GetObjectInfo(ctx, bucket, object, opts)
if err == nil || (oi.DeleteMarker && (isErrObjectNotFound(err) || isErrMethodNotAllowed(err))) {
copies = append(copies, PoolObjInfo{Index: i, ObjInfo: oi})
continue
}
if !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
return nil, err
}
}
sort.Slice(copies, func(i, j int) bool {
a, b := copies[i], copies[j]
if a.ObjInfo.ModTime.Equal(b.ObjInfo.ModTime) {
return a.Index < b.Index
}
return a.ObjInfo.ModTime.After(b.ObjInfo.ModTime)
})
if len(copies) == 0 {
if opts.VersionID != "" {
return nil, VersionNotFound{Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID}
}
return nil, ObjectNotFound{Bucket: bucket, Object: decodeDirObject(object)}
}
return copies, nil
}
// Healing also commits object metadata. Both queued healing and drive healing
// must use the same namespace as pooled writes, rather than racing them under
// a destination set's independent namespace.
func (z *erasureServerPools) healObjectInPool(ctx context.Context, er *erasureObjects, bucket, object, versionID string, opts madmin.HealOpts) (madmin.HealResultItem, error) {
if !z.SinglePool() && !opts.NoLock {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return madmin.HealResultItem{}, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
opts.NoLock = true
}
return er.HealObject(ctx, bucket, object, versionID, opts)
}
// mergePoolLockState resolves each independently ordered field. An ordered
// removal is a value in its own right; an absent unordered field is not one.
func mergePoolLockState(copies []PoolObjInfo) objectLockState {
state := storedObjectLockState(copies[0].ObjInfo.UserDefined)
for _, copy := range copies[1:] {
next := storedObjectLockState(copy.ObjInfo.UserDefined)
retentionTime, _ := time.Parse(time.RFC3339Nano, next.retentionTimestamp)
if state.retentionIsOlderThan(retentionTime) || (state.retentionTimestamp == "" && state.mode == "" && next.mode != "") {
state.mode, state.retainUntil, state.retentionTimestamp = next.mode, next.retainUntil, next.retentionTimestamp
}
holdTime, _ := time.Parse(time.RFC3339Nano, next.legalHoldTimestamp)
if state.legalHoldIsOlderThan(holdTime) || (state.legalHoldTimestamp == "" && state.legalHold == "" && next.legalHold != "") {
state.legalHold, state.legalHoldTimestamp = next.legalHold, next.legalHoldTimestamp
}
}
return state
}
func replaceObjectLockMetadata(metadata map[string]string, state objectLockState) {
for _, key := range []string{
strings.ToLower(xhttp.AmzObjectLockMode), strings.ToLower(xhttp.AmzObjectLockRetainUntilDate),
strings.ToLower(xhttp.AmzObjectLockLegalHold),
ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp,
ReservedMetadataPrefixLower + ObjectLockLegalHoldTimestamp,
} {
// Metadata updates use maps.Copy at the set layer. Empty values also
// clear a previous value there, including timestamp-only removals.
if _, exists := metadata[key]; exists {
metadata[key] = ""
}
}
state.restoreRetention(metadata)
state.restoreLegalHold(metadata)
}
func mergedPoolObjectInfo(copies []PoolObjInfo) ObjectInfo {
oi := copies[0].ObjInfo
oi.UserDefined = maps.Clone(oi.UserDefined)
if oi.UserDefined == nil {
oi.UserDefined = make(map[string]string)
}
replaceObjectLockMetadata(oi.UserDefined, mergePoolLockState(copies))
for _, copy := range copies[1:] {
ts := copy.ObjInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]
stamp, _ := time.Parse(time.RFC3339Nano, ts)
if olderThan(oi.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], stamp) {
oi.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = ts
oi.UserTags = copy.ObjInfo.UserTags
}
}
return oi
}
func (z *erasureServerPools) replicaObjectInfo(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
copies, err := z.objectPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
if copies[0].ObjInfo.DeleteMarker {
return copies[0].ObjInfo, MethodNotAllowed{Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID}
}
return mergedPoolObjectInfo(copies), nil
}
// metadataPoolInfos resolves an unqualified metadata request to one logical
// version before collecting its copies, rather than mixing different versions.
func (z *erasureServerPools) metadataPoolInfos(ctx context.Context, bucket, object string, opts ObjectOptions) ([]PoolObjInfo, error) {
copies, err := z.objectPoolInfos(ctx, bucket, object, opts)
if err != nil {
return nil, err
}
if copies[0].ObjInfo.DeleteMarker {
return nil, MethodNotAllowed{Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID}
}
if opts.VersionID == "" {
opts.VersionID = copies[0].ObjInfo.VersionID
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
return z.objectPoolInfos(ctx, bucket, object, opts)
}
return copies, nil
}
func (z *erasureServerPools) updatePoolMetadata(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
copies, err := z.metadataPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
updated := mergedPoolObjectInfo(copies)
before := maps.Clone(updated.UserDefined)
if opts.EvalMetadataFn != nil {
if _, err := opts.EvalMetadataFn(&updated, nil); err != nil {
return ObjectInfo{}, err
}
}
changes := make(map[string]string)
for key, value := range updated.UserDefined {
if prior, ok := before[key]; !ok || prior != value {
changes[key] = value
}
}
for key := range before {
if _, ok := updated.UserDefined[key]; !ok {
changes[key] = ""
}
}
// Metadata callbacks may explicitly replace tags in the raw write map.
// Otherwise retain the merged value from the ObjectInfo read model.
tags, ok := updated.UserDefined[xhttp.AmzObjectTagging]
if !ok {
tags = updated.UserTags
}
state := storedObjectLockState(updated.UserDefined)
opts.VersionID = updated.VersionID
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
opts.NoLock = true
opts.EvalMetadataFn = func(oi *ObjectInfo, _ error) (ReplicateDecision, error) {
maps.Copy(oi.UserDefined, changes)
replaceObjectLockMetadata(oi.UserDefined, state)
// Reassemble the tag value and its ordering timestamp for storage.
if tags != "" || oi.UserTags != "" {
oi.UserDefined[xhttp.AmzObjectTagging] = tags
}
key := ReservedMetadataPrefixLower + TaggingTimestamp
if value, exists := updated.UserDefined[key]; exists || oi.UserDefined[key] != "" {
oi.UserDefined[key] = value
}
return ReplicateDecision{}, nil
}
var primary ObjectInfo
for _, copy := range copies {
oi, err := z.serverPools[copy.Index].PutObjectMetadata(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
if copy.Index == copies[0].Index {
primary = oi
}
}
return primary, nil
}
// Pass the stored tag value explicitly: ObjectInfo.UserDefined excludes it,
// whereas FileInfo.Metadata retains the raw storage key.
func reconcileStoredObjectTags(metadata map[string]string, storedTags, storedTimestamp string) {
key := ReservedMetadataPrefixLower + TaggingTimestamp
stamp, err := time.Parse(time.RFC3339Nano, storedTimestamp)
if err != nil {
return
}
incoming, err := time.Parse(time.RFC3339Nano, metadata[key])
if err != nil || !stamp.Before(incoming) {
metadata[key] = storedTimestamp
metadata[xhttp.AmzObjectTagging] = storedTags
}
}
// Local tagging mutations must advance the revision they overwrite, even when
// a request's clock or lock acquisition order is behind the stored revision.
// Replica writes use reconcileStoredObjectTags instead of minting a revision.
func monotonicTaggingTimestamp(incoming, stored string) string {
requested, err := time.Parse(time.RFC3339Nano, incoming)
if err != nil {
return incoming
}
current, err := time.Parse(time.RFC3339Nano, stored)
if err != nil || requested.After(current) {
return incoming
}
return current.Add(time.Nanosecond).UTC().Format(time.RFC3339Nano)
}
// A restored version still owns its tier reference even while IsRemote is
// false. Only the last copy of a reference may schedule its contents for GC.
func sharesTierObject(oi ObjectInfo, copies []PoolObjInfo) bool {
ref := oi.TransitionedObject
if ref.Status != lifecycle.TransitionComplete {
return false
}
for _, copy := range copies {
other := copy.ObjInfo.TransitionedObject
if other.Status == lifecycle.TransitionComplete && ref.Tier == other.Tier && ref.Name == other.Name && ref.VersionID == other.VersionID {
return true
}
}
return false
}
// retireReplicaCopies runs only after committing a replacement. Failures are
// returned to the caller, so a stale copy cannot be hidden behind a successful
// response. Data movement owns its source cleanup and does not use this helper.
func (z *erasureServerPools) retireReplicaCopies(ctx context.Context, bucket, object string, keep int, oi ObjectInfo) error {
versionID := oi.VersionID
if versionID == "" {
versionID = nullVersionID
}
copies, err := z.objectPoolInfos(ctx, bucket, object, ObjectOptions{VersionID: versionID})
if err != nil {
return err
}
var retained []PoolObjInfo
for _, copy := range copies {
if copy.Index == keep {
retained = []PoolObjInfo{copy}
break
}
}
if len(retained) == 0 {
return VersionNotFound{Bucket: bucket, Object: decodeDirObject(object), VersionID: versionID}
}
for i, copy := range copies {
if copy.Index == keep {
continue
}
_, err := z.serverPools[copy.Index].DeleteObject(ctx, bucket, object,
ObjectOptions{
VersionID: versionID, NoLock: true, NoAuditLog: true,
SkipFreeVersion: sharesTierObject(copy.ObjInfo, retained) || sharesTierObject(copy.ObjInfo, copies[i+1:]),
})
if err != nil && !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
return err
}
}
return nil
}
// deleteObjectReconciled evaluates any condition once against the logical
// version, then removes all its copies under the same lock as pooled writers.
func (z *erasureServerPools) deleteObjectReconciled(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
copies, err := z.objectPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
primary := copies[0]
if opts.CheckPrecondFn != nil && opts.CheckPrecondFn(primary.ObjInfo) {
return ObjectInfo{}, PreConditionFailed{}
}
opts.CheckPrecondFn = nil
opts.NoLock = true
if opts.EvalRetentionBypassFn != nil || opts.EvalMetadataFn != nil {
logical := primary.ObjInfo
var gerr error
switch {
case logical.DeleteMarker:
// Markers can be deleted by version ID. Match the set layer's
// callback inputs instead of rejecting them as metadata updates.
gerr = toObjectErr(errMethodNotAllowed, bucket, object)
if opts.VersionID == "" || opts.DeleteMarker {
gerr = toObjectErr(errFileNotFound, bucket, object)
}
case opts.VersionID != "":
// An addressed version already resolved every copy above.
logical = mergedPoolObjectInfo(copies)
default:
versions, err := z.metadataPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
logical = mergedPoolObjectInfo(versions)
}
// Keep the retention gate first. These callbacks independently evaluate
// the logical version and run once before any deletion; the handler only
// sweeps metadata's transition state after a successful delete.
if opts.EvalRetentionBypassFn != nil {
if err := opts.EvalRetentionBypassFn(logical, gerr); err != nil {
return ObjectInfo{}, err
}
opts.EvalRetentionBypassFn = nil
}
if opts.EvalMetadataFn != nil {
decision, err := opts.EvalMetadataFn(&logical, gerr)
if err != nil {
return ObjectInfo{}, err
}
if decision.ReplicateAny() {
opts.SetDeleteReplicationState(decision, opts.VersionID)
}
opts.EvalMetadataFn = nil
}
}
if opts.VersionID == "" && (opts.Versioned || opts.VersionSuspended) {
// A single new delete marker hides the current version. Older versions
// remain history and must not be removed by an unqualified DELETE.
oi, err := z.serverPools[primary.Index].DeleteObject(ctx, bucket, object, opts)
oi.Name = decodeDirObject(object)
oi.replicationDecision = opts.DeleteReplication.ReplicateDecisionStr
return oi, err
}
// Retire non-authoritative copies first. If cleanup fails, retain the
// authoritative version and report the error instead of acknowledging a
// deletion that would expose an older copy.
for i := 1; i < len(copies); i++ {
candidate := copies[i]
deleteOpts := opts
deleteOpts.SkipFreeVersion = opts.SkipFreeVersion || sharesTierObject(candidate.ObjInfo, copies[:1]) || sharesTierObject(candidate.ObjInfo, copies[i+1:])
_, err := z.serverPools[candidate.Index].DeleteObject(ctx, bucket, object, deleteOpts)
if err != nil && !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
return ObjectInfo{}, err
}
}
oi, err := z.serverPools[primary.Index].DeleteObject(ctx, bucket, object, opts)
oi.Name = decodeDirObject(object)
oi.replicationDecision = opts.DeleteReplication.ReplicateDecisionStr
return oi, err
}
File diff suppressed because it is too large Load Diff
+5 -1
View File
@@ -23,6 +23,10 @@ import (
)
func prepareErasurePools() (ObjectLayer, []string, error) {
return prepareErasurePoolsWithContext(context.Background())
}
func prepareErasurePoolsWithContext(ctx context.Context) (ObjectLayer, []string, error) {
nDisks := 32
fsDirs, err := getRandomDisks(nDisks)
if err != nil {
@@ -32,7 +36,7 @@ func prepareErasurePools() (ObjectLayer, []string, error) {
pools := mustGetPoolEndpoints(0, fsDirs[:16]...)
pools = append(pools, mustGetPoolEndpoints(1, fsDirs[16:]...)...)
objLayer, _, err := initObjectLayer(context.Background(), pools)
objLayer, _, err := initObjectLayer(ctx, pools)
if err != nil {
removeRoots(fsDirs)
return nil, nil, err
@@ -0,0 +1,172 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"testing"
"time"
)
// A multipart completion's If-Match must be evaluated against the logical
// latest object across pools, not against the copy local to a pool.
//
// Multi-pool write placement is not sticky (getPoolIdx picks by available
// space even for existing objects), so an upload and a newer overwrite of
// the same name routinely end up in different pools. Uploads are pinned to
// their pools directly: routing through z.NewMultipartUpload would make the
// placement depend on the space-weighted random choice.
func TestPoolsMultipartConditionalUsesLogicalLatest(t *testing.T) {
z, bucket := consistencyPools(t)
ctx := t.Context()
ifMatch := func(etag string) (opts ObjectOptions) {
return ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return oi.ETag != etag
},
}
}
uploadPart := func(t *testing.T, bucket, object, uploadID string) []CompletePart {
t.Helper()
pi, err := z.PutObjectPart(ctx, bucket, object, uploadID, 1,
mustGetPutObjReader(t, bytes.NewBufferString("part"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
return []CompletePart{{PartNumber: 1, ETag: pi.ETag}}
}
// Scenario A: the uploaded If-Match carries the stale ETag of the pool-0
// copy while the logical latest object lives in pool 1. The completion
// must fail with 412 instead of shadowing the newer logical state.
objectA := "cond-mp-stale-etag"
base := time.Now()
oldA := putConsistencyObject(t, z, bucket, objectA, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpA, err := z.serverPools[0].NewMultipartUpload(ctx, bucket, objectA, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerA := putConsistencyObject(t, z, bucket, objectA, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
latest, _, err := z.getLatestObjectInfoWithIdx(ctx, bucket, objectA, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if latest.ETag != newerA.ETag {
t.Fatalf("logical latest should be the pool-1 copy: got %s want %s", latest.ETag, newerA.ETag)
}
if _, err = z.CompleteMultipartUpload(ctx, bucket, objectA, mpA.UploadID,
uploadPart(t, bucket, objectA, mpA.UploadID), ifMatch(oldA.ETag)); err == nil {
t.Fatal("If-Match with the stale pool-0 ETag must not complete over the newer pool-1 object")
} else if _, ok := err.(PreConditionFailed); !ok {
t.Fatalf("expected PreconditionFailed, got %v", err)
}
if latest, _, err = z.getLatestObjectInfoWithIdx(ctx, bucket, objectA, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if latest.ETag != newerA.ETag {
t.Fatalf("the newer pool-1 object must remain the logical latest, got %s", latest.ETag)
}
// Scenario B: the uploaded If-Match carries the logical latest ETag (the
// pool-1 copy) while the upload sits next to the stale pool-0 copy. The
// precondition is satisfied, so completion must succeed and its result
// must become the logical latest.
objectB := "cond-mp-latest-etag"
putConsistencyObject(t, z, bucket, objectB, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpB, err := z.serverPools[0].NewMultipartUpload(ctx, bucket, objectB, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerB := putConsistencyObject(t, z, bucket, objectB, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
oiB, err := z.CompleteMultipartUpload(ctx, bucket, objectB, mpB.UploadID,
uploadPart(t, bucket, objectB, mpB.UploadID), ifMatch(newerB.ETag))
if err != nil {
t.Fatalf("If-Match with the logical latest ETag must complete, got %v", err)
}
if oiB.ETag == "" {
t.Fatal("completion returned an empty ETag")
}
if latest, _, err = z.getLatestObjectInfoWithIdx(ctx, bucket, objectB, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if latest.ETag != oiB.ETag {
t.Fatalf("the completed object must be the logical latest: got %s want %s", latest.ETag, oiB.ETag)
}
// Scenario C: the upload lives in pool 1 with the newer copy while pool 0
// holds the stale one. A set-local evaluation order would let pool 0's
// stale copy fail the request before pool 1 is reached; the logical
// latest ETag must complete.
objectC := "cond-mp-upload-other-pool"
putConsistencyObject(t, z, bucket, objectC, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpC, err := z.serverPools[1].NewMultipartUpload(ctx, bucket, objectC, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerC := putConsistencyObject(t, z, bucket, objectC, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
if _, err = z.CompleteMultipartUpload(ctx, bucket, objectC, mpC.UploadID,
uploadPart(t, bucket, objectC, mpC.UploadID), ifMatch(newerC.ETag)); err != nil {
t.Fatalf("If-Match with the logical latest ETag must complete regardless of upload pool, got %v", err)
}
// Scenario D: pool 1 holds the newer copy but cannot be read. An
// unreadable pool may contain the newest state, so the unverifiable
// condition must fail the request rather than pass it against pool 0's
// stale ETag.
objectD := "cond-mp-unreadable-pool"
oldD := putConsistencyObject(t, z, bucket, objectD, 0, "old", ObjectOptions{MTime: base.Add(-2 * time.Minute)})
mpD, err := z.serverPools[0].NewMultipartUpload(ctx, bucket, objectD, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
newerD := putConsistencyObject(t, z, bucket, objectD, 1, "new", ObjectOptions{
MTime: base.Add(-time.Minute),
})
if latest, _, err = z.getLatestObjectInfoWithIdx(ctx, bucket, objectD, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if latest.ETag != newerD.ETag {
t.Fatalf("logical latest before faulting pool 1 should be its copy: got %s want %s", latest.ETag, newerD.ETag)
}
set := z.serverPools[1].getHashedSet(objectD)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = consistencyReadFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: objectD}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
_, err = z.CompleteMultipartUpload(ctx, bucket, objectD, mpD.UploadID,
uploadPart(t, bucket, objectD, mpD.UploadID), ifMatch(oldD.ETag))
if !isErrReadQuorum(err) {
t.Fatalf("expected an insufficient read quorum error, got %v", err)
}
}
@@ -0,0 +1,341 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"encoding/xml"
"errors"
"fmt"
"io"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
)
func multipartConditionRequest(t *testing.T, router http.Handler, method, target, body string, headers map[string]string) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, target, int64(len(body)), strings.NewReader(body), globalActiveCred.AccessKey, globalActiveCred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
return rec
}
// The server writes the wire spelling ETag directly into Header; unlike a
// network response, httptest's map has not canonicalized it to Etag.
func multipartConditionResponseETag(rec *httptest.ResponseRecorder) string {
for key, values := range rec.Header() {
if strings.EqualFold(key, xhttp.ETag) && len(values) > 0 {
return strings.Trim(values[0], "\"")
}
}
return ""
}
func multipartConditionUpload(t *testing.T, z *erasureServerPools, bucket, object string, owner int, opts ObjectOptions) (string, []CompletePart) {
t.Helper()
mp, err := z.serverPools[owner].NewMultipartUpload(t.Context(), bucket, object, opts)
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[owner].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1, mustGetPutObjReader(t, bytes.NewBufferString("replacement"), 11, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
return mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}
}
func multipartConditionCompleteBody(parts []CompletePart) string {
data, err := xml.Marshal(CompleteMultipartUpload{Parts: parts})
if err != nil {
panic(err)
}
return string(data)
}
func multipartConditionError(t *testing.T, rec *httptest.ResponseRecorder, code string) {
t.Helper()
decoder := xml.NewDecoder(strings.NewReader(rec.Body.String()))
var response APIErrorResponse
if err := decoder.Decode(&response); err != nil || response.Code != code {
t.Fatalf("expected %s error, got %q: %v", code, rec.Body.String(), err)
}
if err := decoder.Decode(&response); err != io.EOF {
t.Fatalf("expected exactly one error response, got %q: %v", rec.Body.String(), err)
}
}
// State is deliberately placed per pool; the final request uses signed HTTP
// and the real handler, precondition callback, erasure metadata and rename.
func TestPoolsMultipartConditionalHTTPMatrix(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
for owner := range 2 {
for _, localOld := range []bool{false, true} {
for _, condition := range []string{"match-old", "match-current", "none-match"} {
t.Run(fmt.Sprintf("owner=%d/old=%t/%s", owner, localOld, condition), func(t *testing.T) {
object := fmt.Sprintf("http-%d-%t-%s", owner, localOld, condition)
oldETag := "old-does-not-exist"
if localOld {
oldETag = putConsistencyObject(t, z, bucket, object, owner, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)}).ETag
}
id, parts := multipartConditionUpload(t, z, bucket, object, owner, ObjectOptions{})
current := putConsistencyObject(t, z, bucket, object, 1-owner, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
head := multipartConditionRequest(t, router, http.MethodHead, getGetObjectURL("", bucket, object), "", nil)
if head.Code != 200 || multipartConditionResponseETag(head) != current.ETag {
t.Fatalf("bad HEAD: %d %v", head.Code, head.Header())
}
h := map[string]string{xhttp.IfMatch: "\"" + oldETag + "\""}
want := 412
if condition == "match-current" {
h[xhttp.IfMatch] = "\"" + current.ETag + "\""
want = 200
}
if condition == "none-match" {
h = map[string]string{xhttp.IfNoneMatch: "*"}
}
rec := multipartConditionRequest(t, router, http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, id), multipartConditionCompleteBody(parts), h)
t.Logf("HEAD current=%s, requested=%v, complete HTTP=%d", current.ETag, h, rec.Code)
if rec.Code != want {
t.Errorf("want HTTP %d, got %d: %s", want, rec.Code, rec.Body.String())
}
if want == 412 {
multipartConditionError(t, rec, "PreconditionFailed")
got := multipartConditionRequest(t, router, http.MethodGet, getGetObjectURL("", bucket, object), "", nil)
if got.Code != 200 || got.Body.String() != "current" {
t.Errorf("rejected request must preserve current data: %d %q", got.Code, got.Body.String())
}
if _, err := z.serverPools[owner].ListObjectParts(t.Context(), bucket, object, id, 0, 10, ObjectOptions{}); err != nil {
t.Errorf("rejected request consumed upload: %v", err)
}
}
})
}
}
}
}
func TestPoolsMultipartConditionalHTTPAbsentObject(t *testing.T) {
for _, deleted := range []bool{false, true} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("delete-marker=%t/if-match=%t", deleted, match), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{VersioningEnabled: deleted})
if err != nil {
t.Fatal(err)
}
object := "http-absent"
if deleted {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
_, err = z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)})
if err != nil {
t.Fatal(err)
}
}
id, parts := multipartConditionUpload(t, z, bucket, object, 0, ObjectOptions{Versioned: deleted})
headers := map[string]string{xhttp.IfNoneMatch: "*"}
want := http.StatusOK
if match {
headers = map[string]string{xhttp.IfMatch: "\"missing\""}
want = http.StatusNotFound
}
rec := multipartConditionRequest(t, router, http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, id), multipartConditionCompleteBody(parts), headers)
if rec.Code != want {
t.Fatalf("expected HTTP %d, got %d: %s", want, rec.Code, rec.Body.String())
}
if match {
multipartConditionError(t, rec, "NoSuchKey")
if _, err := z.serverPools[0].ListObjectParts(t.Context(), bucket, object, id, 0, 10, ObjectOptions{}); err != nil {
t.Errorf("rejected request consumed upload: %v", err)
}
}
})
}
}
}
// All writes below use ordinary signed S3 requests. The result must be correct
// for every placement; the matrix above deterministically covers split pools.
func TestPoolsMultipartConditionalHTTPNormalRouting(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
object := "normal-routing"
url := getPutObjectURL("", bucket, object)
old := multipartConditionRequest(t, router, http.MethodPut, url, "old-data", nil)
if old.Code != http.StatusOK {
t.Fatalf("initial PUT %d: %s", old.Code, old.Body.String())
}
init := multipartConditionRequest(t, router, http.MethodPost, url+"?uploads", "", nil)
if init.Code != http.StatusOK {
t.Fatalf("init %d: %s", init.Code, init.Body.String())
}
var mp InitiateMultipartUploadResponse
if err := xml.Unmarshal(init.Body.Bytes(), &mp); err != nil {
t.Fatal(err)
}
part := multipartConditionRequest(t, router, http.MethodPut, getPutObjectPartURL("", bucket, object, mp.UploadID, "1"), "replacement", nil)
if part.Code != http.StatusOK {
t.Fatalf("part %d: %s", part.Code, part.Body.String())
}
parts := []CompletePart{{PartNumber: 1, ETag: multipartConditionResponseETag(part)}}
newer := multipartConditionRequest(t, router, http.MethodPut, url, "newer-data", nil)
if newer.Code != http.StatusOK {
t.Fatalf("new PUT %d: %s", newer.Code, newer.Body.String())
}
head := multipartConditionRequest(t, router, http.MethodHead, url, "", nil)
if head.Code != http.StatusOK || multipartConditionResponseETag(head) != multipartConditionResponseETag(newer) {
t.Fatalf("HEAD did not pick new object: %d %v", head.Code, head.Header())
}
rec := multipartConditionRequest(t, router, http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, mp.UploadID), multipartConditionCompleteBody(parts), map[string]string{xhttp.IfMatch: "\"" + multipartConditionResponseETag(old) + "\""})
if rec.Code != http.StatusPreconditionFailed {
t.Errorf("stale If-Match should be HTTP 412, got %d: %s", rec.Code, rec.Body.String())
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "newer-data" {
t.Errorf("conditional completion changed newer data: %d %q", get.Code, get.Body.String())
}
}
func TestPoolsMultipartConditionalUnreadablePool(t *testing.T) {
z, bucket := consistencyPools(t)
object := "quorum-with-readable-copy"
old := putConsistencyObject(t, z, bucket, object, 0, "readable", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "hidden-newer", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
id, parts := multipartConditionUpload(t, z, bucket, object, 0, ObjectOptions{})
set := z.serverPools[1].getHashedSet(object)
original := set.getDisks
disks := append([]StorageAPI(nil), original()...)
for i := range disks {
disks[i] = consistencyReadFaultDisk{StorageAPI: disks[i], bucket: bucket, object: object}
}
set.getDisks = func() []StorageAPI { return disks }
defer func() { set.getDisks = original }()
read, readErr := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
t.Logf("ordinary GET lookup: etag=%s err=%v", read.ETag, readErr)
called := 0
_, err := z.CompleteMultipartUpload(t.Context(), bucket, object, id, parts, ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { called++; return oi.ETag != old.ETag }})
if !isErrReadQuorum(err) {
t.Errorf("conditional write must fail on unreadable pool, got %v", err)
}
if called != 0 {
t.Errorf("callback evaluated without complete state: %d calls", called)
}
}
func TestPoolsMultipartConditionalLatestVersionAndCallbackOnce(t *testing.T) {
for _, explicit := range []bool{false, true} {
t.Run(fmt.Sprintf("explicit=%t", explicit), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "tie-version"
opts := ObjectOptions{MTime: UTCNow().Add(-time.Hour), Versioned: explicit}
if explicit {
opts.VersionID = mustGetUUID()
}
current := putConsistencyObject(t, z, bucket, object, 0, "first", opts)
putConsistencyObject(t, z, bucket, object, 1, "second", opts)
if explicit {
current = putConsistencyObject(t, z, bucket, object, 1, "latest-other-version", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
}
id, parts := multipartConditionUpload(t, z, bucket, object, 1, opts)
called := 0
opts.CheckPrecondFn = func(oi ObjectInfo) bool { called++; return oi.ETag != current.ETag }
opts.MTime = time.Time{}
opts.HasIfMatch = true
_, err := z.CompleteMultipartUpload(t.Context(), bucket, object, id, parts, opts)
if err != nil {
t.Errorf("logical current object should match: %v", err)
}
if called != 1 {
t.Errorf("condition evaluated %d times; want exactly once", called)
}
})
}
}
func TestPoolsMultipartConditionalConcurrentCompletes(t *testing.T) {
z, bucket := consistencyPools(t)
object := "concurrent-completes"
old := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "old", ObjectOptions{MTime: old.ModTime})
ids := make([]string, 2)
parts := make([][]CompletePart, 2)
for i := range 2 {
ids[i], parts[i] = multipartConditionUpload(t, z, bucket, object, i, ObjectOptions{})
}
entered := make(chan struct{})
release := make(chan struct{})
var gate, releaseOnce sync.Once
defer releaseOnce.Do(func() { close(release) })
errs := make([]error, 2)
var wg sync.WaitGroup
// Let pool 1's completion hold the object lock before pool 0's starts.
// Once pool 1 commits, a set-local read in pool 0 would still see the
// old ETag and incorrectly accept the second completion.
wg.Add(1)
go func() {
defer wg.Done()
_, errs[1] = z.CompleteMultipartUpload(t.Context(), bucket, object, ids[1], parts[1], ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool {
gate.Do(func() { close(entered); <-release })
return oi.ETag != old.ETag
}})
}()
select {
case <-entered:
case <-time.After(10 * time.Second):
t.Fatal("first completion did not enter its condition callback")
}
started := make(chan struct{})
wg.Add(1)
go func() {
defer wg.Done()
close(started)
_, errs[0] = z.CompleteMultipartUpload(t.Context(), bucket, object, ids[0], parts[0], ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != old.ETag }})
}()
<-started
releaseOnce.Do(func() { close(release) })
wg.Wait()
success, failed := 0, 0
for _, err := range errs {
var p PreConditionFailed
switch {
case err == nil:
success++
case errors.As(err, &p):
failed++
default:
t.Errorf("unexpected completion error %v", err)
}
}
if success != 1 || failed != 1 {
t.Errorf("CAS writers: success=%d conditional failures=%d errors=%v", success, failed, errs)
}
}
@@ -0,0 +1,135 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"errors"
"fmt"
"testing"
"time"
)
func TestPoolsMultipartConditionMatrix(t *testing.T) {
for owner := range 2 {
for _, withOld := range []bool{false, true} {
for _, condition := range []string{"match-current", "match-old", "none-match-any"} {
t.Run(fmt.Sprintf("upload-pool=%d/old-copy=%t/%s", owner, withOld, condition), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-multipart"
oldETag := "arbitrary-old"
if withOld {
old := putConsistencyObject(t, z, bucket, object, owner, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
oldETag = old.ETag
}
mp, err := z.serverPools[owner].NewMultipartUpload(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[owner].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1, mustGetPutObjReader(t, bytes.NewBufferString("replacement"), 11, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
current := putConsistencyObject(t, z, bucket, object, 1-owner, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
visible, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil || visible.ETag != current.ETag {
t.Fatalf("invalid current state: %v", err)
}
opts := ObjectOptions{HasIfMatch: condition != "none-match-any", CheckPrecondFn: func(oi ObjectInfo) bool {
switch condition {
case "match-current":
return oi.ETag != current.ETag
case "match-old":
return oi.ETag != oldETag
default:
return true
}
}}
_, err = z.CompleteMultipartUpload(t.Context(), bucket, object, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, opts)
if condition == "match-current" {
if err != nil {
t.Errorf("correct logical If-Match rejected: %v", err)
}
} else {
var expected PreConditionFailed
if !errors.As(err, &expected) {
t.Errorf("logical condition must fail, got %v", err)
}
}
})
}
}
}
}
func TestPoolsMultipartConditionBoundaries(t *testing.T) {
for _, state := range []string{"missing", "latest-delete-marker", "unreadable-other-pool"} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("%s/if-match=%t", state, match), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-boundary"
versioned := state == "latest-delete-marker"
if versioned {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
if _, err := z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)}); err != nil {
t.Fatal(err)
}
}
mp, err := z.serverPools[0].NewMultipartUpload(t.Context(), bucket, object, ObjectOptions{Versioned: versioned})
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[0].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1, mustGetPutObjReader(t, bytes.NewBufferString("new"), 3, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if state == "unreadable-other-pool" {
set := z.serverPools[1].getHashedSet(object)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = consistencyReadFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: object}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
}
opts := ObjectOptions{Versioned: versioned, HasIfMatch: match, CheckPrecondFn: func(ObjectInfo) bool { return true }}
_, err = z.CompleteMultipartUpload(t.Context(), bucket, object, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, opts)
switch {
case state == "unreadable-other-pool":
if !isErrReadQuorum(err) {
t.Errorf("unreadable pool must not mean absence: %v", err)
}
case match:
if !isErrObjectNotFound(err) {
t.Errorf("If-Match against logical absence should report absence: %v", err)
}
default:
if err != nil {
t.Errorf("If-None-Match against logical absence must succeed: %v", err)
}
}
if err != nil {
if _, lerr := z.serverPools[0].ListObjectParts(t.Context(), bucket, object, mp.UploadID, 0, 10, ObjectOptions{}); lerr != nil {
t.Errorf("failed condition consumed upload: %v", lerr)
}
}
})
}
}
}
@@ -0,0 +1,721 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"crypto/md5"
"encoding/base64"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// Change allocation capacity only; metadata and data use real fixture disks.
// Restoring the adapters also lets encrypted fixtures move the write target
// between requests without changing the production allocation policy.
type conditionalPutCapacityDisk struct {
StorageAPI
full bool
}
func (d conditionalPutCapacityDisk) DiskInfo(ctx context.Context, opts DiskInfoOptions) (DiskInfo, error) {
info, err := d.StorageAPI.DiskInfo(ctx, opts)
info.Total, info.Used = 1<<40, 0
if d.full {
info.Used = info.Total - (1 << 20)
}
info.Free = info.Total - info.Used
return info, err
}
// Keep getDisks immutable while background IAM and storage readers use it.
// GetDisks takes this same mutex when it copies the backing disk list.
func conditionalPutSwapDisks(pool *erasureSets, object string, wrap func(StorageAPI) StorageAPI) func() {
setIndex := pool.getHashedSet(object).setIndex
pool.erasureDisksMu.Lock()
previous := pool.erasureDisks[setIndex]
disks := append([]StorageAPI(nil), previous...)
for i, disk := range disks {
disks[i] = wrap(disk)
}
pool.erasureDisks[setIndex] = disks
pool.erasureDisksMu.Unlock()
return func() {
pool.erasureDisksMu.Lock()
pool.erasureDisks[setIndex] = previous
pool.erasureDisksMu.Unlock()
}
}
func conditionalPutPool(t *testing.T, z *erasureServerPools, object string, target int) func() {
t.Helper()
var restore []func()
for i, pool := range z.serverPools {
restore = append(restore, conditionalPutSwapDisks(pool, object, func(disk StorageAPI) StorageAPI {
return conditionalPutCapacityDisk{StorageAPI: disk, full: i != target}
}))
}
return func() {
for _, fn := range restore {
fn()
}
}
}
func conditionalPutBucket(t *testing.T, z *erasureServerPools, mode string) (string, http.Handler) {
t.Helper()
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{VersioningEnabled: mode != "unversioned"})
if err != nil {
t.Fatal(err)
}
if mode == "suspended" {
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig,
[]byte(`<VersioningConfiguration><Status>Suspended</Status></VersioningConfiguration>`)); err != nil {
t.Fatal(err)
}
}
return bucket, router
}
func TestPoolsConditionalPutHTTP(t *testing.T) {
for _, mode := range []string{"unversioned", "versioned", "suspended"} {
t.Run(mode, func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, mode)
for target := range 2 {
for _, tc := range []struct {
name string
oldCopy, localCurrent bool
condition string
status int
}{
{"stale-etag-accepted", true, false, "old", http.StatusPreconditionFailed},
{"current-etag-rejected", true, false, "current", http.StatusOK},
{"create-only-overwrites-other-pool", false, false, "none", http.StatusPreconditionFailed},
{"current-etag-missing-in-write-pool", false, false, "current", http.StatusOK},
{"control-current-in-write-pool", true, true, "current", http.StatusOK},
{"control-stale-etag-rejected", true, true, "old", http.StatusPreconditionFailed},
{"none-match-current-etag", true, false, "none-current", http.StatusPreconditionFailed},
{"none-match-old-etag", true, false, "none-old", http.StatusOK},
} {
t.Run(fmt.Sprintf("target=%d/%s", target, tc.name), func(t *testing.T) {
object := fmt.Sprintf("%d-%s", target, tc.name)
currentPool := 1 - target
if tc.localCurrent {
currentPool = target
}
opts := ObjectOptions{Versioned: mode == "versioned", VersionSuspended: mode == "suspended", MTime: UTCNow().Add(-time.Hour)}
oldETag := "absent-old"
if tc.oldCopy {
oldETag = putConsistencyObject(t, z, bucket, object, 1-currentPool, "old", opts).ETag
}
opts.MTime = UTCNow().Add(-time.Minute)
current := putConsistencyObject(t, z, bucket, object, currentPool, "current", opts)
defer conditionalPutPool(t, z, object, target)()
idx, err := z.getWritePoolIdx(t.Context(), bucket, object, 11, false)
if err != nil || idx != target {
t.Fatalf("allocation target=%d: idx=%d err=%v", target, idx, err)
}
headers := map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", current.ETag)}
if tc.condition == "old" {
headers[xhttp.IfMatch] = fmt.Sprintf("%q", oldETag)
}
if tc.condition == "none" {
headers = map[string]string{xhttp.IfNoneMatch: "*"}
}
if tc.condition == "none-current" {
headers = map[string]string{xhttp.IfNoneMatch: fmt.Sprintf("%q", current.ETag)}
}
if tc.condition == "none-old" {
headers = map[string]string{xhttp.IfNoneMatch: fmt.Sprintf("%q", oldETag)}
}
url := getPutObjectURL("", bucket, object)
before := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if before.Code != http.StatusOK || before.Body.String() != "current" || multipartConditionResponseETag(before) != current.ETag {
t.Fatalf("invalid current object: %d %q %v", before.Code, before.Body.String(), before.Header())
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != tc.status {
t.Errorf("PUT status=%d want=%d: %s", put.Code, tc.status, put.Body.String())
}
if put.Code == http.StatusPreconditionFailed {
multipartConditionError(t, put, "PreconditionFailed")
if multipartConditionResponseETag(put) != current.ETag || put.Header().Get(xhttp.LastModified) != before.Header().Get(xhttp.LastModified) {
t.Errorf("412 headers do not describe current object: %v", put.Header())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
wantBody, wantETag := "current", current.ETag
if tc.status == http.StatusOK {
wantBody, wantETag = "replacement", fmt.Sprintf("%x", md5.Sum([]byte("replacement")))
}
t.Logf("PUT %d; GET %d bytes=%q ETag=%s", put.Code, get.Code, get.Body.String(), multipartConditionResponseETag(get))
if get.Code != http.StatusOK || get.Body.String() != wantBody || multipartConditionResponseETag(get) != wantETag {
t.Errorf("GET=%d bytes=%q ETag=%s; want %q %s", get.Code, get.Body.String(), multipartConditionResponseETag(get), wantBody, wantETag)
}
})
}
}
})
}
}
func TestPoolsConditionalPutHTTPAbsence(t *testing.T) {
for _, state := range []string{"missing", "uuid-marker", "null-marker"} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("%s/match=%t", state, match), func(t *testing.T) {
z, _ := consistencyPools(t)
mode := "versioned"
if state == "null-marker" {
mode = "suspended"
}
bucket, router := conditionalPutBucket(t, z, mode)
object := "absent-key"
if state != "missing" {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
vid := mustGetUUID()
if state == "null-marker" {
vid = nullVersionID
}
if _, err := z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: vid, DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)}); err != nil {
t.Fatal(err)
}
}
defer conditionalPutPool(t, z, object, 0)()
headers := map[string]string{xhttp.IfNoneMatch: "*"}
want := http.StatusOK
if match {
headers = map[string]string{xhttp.IfMatch: "*"}
want = http.StatusNotFound
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url, "new", headers)
if put.Code != want {
t.Fatalf("PUT %d want %d: %s", put.Code, want, put.Body.String())
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if match {
multipartConditionError(t, put, "NoSuchKey")
if get.Code != http.StatusNotFound {
t.Fatalf("failed PUT exposed data: %d %s", get.Code, get.Body.String())
}
} else if get.Code != http.StatusOK || get.Body.String() != "new" || multipartConditionResponseETag(get) != multipartConditionResponseETag(put) {
t.Fatalf("successful create GET: %d %q %v", get.Code, get.Body.String(), get.Header())
}
})
}
}
}
func TestPoolsConditionalPutUnreadable(t *testing.T) {
for faultPool := range 2 {
for _, present := range []bool{false, true} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("fault-pool=%d/present=%t/match=%t", faultPool, present, match), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
object := "unreadable-key"
if present {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
}
defer conditionalPutPool(t, z, object, 0)()
restoreFault := conditionalPutSwapDisks(z.serverPools[faultPool], object, func(disk StorageAPI) StorageAPI {
return consistencyReadFaultDisk{StorageAPI: disk, bucket: bucket, object: object}
})
defer restoreFault()
headers := map[string]string{xhttp.IfNoneMatch: "*"}
if match {
headers = map[string]string{xhttp.IfMatch: "*"}
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != http.StatusServiceUnavailable {
t.Errorf("unverified PUT must fail: %d %s", put.Code, put.Body.String())
}
called := 0
_, err := z.PutObject(t.Context(), bucket, object, mustGetPutObjReader(t, strings.NewReader("replacement"), 11, "", ""), ObjectOptions{HasIfMatch: match, CheckPrecondFn: func(ObjectInfo) bool { called++; return false }})
if !isErrReadQuorum(err) || called != 0 {
t.Errorf("lookup error=%v callback calls=%d", err, called)
}
restoreFault()
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if present {
if get.Code != http.StatusOK || get.Body.String() != "current" || multipartConditionResponseETag(get) != fmt.Sprintf("%x", md5.Sum([]byte("current"))) {
t.Fatalf("failed PUT changed object: %d %q", get.Code, get.Body.String())
}
} else if get.Code != http.StatusNotFound {
t.Fatalf("failed PUT created object: %d %q", get.Code, get.Body.String())
}
})
}
}
}
}
func TestPoolsConditionalPutEncryptedETag(t *testing.T) {
for _, kind := range []string{"SSE-C", "SSE-S3", "SSE-KMS"} {
t.Run(kind, func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
oldKMS, oldTLS := GlobalKMS, globalIsTLS
GlobalKMS, globalIsTLS = kms.NewStub("conditional-put-key"), true
defer func() { GlobalKMS, globalIsTLS = oldKMS, oldTLS }()
headers := map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}
readHeaders := map[string]string{}
if kind == "SSE-C" {
key := bytes.Repeat([]byte{0x42}, 32)
digest := md5.Sum(key)
headers = map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(digest[:]),
}
readHeaders = maps.Clone(headers)
}
if kind == "SSE-KMS" {
headers[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionKMS
headers[xhttp.AmzServerSideEncryptionKmsID] = "conditional-put-key"
}
object := "encrypted-key"
url := getPutObjectURL("", bucket, object)
restore := conditionalPutPool(t, z, object, 0)
old := multipartConditionRequest(t, router, http.MethodPut, url, "old", headers)
restore()
restore = conditionalPutPool(t, z, object, 1)
current := multipartConditionRequest(t, router, http.MethodPut, url, "current", headers)
restore()
if old.Code != http.StatusOK || current.Code != http.StatusOK {
t.Fatalf("encrypted setup: %d %s / %d %s", old.Code, old.Body.String(), current.Code, current.Body.String())
}
defer conditionalPutPool(t, z, object, 0)()
for _, stale := range []bool{true, false} {
h := maps.Clone(headers)
h[xhttp.IfMatch] = fmt.Sprintf("%q", multipartConditionResponseETag(current))
wantStatus, wantBody, wantETag := http.StatusOK, "replacement", ""
if stale {
h[xhttp.IfMatch] = fmt.Sprintf("%q", multipartConditionResponseETag(old))
wantStatus, wantBody, wantETag = http.StatusPreconditionFailed, "current", multipartConditionResponseETag(current)
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", h)
if put.Code != wantStatus {
t.Fatalf("encrypted condition: %d want %d: %s", put.Code, wantStatus, put.Body.String())
}
if !stale {
wantETag = multipartConditionResponseETag(put)
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", readHeaders)
if get.Code != http.StatusOK || get.Body.String() != wantBody || multipartConditionResponseETag(get) != wantETag {
t.Fatalf("encrypted GET: %d %q %v", get.Code, get.Body.String(), get.Header())
}
}
})
}
}
func TestPoolsConditionalPutVersionSelection(t *testing.T) {
for _, kind := range []string{"public-version", "replica", "replica-preserve-etag", "movement", "no-lock", "tie", "draining"} {
t.Run(kind, func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "version-selection"
addressed := putConsistencyObject(t, z, bucket, object, 1, "addressed", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
opts := ObjectOptions{Versioned: true, VersionID: addressed.VersionID, HasIfMatch: true}
want := current
switch kind {
case "replica":
opts.ReplicaLockReconcile, opts.ReplicationRequest = true, true
want = addressed
case "replica-preserve-etag":
opts.PreserveETag = addressed.ETag
opts.ReplicaLockReconcile, opts.ReplicationRequest = true, true
want = addressed
case "movement":
opts.DataMovement, opts.SrcPoolIdx = true, 1
want = addressed
case "tie":
want = putConsistencyObject(t, z, bucket, object, 0, "tie-winner", ObjectOptions{Versioned: true, MTime: current.ModTime})
case "draining":
z.poolMetaMutex.Lock()
z.poolMeta.Pools[1].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
}
defer conditionalPutPool(t, z, object, 0)()
ctx := t.Context()
if kind == "no-lock" {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
t.Fatal(err)
}
defer lk.Unlock(lkctx)
ctx, opts.NoLock = lkctx.Context(), true
}
called := 0
opts.UserDefined = make(map[string]string)
opts.CheckPrecondFn = func(oi ObjectInfo) bool {
called++
if oi.ETag != want.ETag || oi.VersionID != want.VersionID {
t.Errorf("comparison ETag/version=%s/%s want %s/%s", oi.ETag, oi.VersionID, want.ETag, want.VersionID)
}
return oi.ETag != want.ETag
}
oi, err := z.PutObject(ctx, bucket, object, mustGetPutObjReader(t, strings.NewReader("replacement"), 11, "", ""), opts)
if err != nil || called != 1 {
t.Fatalf("PUT err=%v callback calls=%d", err, called)
}
if oi.VersionID != addressed.VersionID {
t.Fatalf("destination version changed: %s", oi.VersionID)
}
if opts.PreserveETag != "" && oi.ETag != opts.PreserveETag {
t.Fatalf("PreserveETag changed: %s", oi.ETag)
}
})
}
}
func TestPoolsConditionalPutReplicaDuplicateHTTP(t *testing.T) {
for _, null := range []bool{false, true} {
t.Run(fmt.Sprintf("null=%t", null), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "replica-duplicate"
vid := mustGetUUID()
if null {
vid = nullVersionID
}
addressed := putConsistencyObject(t, z, bucket, object, 1, "addressed", ObjectOptions{Versioned: true, VersionID: vid, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
xhttp.MinIOSourceETag: addressed.ETag,
xhttp.MinIOSourceMTime: addressed.ModTime.Format(time.RFC3339Nano),
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url+"?versionId="+vid, "addressed", headers)
if put.Code != http.StatusPreconditionFailed {
t.Fatalf("replica duplicate: %d %s", put.Code, put.Body.String())
}
multipartConditionError(t, put, "PreconditionFailed")
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "current" || multipartConditionResponseETag(get) != current.ETag {
t.Fatalf("duplicate changed current object: %d %q", get.Code, get.Body.String())
}
get = multipartConditionRequest(t, router, http.MethodGet, url+"?versionId="+vid, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "addressed" || multipartConditionResponseETag(get) != addressed.ETag {
t.Fatalf("duplicate changed addressed version: %d %q", get.Code, get.Body.String())
}
})
}
}
func TestPoolsConditionalPutConcurrentHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
for _, match := range []bool{false, true} {
for iteration := range 5 {
t.Run(fmt.Sprintf("match=%t/iteration=%d", match, iteration), func(t *testing.T) {
object := fmt.Sprintf("concurrent-%t-%d", match, iteration)
headers := map[string]string{xhttp.IfNoneMatch: "*"}
if match {
oi := putConsistencyObject(t, z, bucket, object, 1, "old", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
headers = map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", oi.ETag)}
}
defer conditionalPutPool(t, z, object, 0)()
start := make(chan struct{})
results := make(chan *httptest.ResponseRecorder, 2)
url := getPutObjectURL("", bucket, object)
for i := range 2 {
body := fmt.Sprintf("writer-%d", i)
req, err := newTestSignedRequestV4(http.MethodPut, url, int64(len(body)), strings.NewReader(body), globalActiveCred.AccessKey, globalActiveCred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
go func() {
<-start
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
results <- rec
}()
}
close(start)
success, failed, winnerETag := 0, 0, ""
for range 2 {
result := <-results
switch result.Code {
case http.StatusOK:
success++
winnerETag = multipartConditionResponseETag(result)
case http.StatusPreconditionFailed:
failed++
multipartConditionError(t, result, "PreconditionFailed")
default:
t.Errorf("unexpected PUT %d: %s", result.Code, result.Body.String())
}
}
if success != 1 || failed != 1 {
t.Fatalf("success=%d precondition failures=%d", success, failed)
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || multipartConditionResponseETag(get) != winnerETag || fmt.Sprintf("%x", md5.Sum(get.Body.Bytes())) != winnerETag {
t.Fatalf("winner lost: %d %q %v", get.Code, get.Body.String(), get.Header())
}
})
}
}
}
func TestPoolsConditionalPutSerializesMutation(t *testing.T) {
for _, deletion := range []bool{false, true} {
t.Run(fmt.Sprintf("delete=%t", deletion), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-mutation"
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
gate := &consistencyGateReader{Reader: strings.NewReader("replacement"), entered: make(chan struct{}), resume: make(chan struct{})}
release := func() { gate.release.Do(func() { close(gate.resume) }) }
defer release()
reader := mustGetPutObjReader(t, gate, 11, "", "")
written := make(chan error, 1)
go func() {
_, err := z.PutObject(ctx, bucket, object, reader, ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != current.ETag }})
written <- err
}()
select {
case <-gate.entered:
case err := <-written:
t.Fatalf("PUT failed before body read: %v", err)
case <-ctx.Done():
t.Fatal(ctx.Err())
}
mutated := make(chan error, 1)
go func() {
var err error
if deletion {
_, err = z.DeleteObject(ctx, bucket, object, ObjectOptions{})
} else {
_, err = z.PutObjectMetadata(ctx, bucket, object, ObjectOptions{EvalMetadataFn: func(oi *ObjectInfo, _ error) (ReplicateDecision, error) {
oi.UserDefined["x-amz-meta-after-put"] = "present"
return ReplicateDecision{}, nil
}})
}
mutated <- err
}()
select {
case err := <-mutated:
t.Fatalf("mutation escaped PUT lock: %v", err)
case <-time.After(100 * time.Millisecond):
}
release()
if err := <-written; err != nil {
t.Fatal(err)
}
if err := <-mutated; err != nil {
t.Fatal(err)
}
oi, err := z.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if deletion {
if !isErrObjectNotFound(err) {
t.Fatalf("delete lost: %v", err)
}
} else if err != nil || oi.UserDefined["x-amz-meta-after-put"] != "present" || oi.ETag != fmt.Sprintf("%x", md5.Sum([]byte("replacement"))) {
t.Fatalf("metadata/PUT lost: %+v %v", oi, err)
}
})
}
}
func TestSinglePoolConditionalPutHTTP(t *testing.T) {
ctx, cancel := context.WithCancel(t.Context())
obj, dirs, err := prepareErasure16(ctx)
if err != nil {
cancel()
t.Fatal(err)
}
z := obj.(*erasureServerPools)
t.Cleanup(func() { cancel(); z.Shutdown(context.Background()); removeRoots(dirs) })
if !z.SinglePool() {
t.Fatal("fixture is not a single pool")
}
bucket, router := conditionalPutBucket(t, z, "unversioned")
object := "single-pool-condition"
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object)
for _, tc := range []struct {
body, match, none string
status int
}{
{"missing", "*", "", http.StatusNotFound},
{"first", "", "*", http.StatusOK},
{"blocked", "", "*", http.StatusPreconditionFailed},
{"blocked", "stale", "", http.StatusPreconditionFailed},
{"second", fmt.Sprintf("%x", md5.Sum([]byte("first"))), "", http.StatusOK},
} {
rec := multipartConditionRequest(t, router, http.MethodPut, url, tc.body, map[string]string{xhttp.IfMatch: tc.match, xhttp.IfNoneMatch: tc.none})
if rec.Code != tc.status {
t.Fatalf("PUT %d want %d: %s", rec.Code, tc.status, rec.Body.String())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "second" {
t.Fatalf("single-pool GET: %d %q", get.Code, get.Body.String())
}
}
// An internal replica callback without an addressed version is not a public
// condition. Preserve its availability when another pool is unreadable. An
// addressed replica already requires all pools for lock/tag reconciliation.
func TestPoolsConditionalPutReplicaAvailability(t *testing.T) {
for _, addressed := range []bool{false, true} {
t.Run(fmt.Sprintf("addressed=%t", addressed), func(t *testing.T) {
z, _ := consistencyPools(t)
mode := "unversioned"
if addressed {
mode = "versioned"
}
bucket, router := conditionalPutBucket(t, z, mode)
object := "replica-availability"
oi := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: addressed, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
restoreFault := conditionalPutSwapDisks(z.serverPools[1], object, func(disk StorageAPI) StorageAPI {
return consistencyReadFaultDisk{StorageAPI: disk, bucket: bucket, object: object}
})
defer restoreFault()
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
xhttp.MinIOSourceETag: fmt.Sprintf("%x", md5.Sum([]byte("replacement"))),
}
url := getPutObjectURL("", bucket, object)
wantStatus, wantBody := http.StatusOK, "replacement"
if addressed {
url += "?versionId=" + oi.VersionID
wantStatus, wantBody = http.StatusServiceUnavailable, "old"
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != wantStatus {
t.Fatalf("replica PUT: %d want %d: %s", put.Code, wantStatus, put.Body.String())
}
restoreFault()
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != wantBody {
t.Fatalf("replica GET: %d %q", get.Code, get.Body.String())
}
})
}
}
func TestPoolsConditionalPutDestinationVersionHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "client-version"
old := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object) + "?versionId=" + old.VersionID
for _, stale := range []bool{true, false} {
tag, status := current.ETag, http.StatusOK
if stale {
tag, status = old.ETag, http.StatusPreconditionFailed
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", tag)})
if put.Code != status {
t.Fatalf("version-addressed public PUT: %d want %d: %s", put.Code, status, put.Body.String())
}
if got := put.Header()[xhttp.AmzVersionID]; !stale && (len(got) != 1 || got[0] != old.VersionID) {
t.Fatalf("write version changed: %v", put.Header())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "replacement" {
t.Fatalf("addressed GET: %d %q", get.Code, get.Body.String())
}
}
func TestPoolsConditionalPutDeleteMarkerTie(t *testing.T) {
for markerPool := range 2 {
t.Run(fmt.Sprintf("marker-pool=%d", markerPool), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "marker-tie"
mtime := UTCNow().Add(-time.Minute)
putConsistencyObject(t, z, bucket, object, 1-markerPool, "live", ObjectOptions{Versioned: true, MTime: mtime})
if _, err := z.serverPools[markerPool].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: mtime}); err != nil {
t.Fatal(err)
}
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object)
before := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
want := http.StatusPreconditionFailed
if markerPool == 0 {
want = http.StatusOK
if before.Code != http.StatusNotFound {
t.Fatalf("GET tie: %d", before.Code)
}
} else if before.Code != http.StatusOK {
t.Fatalf("GET tie: %d", before.Code)
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", map[string]string{xhttp.IfNoneMatch: "*"})
if put.Code != want {
t.Fatalf("PUT tie: %d want %d: %s", put.Code, want, put.Body.String())
}
})
}
}
// Overlap fixture changes with the real IAM Walk reader instead of relying on
// its periodic refresh timer to expose an unsynchronized disk-adapter swap.
func TestPoolsConditionalPutFixtureConcurrentIAM(t *testing.T) {
z, _ := consistencyPools(t)
conditionalPutBucket(t, z, "unversioned")
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
started, finished := make(chan struct{}), make(chan error, 1)
iam := globalIAMSys
go func() {
close(started)
for range 20 {
if err := iam.Load(ctx, false); err != nil {
finished <- err
return
}
}
finished <- nil
}()
<-started
for i := range 5000 {
restore := conditionalPutPool(t, z, "fixture-concurrent-iam", i%2)
restore()
}
if err := <-finished; err != nil {
t.Fatal(err)
}
}
+459
View File
@@ -0,0 +1,459 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"context"
"fmt"
"maps"
"strings"
"testing"
xhttp "github.com/minio/minio/internal/http"
)
// The fixture places copies directly in real erasure pools. It models a
// duplicated version; it does not claim to exercise a rebalance workflow.
func TestPoolsMetadataUpdatePreservesTags(t *testing.T) {
z, bucket := consistencyPools(t)
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, test := range []struct {
name string
tags, stamps [2]string
winner int
single bool
}{
{name: "newer-secondary", tags: [2]string{"key=old", "key=new"}, stamps: [2]string{old, recent}, winner: 1},
{name: "newer-primary", tags: [2]string{"key=new", "key=old"}, stamps: [2]string{recent, old}},
{name: "empty-secondary", tags: [2]string{"key=old", ""}, stamps: [2]string{old, recent}, winner: 1},
{name: "empty-primary", tags: [2]string{"", "key=old"}, stamps: [2]string{recent, old}},
{name: "single-copy-primary", tags: [2]string{"key=only", ""}, stamps: [2]string{recent, ""}, single: true},
{name: "single-copy-secondary", tags: [2]string{"", "key=only"}, stamps: [2]string{"", recent}, winner: 1, single: true},
{name: "legacy-single-copy", tags: [2]string{"key=legacy", ""}, single: true},
{name: "legacy-duplicates", tags: [2]string{"key=legacy", "key=legacy"}},
} {
t.Run(test.name, func(t *testing.T) {
object := test.name
var original ObjectInfo
for pool := range 2 {
if test.single && pool != test.winner {
continue
}
metadata := map[string]string{
xhttp.AmzObjectTagging: test.tags[pool],
"copy-local": fmt.Sprint(pool),
}
if test.stamps[pool] != "" {
metadata[timestamp] = test.stamps[pool]
}
original = putConsistencyObject(t, z, bucket, object, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, UserDefined: metadata,
})
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: original.VersionID})
if err != nil || got.UserTags != test.tags[pool] || got.UserDefined[timestamp] != test.stamps[pool] {
t.Fatalf("pool %d fixture: tags=%q timestamp=%q err=%v", pool, got.UserTags, got.UserDefined[timestamp], err)
}
if _, exists := got.UserDefined[xhttp.AmzObjectTagging]; exists {
t.Fatal("fixture must use the cleaned ObjectInfo representation")
}
}
wantTags, wantStamp := test.tags[test.winner], test.stamps[test.winner]
called := 0
got, err := z.PutObjectMetadata(t.Context(), bucket, object, ObjectOptions{
VersionID: original.VersionID, MTime: original.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
called++
if current.UserTags != wantTags || current.UserDefined[timestamp] != wantStamp {
t.Errorf("callback tags=%q timestamp=%q; want %q %q", current.UserTags, current.UserDefined[timestamp], wantTags, wantStamp)
}
current.UserDefined["unrelated-update"] = "preserved"
return ReplicateDecision{}, nil
},
})
if err != nil || called != 1 {
t.Fatalf("metadata update: %v, callbacks=%d", err, called)
}
if got.UserTags != wantTags || got.UserDefined[timestamp] != wantStamp {
t.Errorf("response tags=%q timestamp=%q; want %q %q", got.UserTags, got.UserDefined[timestamp], wantTags, wantStamp)
}
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: original.VersionID})
if test.single && pool != test.winner {
if !isErrVersionNotFound(err) {
t.Errorf("metadata update created another copy: %v", err)
}
continue
}
if err != nil {
t.Fatal(err)
}
t.Logf("pool %d persisted tags=%q timestamp=%q", pool, got.UserTags, got.UserDefined[timestamp])
if got.UserTags != wantTags || got.UserDefined[timestamp] != wantStamp {
t.Errorf("pool %d persisted tags=%q timestamp=%q; want %q %q", pool, got.UserTags, got.UserDefined[timestamp], wantTags, wantStamp)
}
if got.UserDefined["unrelated-update"] != "preserved" || got.UserDefined["copy-local"] != fmt.Sprint(pool) {
t.Errorf("pool %d lost unrelated metadata: %v", pool, got.UserDefined)
}
}
})
}
}
func TestReplicaWritesPreserveTagOrdering(t *testing.T) {
z, bucket := consistencyPools(t)
// Pool allocation checks the host's used-space percentage. Present only
// its free space as fixture capacity; all reads and writes still use the
// real disks. This keeps unrelated host disk usage out of the tag test.
for _, pool := range z.serverPools {
for _, set := range pool.sets {
getDisks := set.getDisks
disks := append([]StorageAPI(nil), getDisks()...)
for i := range disks {
disks[i] = tagTestCapacityDisk{StorageAPI: disks[i]}
}
set.getDisks = func() []StorageAPI { return disks }
t.Cleanup(func() { set.getDisks = getDisks })
}
}
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, path := range []struct {
name, operation string
direct bool
winner int
}{
{name: "set-put", operation: "put", direct: true},
{name: "set-multipart", operation: "multipart", direct: true},
{name: "set-copy", operation: "copy", direct: true},
{name: "pools-put", operation: "put", winner: 1},
{name: "pools-multipart", operation: "multipart", winner: 1},
{name: "pools-copy-primary", operation: "copy"},
{name: "pools-copy-secondary", operation: "copy", winner: 1},
} {
for _, test := range []struct {
name, storedTags, incomingTags, storedStamp, incomingStamp, wantTags string
}{
{"stored-newer", "key=stored", "key=incoming", recent, old, "key=stored"},
{"stored-deleted", "", "key=incoming", recent, old, ""},
{"incoming-newer", "key=stored", "key=incoming", old, recent, "key=incoming"},
{"incoming-deleted", "key=stored", "", old, recent, ""},
} {
t.Run(path.name+"/"+test.name, func(t *testing.T) {
object := path.name + "-" + test.name
var original ObjectInfo
for pool := range 2 {
if path.direct && pool != 0 {
continue
}
metadata := map[string]string{
xhttp.AmzObjectTagging: "key=older-copy",
timestamp: "2026-09-09T08:00:00Z",
}
if pool == path.winner {
metadata[xhttp.AmzObjectTagging] = test.storedTags
metadata[timestamp] = test.storedStamp
}
original = putConsistencyObject(t, z, bucket, object, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, UserDefined: metadata,
})
}
opts := ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, ReplicaLockReconcile: true,
UserDefined: map[string]string{
xhttp.AmzObjectTagging: test.incomingTags,
timestamp: test.incomingStamp,
},
}
var got ObjectInfo
var err error
switch path.operation {
case "put":
put := z.PutObject
if path.direct {
put = z.serverPools[0].PutObject
}
got, err = put(t.Context(), bucket, object, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), opts)
case "multipart":
// Persist the incoming tags with the upload, before completion
// reconciles the destination version through its real resolver.
mp, err := z.serverPools[0].NewMultipartUpload(t.Context(), bucket, object, opts)
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[0].PutObjectPart(t.Context(), bucket, object, mp.UploadID, 1,
mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
complete := z.CompleteMultipartUpload
if path.direct {
complete = z.serverPools[0].CompleteMultipartUpload
}
got, err = complete(t.Context(), bucket, object, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, ObjectOptions{
Versioned: true, MTime: original.ModTime, ReplicaLockReconcile: true,
})
if err != nil {
t.Fatal(err)
}
case "copy":
src := original
src.metadataOnly = true
src.UserDefined = maps.Clone(opts.UserDefined)
copyObject := z.CopyObject
if path.direct {
copyObject = z.serverPools[0].CopyObject
}
got, err = copyObject(t.Context(), bucket, object, bucket, object, src, ObjectOptions{VersionID: original.VersionID}, opts)
}
if err != nil {
t.Fatal(err)
}
if got.UserTags != test.wantTags || got.UserDefined[timestamp] != recent {
t.Errorf("response tags=%q timestamp=%q; want %q %q", got.UserTags, got.UserDefined[timestamp], test.wantTags, recent)
}
copies := 0
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: original.VersionID})
if isErrVersionNotFound(err) {
continue
}
if err != nil {
t.Fatal(err)
}
copies++
if got.UserTags != test.wantTags || got.UserDefined[timestamp] != recent {
t.Errorf("pool %d persisted tags=%q timestamp=%q; want %q %q", pool, got.UserTags, got.UserDefined[timestamp], test.wantTags, recent)
}
}
if copies != 1 {
t.Errorf("replacement left %d copies; want 1", copies)
}
})
}
}
}
type tagTestCapacityDisk struct{ StorageAPI }
func (d tagTestCapacityDisk) DiskInfo(ctx context.Context, opts DiskInfoOptions) (DiskInfo, error) {
info, err := d.StorageAPI.DiskInfo(ctx, opts)
info.Total, info.Used = info.Free, 0
return info, err
}
func TestMergedPoolObjectInfoTagOrdering(t *testing.T) {
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, test := range []struct {
name, firstStamp, secondStamp, secondTags string
winner int
}{
{"newer", old, recent, "key=second", 1},
{"newer-removal", old, recent, "", 1},
{"equal", recent, recent, "key=second", 0},
{"unordered", "", "", "key=second", 0},
{"missing-first", "", recent, "key=second", 1},
{"missing-second", recent, "", "key=second", 0},
{"invalid-first", "invalid", recent, "key=second", 1},
{"invalid-second", recent, "invalid", "key=second", 0},
} {
t.Run(test.name, func(t *testing.T) {
tagValues := []string{"key=first", test.secondTags}
stamps := []string{test.firstStamp, test.secondStamp}
copies := make([]PoolObjInfo, 2)
before := make([]map[string]string, 2)
for i := range copies {
fi := FileInfo{Metadata: map[string]string{xhttp.AmzObjectTagging: tagValues[i]}}
if stamps[i] != "" {
fi.Metadata[timestamp] = stamps[i]
}
copies[i] = PoolObjInfo{Index: i, ObjInfo: fi.ToObjectInfo("bucket", "object", true)}
before[i] = maps.Clone(copies[i].ObjInfo.UserDefined)
}
got := mergedPoolObjectInfo(copies)
if got.UserTags != tagValues[test.winner] || got.UserDefined[timestamp] != stamps[test.winner] {
t.Errorf("merged tags=%q timestamp=%q; want %q %q", got.UserTags, got.UserDefined[timestamp], tagValues[test.winner], stamps[test.winner])
}
if _, exists := got.UserDefined[xhttp.AmzObjectTagging]; exists {
t.Error("merged ObjectInfo leaked the raw tagging key into UserDefined")
}
for i := range copies {
if !maps.Equal(copies[i].ObjInfo.UserDefined, before[i]) || copies[i].ObjInfo.UserTags != tagValues[i] {
t.Errorf("merge mutated input copy %d", i)
}
}
})
}
}
func TestPoolsMetadataCallbackReplacesTags(t *testing.T) {
z, bucket := consistencyPools(t)
const timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
for _, test := range []struct{ name, tags string }{
{"replace", "key=callback"},
{"remove", ""},
} {
t.Run(test.name, func(t *testing.T) {
var original ObjectInfo
for pool := range 2 {
original = putConsistencyObject(t, z, bucket, test.name, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime,
UserDefined: map[string]string{
xhttp.AmzObjectTagging: []string{"key=old", "key=new"}[pool],
timestamp: []string{"2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"}[pool],
},
})
}
const updatedStamp = "2026-09-09T11:00:00Z"
got, err := z.PutObjectMetadata(t.Context(), bucket, test.name, ObjectOptions{
VersionID: original.VersionID, MTime: original.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
if current.UserTags != "key=new" {
t.Errorf("callback read tags=%q; want key=new", current.UserTags)
}
if _, exists := current.UserDefined[xhttp.AmzObjectTagging]; exists {
t.Error("callback received the raw tagging key")
}
current.UserDefined[xhttp.AmzObjectTagging] = test.tags
current.UserDefined[timestamp] = updatedStamp
return ReplicateDecision{}, nil
},
})
if err != nil {
t.Fatal(err)
}
if got.UserTags != test.tags || got.UserDefined[timestamp] != updatedStamp {
t.Errorf("callback update response tags=%q timestamp=%q", got.UserTags, got.UserDefined[timestamp])
}
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, test.name, ObjectOptions{VersionID: original.VersionID})
if err != nil || got.UserTags != test.tags || got.UserDefined[timestamp] != updatedStamp {
t.Errorf("pool %d did not persist callback tags: tags=%q timestamp=%q err=%v", pool, got.UserTags, got.UserDefined[timestamp], err)
}
}
})
}
}
func TestReconcileStoredObjectTagOrdering(t *testing.T) {
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, test := range []struct {
name, storedStamp, incomingStamp, storedTags string
wantStored bool
}{
{"stored-newer", recent, old, "key=stored", true},
{"incoming-newer", old, recent, "key=stored", false},
{"equal", recent, recent, "key=stored", true},
{"equal-removal", recent, recent, "", true},
{"missing-stored", "", recent, "key=stored", false},
{"missing-incoming", recent, "", "key=stored", true},
{"invalid-stored", "invalid", recent, "key=stored", false},
{"invalid-incoming", recent, "invalid", "key=stored", true},
} {
t.Run(test.name, func(t *testing.T) {
metadata := map[string]string{
xhttp.AmzObjectTagging: "key=incoming",
timestamp: test.incomingStamp,
"unrelated": "preserved",
}
reconcileStoredObjectTags(metadata, test.storedTags, test.storedStamp)
wantTags, wantStamp := "key=incoming", test.incomingStamp
if test.wantStored {
wantTags, wantStamp = test.storedTags, test.storedStamp
}
if metadata[xhttp.AmzObjectTagging] != wantTags || metadata[timestamp] != wantStamp || metadata["unrelated"] != "preserved" {
t.Errorf("reconciled metadata=%v; want tags=%q timestamp=%q and unrelated field preserved", metadata, wantTags, wantStamp)
}
})
}
}
func TestPoolsMetadataUpdatePreservesAbsentTags(t *testing.T) {
z, bucket := consistencyPools(t)
const (
old = "2026-09-09T09:00:00Z"
recent = "2026-09-09T10:00:00Z"
timestamp = ReservedMetadataPrefixLower + TaggingTimestamp
)
for _, removal := range []bool{false, true} {
t.Run(fmt.Sprintf("timestamp-only-removal=%t", removal), func(t *testing.T) {
object := fmt.Sprintf("absent-tags-%t", removal)
var original ObjectInfo
checkStored := func(pool int, wantKey bool, wantTags, wantStamp string) {
t.Helper()
infos, errs := readAllFileInfo(t.Context(), z.serverPools[pool].getHashedSet(object).getDisks(), "", bucket, object, original.VersionID, false, false)
for disk, info := range infos {
if errs[disk] != nil {
t.Fatal(errs[disk])
}
tags, exists := info.Metadata[xhttp.AmzObjectTagging]
if exists != wantKey || tags != wantTags || info.Metadata[timestamp] != wantStamp {
t.Errorf("pool %d disk %d raw tagging key=%t value=%q stamp=%q; want %t %q %q", pool, disk, exists, tags, info.Metadata[timestamp], wantKey, wantTags, wantStamp)
}
}
}
for pool := range 2 {
metadata := map[string]string{}
if removal {
metadata[timestamp] = recent
if pool == 1 {
metadata[xhttp.AmzObjectTagging] = "key=old"
metadata[timestamp] = old
}
}
original = putConsistencyObject(t, z, bucket, object, pool, "data", ObjectOptions{
Versioned: true, VersionID: original.VersionID, MTime: original.ModTime, UserDefined: metadata,
})
checkStored(pool, removal && pool == 1, metadata[xhttp.AmzObjectTagging], metadata[timestamp])
}
_, err := z.PutObjectMetadata(t.Context(), bucket, object, ObjectOptions{
VersionID: original.VersionID, MTime: original.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
current.UserDefined["unrelated-update"] = "preserved"
return ReplicateDecision{}, nil
},
})
if err != nil {
t.Fatal(err)
}
wantStamp := ""
if removal {
wantStamp = recent
}
for pool := range 2 {
// A previously non-empty key needs an explicit empty value to
// propagate deletion; an absent key should remain absent.
checkStored(pool, removal && pool == 1, "", wantStamp)
}
})
}
}
+261 -51
View File
@@ -635,7 +635,14 @@ func (z *erasureServerPools) getPoolIdxNoLock(ctx context.Context, bucket, objec
// if none are found falls back to most available space pool, this function is
// designed to be only used by PutObject, CopyObject (newObject creation) and NewMultipartUpload.
func (z *erasureServerPools) getPoolIdx(ctx context.Context, bucket, object string, size int64) (idx int, err error) {
return z.getWritePoolIdx(ctx, bucket, object, size, false)
}
// getWritePoolIdx keeps the write-allocation policy when the caller already
// holds the pools-layer object lock and must not reacquire a set read lock.
func (z *erasureServerPools) getWritePoolIdx(ctx context.Context, bucket, object string, size int64, noLock bool) (idx int, err error) {
pinfo, _, err := z.getPoolInfoExistingWithOpts(ctx, bucket, object, ObjectOptions{
NoLock: noLock,
SkipDecommissioned: true,
SkipRebalancing: true,
})
@@ -1131,7 +1138,52 @@ func (z *erasureServerPools) PutObject(ctx context.Context, bucket string, objec
return z.serverPools[0].PutObject(ctx, bucket, object, data, opts)
}
idx, err := z.getPoolIdx(ctx, bucket, object, data.Size())
if !opts.NoLock {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return ObjectInfo{}, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
}
opts.NoLock = true
// Public write conditions compare the logical current object while the
// pools-layer write lock is held. The destination selected by capacity may
// be empty or stale, and draining pools can still hold the current object.
// Replica callbacks retain their existing addressed-version semantics and
// metadata reconciliation at the set layer.
if opts.CheckPrecondFn != nil && !opts.ReplicationRequest &&
!opts.ReplicaLockReconcile && !opts.DataMovement {
copies, lerr := z.objectPoolInfos(ctx, bucket, object, ObjectOptions{
VersionID: "", // Compare the current object, not the write's version.
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoAuditLog: true,
})
var latest ObjectInfo
if lerr == nil {
latest = copies[0].ObjInfo
if latest.DeleteMarker {
lerr = toObjectErr(errFileNotFound, bucket, object)
}
}
// An unreadable pool may hold the newest object; it is not absence.
if lerr != nil && !isErrObjectNotFound(lerr) && !isErrVersionNotFound(lerr) {
return ObjectInfo{}, lerr
}
if lerr == nil && opts.CheckPrecondFn(latest) {
return ObjectInfo{}, PreConditionFailed{}
}
if lerr != nil && opts.HasIfMatch {
return ObjectInfo{}, lerr
}
// Do not repeat an accepted condition against the destination's copy.
opts.CheckPrecondFn = nil
}
idx, err := z.getWritePoolIdx(ctx, bucket, object, data.Size(), true)
if err != nil {
return ObjectInfo{}, err
}
@@ -1145,7 +1197,15 @@ func (z *erasureServerPools) PutObject(ctx context.Context, bucket string, objec
}
}
return z.serverPools[idx].PutObject(ctx, bucket, object, data, opts)
if opts.ReplicaLockReconcile || opts.DataMovement {
opts.ReplicaLockReconcile = true
opts.replicaObjectInfo = z.replicaObjectInfo
}
oi, err := z.serverPools[idx].PutObject(ctx, bucket, object, data, opts)
if err == nil && opts.ReplicaLockReconcile && !opts.DataMovement {
err = z.retireReplicaCopies(ctx, bucket, object, idx, oi)
}
return oi, err
}
func (z *erasureServerPools) deletePrefix(ctx context.Context, bucket string, prefix string) error {
@@ -1166,7 +1226,7 @@ func (z *erasureServerPools) DeleteObject(ctx context.Context, bucket string, ob
object = encodeDirObject(object)
}
// Acquire a write lock before deleting the object.
// Serialize logical deletion with pooled writes and metadata updates.
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalDeleteOperationTimeout)
if err != nil {
@@ -1179,6 +1239,13 @@ func (z *erasureServerPools) DeleteObject(ctx context.Context, bucket string, ob
return ObjectInfo{}, z.deletePrefix(ctx, bucket, object)
}
// Reconcile ordinary addressed-version deletes independently of pool movement.
reconcileVersion := opts.VersionID != "" && !opts.DataMovement &&
!opts.ReplicationRequest && !opts.Expiration.Expire && !opts.InclFreeVersions
if !z.SinglePool() && (opts.CheckPrecondFn != nil || reconcileVersion) {
return z.deleteObjectReconciled(ctx, bucket, object, opts)
}
gopts := opts
gopts.NoLock = true
@@ -1187,9 +1254,55 @@ func (z *erasureServerPools) DeleteObject(ctx context.Context, bucket string, ob
if _, ok := err.(InsufficientReadQuorum); ok {
return objInfo, InsufficientWriteQuorum{}
}
// A conditional (If-Match) delete addressing a specific version treats an
// absent key as an absent version. getPoolInfoExistingWithOpts strips
// VersionID, so a missing key surfaces ObjectNotFound here even for a
// version-scoped delete; normalize it to VersionNotFound (NoSuchVersion),
// matching this function's tail. The unconditional path is unchanged.
if opts.CheckPrecondFn != nil && opts.VersionID != "" && isErrObjectNotFound(err) {
return objInfo, VersionNotFound{Bucket: bucket, Object: object, VersionID: opts.VersionID}
}
return objInfo, err
}
// Evaluate the conditional (If-Match) precondition while the write lock
// acquired above is held, before the delete-marker short-circuit and before
// any version is removed, so the object cannot change between the check and
// the delete. This is scoped to a single erasure set (see the note at the
// lock above): only there do the delete lock and the write path share the
// same lock namespace, making the check-then-delete atomic.
if opts.CheckPrecondFn != nil {
// pinfo.ObjInfo is the current latest version. getPoolInfoExistingWithOpts
// intentionally strips VersionID, so for a version-scoped delete read the
// specifically addressed version and evaluate the precondition against it.
checkInfo := pinfo.ObjInfo
if opts.VersionID != "" {
vopts := opts
vopts.NoLock = true // delete lock already held above
vopts.CheckPrecondFn = nil
vi, verr := z.serverPools[pinfo.Index].GetObjectInfo(ctx, bucket, object, vopts)
if verr != nil && (!isErrMethodNotAllowed(verr) || !vi.DeleteMarker) {
// Genuine read failure for the addressed version: a missing
// version -> VersionNotFound (NoSuchVersion), read-quorum loss, etc.
return objInfo, verr
}
// verr is nil for a live version, or MethodNotAllowed with a populated
// delete-marker ObjectInfo when the addressed version is a delete
// marker. In the latter case evaluate the precondition against the
// marker, which fails any If-Match (-> 412), rather than surfacing 405.
checkInfo = vi
} else if checkInfo.Name == "" {
// The current state could not be read (e.g. read-quorum loss); refuse
// the conditional delete rather than act on an unverified precondition.
return objInfo, InsufficientReadQuorum{}
}
if opts.CheckPrecondFn(checkInfo) {
return objInfo, PreConditionFailed{}
}
// Precondition satisfied; lower layers must not re-evaluate it.
opts.CheckPrecondFn = nil
}
// Delete marker already present we are not going to create new delete markers.
if pinfo.ObjInfo.DeleteMarker && opts.VersionID == "" {
pinfo.ObjInfo.Name = decodeDirObject(object)
@@ -1268,8 +1381,12 @@ func (z *erasureServerPools) deleteObjectFromAllPools(ctx context.Context, bucke
// the delete call tries to clean up other pools during DeleteObject call.
objInfo = dobjects[0]
objInfo.Name = decodeDirObject(object)
err = derrs[0]
return objInfo, err
for _, derr := range derrs {
if derr != nil && !isErrObjectNotFound(derr) && !isErrVersionNotFound(derr) {
return objInfo, derr
}
}
return objInfo, derrs[0]
}
func (z *erasureServerPools) DeleteObjects(ctx context.Context, bucket string, objects []ObjectToDelete, opts ObjectOptions) ([]DeletedObject, []error) {
@@ -1357,11 +1474,38 @@ func (z *erasureServerPools) CopyObject(ctx context.Context, srcBucket, srcObjec
dstOpts.NoLock = true
}
if !z.SinglePool() && cpSrcDstSame && srcInfo.metadataOnly && dstOpts.ReplicaLockReconcile && srcOpts.VersionID == dstOpts.VersionID {
copies, err := z.metadataPoolInfos(ctx, dstBucket, dstObject, dstOpts)
if err != nil {
return ObjectInfo{}, err
}
stored := mergedPoolObjectInfo(copies)
reconcileStoredObjectLock(srcInfo.UserDefined, storedObjectLockState(stored.UserDefined))
reconcileStoredObjectTags(srcInfo.UserDefined, stored.UserTags, stored.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
idx := copies[0].Index
oi, err := z.serverPools[idx].CopyObject(ctx, srcBucket, srcObject, dstBucket, dstObject, srcInfo, srcOpts, dstOpts)
if err == nil {
err = z.retireReplicaCopies(ctx, dstBucket, dstObject, idx, oi)
}
return oi, err
}
poolIdx, err := z.getPoolIdxNoLock(ctx, dstBucket, dstObject, srcInfo.Size)
if err != nil {
return objInfo, err
}
if !z.SinglePool() && dstOpts.ReplicaLockReconcile {
dstOpts.replicaObjectInfo = z.replicaObjectInfo
if !dstOpts.DataMovement {
defer func() {
if err == nil {
err = z.retireReplicaCopies(ctx, dstBucket, dstObject, poolIdx, objInfo)
}
}()
}
}
// CopyObjectHandler predicts the outcome of this decision in
// copyRewritesObjectData(); keep the two in sync.
if cpSrcDstSame && srcInfo.metadataOnly {
@@ -1396,6 +1540,8 @@ func (z *erasureServerPools) CopyObject(ctx context.Context, srcBucket, srcObjec
EncryptFn: dstOpts.EncryptFn,
WantChecksum: dstOpts.WantChecksum,
WantServerSideChecksumType: dstOpts.WantServerSideChecksumType,
ReplicaLockReconcile: dstOpts.ReplicaLockReconcile,
replicaObjectInfo: dstOpts.replicaObjectInfo,
}
return z.serverPools[poolIdx].PutObject(ctx, dstBucket, dstObject, srcInfo.PutObjReader, putOpts)
@@ -2031,6 +2177,60 @@ func (z *erasureServerPools) CompleteMultipartUpload(ctx context.Context, bucket
}
}()
if !z.SinglePool() {
if !opts.NoLock {
lk := z.NewNSLock(bucket, encodeDirObject(object))
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return objInfo, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
}
opts.NoLock = true
// A conditional completion must be evaluated against the logical
// latest object across pools, under the object write lock held for
// this operation. The pool hosting the upload may only hold a stale
// duplicate, so its set-local check would both accept an outdated
// ETag and reject the current one. An unreadable pool is not
// absence: it may hold the newest copy, so a read that cannot be
// verified fails the request instead of passing the condition.
// Once satisfied, the callback is cleared so the set layer does not
// re-evaluate it against its local copy.
if opts.CheckPrecondFn != nil {
copies, lerr := z.objectPoolInfos(ctx, bucket, encodeDirObject(object), ObjectOptions{
// Conditions always compare the logical current object,
// independently of the completion's destination version.
VersionID: "",
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoAuditLog: true,
})
var latest ObjectInfo
if lerr == nil {
latest = copies[0].ObjInfo
if latest.DeleteMarker {
// A delete-marker latest reads as an absent key, matching
// the set layer's getObjectInfo.
lerr = toObjectErr(errFileNotFound, bucket, object)
}
}
if lerr == nil && opts.CheckPrecondFn(latest) {
return ObjectInfo{}, PreConditionFailed{}
}
if lerr != nil && !isErrVersionNotFound(lerr) && !isErrObjectNotFound(lerr) {
return ObjectInfo{}, lerr
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if lerr != nil && opts.HasIfMatch && (isErrObjectNotFound(lerr) || isErrVersionNotFound(lerr)) {
return ObjectInfo{}, lerr
}
opts.CheckPrecondFn = nil
}
}
// Hold write locks to verify uploaded parts, also disallows any
// parallel PutObjectPart() requests.
uploadIDLock := z.NewNSLock(bucket, pathJoin(object, uploadID))
@@ -2045,13 +2245,21 @@ func (z *erasureServerPools) CompleteMultipartUpload(ctx context.Context, bucket
return z.serverPools[0].CompleteMultipartUpload(ctx, bucket, object, uploadID, uploadedParts, opts)
}
if opts.ReplicaLockReconcile || opts.DataMovement {
opts.ReplicaLockReconcile = true
opts.replicaObjectInfo = z.replicaObjectInfo
}
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) {
continue
}
objInfo, err = pool.CompleteMultipartUpload(ctx, bucket, object, uploadID, uploadedParts, opts)
if err == nil {
return objInfo, nil
if opts.ReplicaLockReconcile && !opts.DataMovement {
err = z.retireReplicaCopies(ctx, bucket, encodeDirObject(object), idx, objInfo)
}
return objInfo, err
}
if _, ok := err.(InvalidUploadID); ok {
// upload id not found move to next pool
@@ -2073,6 +2281,11 @@ func (z *erasureServerPools) GetBucketInfo(ctx context.Context, bucket string, o
if err != nil {
return bucketInfo, toObjectErr(err, bucket)
}
// Physical existence/creation probes must not be overwritten by cached
// metadata, which can legitimately lack Created on an unmigrated bucket.
if opts.NoMetadata {
return bucketInfo, nil
}
meta, err := globalBucketMetadataSys.Get(bucket)
if err == nil {
@@ -2141,7 +2354,15 @@ func (z *erasureServerPools) DeleteBucket(ctx context.Context, bucket string, op
opts.Force = true
}
err := z.s3Peer.DeleteBucket(ctx, bucket, opts)
// Take the metadata writer lock before deleting anything. Failure or
// cancellation must leave both the bucket and its metadata intact.
ctx, unlock, err := lockBucketMetadata(ctx, z, bucket)
if err != nil {
return toObjectErr(err, bucket)
}
defer unlock()
err = z.s3Peer.DeleteBucket(ctx, bucket, opts)
if err == nil || isErrBucketNotFound(err) {
// If site replication is configured, hold on to deleted bucket state until sites sync
if opts.SRDeleteOp == MarkDelete {
@@ -2150,8 +2371,9 @@ func (z *erasureServerPools) DeleteBucket(ctx context.Context, bucket string, op
}
if err == nil {
// Purge the entire bucket metadata entirely.
z.deleteAll(context.Background(), minioMetaBucket, pathJoin(bucketMetaPrefix, bucket))
// Finish cleanup after a committed delete even if the client disconnects.
// Both the bucket-name and metadata locks remain held until return.
z.deleteAll(context.WithoutCancel(ctx), minioMetaBucket, pathJoin(bucketMetaPrefix, bucket))
}
return toObjectErr(err, bucket)
@@ -2569,7 +2791,7 @@ func (z *erasureServerPools) HealObject(ctx context.Context, bucket, object, ver
wg.Add(1)
go func(idx int, pool *erasureSets) {
defer wg.Done()
result, err := pool.HealObject(ctx, bucket, object, versionID, opts)
result, err := z.healObjectInPool(ctx, pool.getHashedSet(object), bucket, object, versionID, opts)
result.Object = decodeDirObject(result.Object)
errs[idx] = err
results[idx] = result
@@ -2838,15 +3060,7 @@ func (z *erasureServerPools) PutObjectMetadata(ctx context.Context, bucket, obje
defer lk.Unlock(lkctx)
}
opts.MetadataChg = true
opts.NoLock = true
// We don't know the size here set 1GiB at least.
idx, err := z.getPoolIdxExistingWithOpts(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
return z.serverPools[idx].PutObjectMetadata(ctx, bucket, object, opts)
return z.updatePoolMetadata(ctx, bucket, object, opts)
}
// PutObjectTags - replace or add tags to an existing object
@@ -2867,44 +3081,40 @@ func (z *erasureServerPools) PutObjectTags(ctx context.Context, bucket, object s
defer lk.Unlock(lkctx)
}
opts.MetadataChg = true
opts.NoLock = true
// We don't know the size here set 1GiB at least.
idx, err := z.getPoolIdxExistingWithOpts(ctx, bucket, object, opts)
copies, err := z.metadataPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
return z.serverPools[idx].PutObjectTags(ctx, bucket, object, tags, opts)
// Ordinary reads and replication can return any owning pool. Persist one
// revision beyond all copies, so the returned value and every pool agree.
if stamp := opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]; stamp != "" {
for _, copy := range copies {
stamp = monotonicTaggingTimestamp(stamp, copy.ObjInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
}
opts.UserDefined = cloneMSS(opts.UserDefined)
opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = stamp
}
opts.NoLock = true
opts.VersionID = copies[0].ObjInfo.VersionID
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
var primary ObjectInfo
for _, copy := range copies {
oi, err := z.serverPools[copy.Index].PutObjectTags(ctx, bucket, object, tags, opts)
if err != nil {
return ObjectInfo{}, err
}
if copy.Index == copies[0].Index {
primary = oi
}
}
return primary, nil
}
// DeleteObjectTags - delete object tags from an existing object
func (z *erasureServerPools) DeleteObjectTags(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
object = encodeDirObject(object)
if z.SinglePool() {
return z.serverPools[0].DeleteObjectTags(ctx, bucket, object, opts)
}
if !opts.NoLock {
// Lock the object before deleting tags.
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return ObjectInfo{}, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
}
opts.MetadataChg = true
opts.NoLock = true
idx, err := z.getPoolIdxExistingWithOpts(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
return z.serverPools[idx].DeleteObjectTags(ctx, bucket, object, opts)
return z.PutObjectTags(ctx, bucket, object, "", opts)
}
// GetObjectTags - get object tags from an existing object
@@ -2914,7 +3124,7 @@ func (z *erasureServerPools) GetObjectTags(ctx context.Context, bucket, object s
return z.serverPools[0].GetObjectTags(ctx, bucket, object, opts)
}
oi, _, err := z.getLatestObjectInfoWithIdx(ctx, bucket, object, opts)
oi, err := z.GetObjectInfo(ctx, bucket, object, opts)
if err != nil {
return nil, err
}
+8 -2
View File
@@ -166,6 +166,12 @@ func (er *erasureObjects) healErasureSet(ctx context.Context, buckets []string,
if objAPI == nil {
return errServerNotInitialized
}
healInPool := er.HealObject
if z, ok := objAPI.(*erasureServerPools); ok && !z.SinglePool() {
healInPool = func(ctx context.Context, bucket, object, versionID string, opts madmin.HealOpts) (madmin.HealResultItem, error) {
return z.healObjectInPool(ctx, er, bucket, object, versionID, opts)
}
}
started := tracker.Started
if started.IsZero() || started.Equal(timeSentinel) {
@@ -419,7 +425,7 @@ func (er *erasureObjects) healErasureSet(ctx context.Context, buckets []string,
var result healEntryResult
fivs, err := entry.fileInfoVersions(bucket)
if err != nil {
res, err := er.HealObject(ctx, bucket, encodedEntryName, "",
res, err := healInPool(ctx, bucket, encodedEntryName, "",
madmin.HealOpts{
ScanMode: scanMode,
Remove: healDeleteDangling,
@@ -455,7 +461,7 @@ func (er *erasureObjects) healErasureSet(ctx context.Context, buckets []string,
continue
}
res, err := er.HealObject(ctx, bucket, encodedEntryName,
res, err := healInPool(ctx, bucket, encodedEntryName,
version.VersionID, madmin.HealOpts{
ScanMode: scanMode,
Remove: healDeleteDangling,
+4 -6
View File
@@ -51,9 +51,8 @@ func initGlobalGrid(ctx context.Context, eps EndpointServerPools) error {
grid.ContextDialer(xhttp.DialContextWithLookupHost(lookupHost, xhttp.NewInternodeDialContext(rest.DefaultTimeout, globalTCPOptions.ForWebsocket()))),
newCachedAuthToken(),
&tls.Config{
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
CurvePreferences: crypto.TLSCurveIDs(),
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
}),
Local: local,
Hosts: hosts,
@@ -84,9 +83,8 @@ func initGlobalLockGrid(ctx context.Context, eps EndpointServerPools) error {
grid.ContextDialer(xhttp.DialContextWithLookupHost(lookupHost, xhttp.NewInternodeDialContext(rest.DefaultTimeout, globalTCPOptions.ForWebsocket()))),
newCachedAuthToken(),
&tls.Config{
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
CurvePreferences: crypto.TLSCurveIDs(),
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
}, grid.RouteLockPath),
Local: local,
Hosts: hosts,
+33 -23
View File
@@ -246,16 +246,6 @@ func extractMetadata(ctx context.Context, mimesHeader ...textproto.MIMEHeader) (
// extractMetadata extracts metadata from map values.
func extractMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
return extractMetadataFromMimeWithReplication(ctx, v, m, false)
}
// extractReplicationMetadataFromMime restores replication-only metadata after the
// caller has validated that the request is a trusted replication write.
func extractReplicationMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
return extractMetadataFromMimeWithReplication(ctx, v, m, true)
}
func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIMEHeader, m map[string]string, allowReplication bool) error {
if v == nil {
bugLogIf(ctx, errInvalidArgument)
return errInvalidArgument
@@ -267,18 +257,14 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
nv[http.CanonicalHeaderKey(k)] = kv
}
// Save all supported headers.
// Save ordinary object metadata. Replication-only headers are restored only
// after the request has been validated as a trusted replication write.
for _, supportedHeader := range supportedHeaders {
value, ok := nv[http.CanonicalHeaderKey(supportedHeader)]
if ok {
if v, ok := replicationToInternalHeaders[supportedHeader]; ok {
if !allowReplication {
continue
}
m[v] = strings.Join(value, ",")
} else {
m[supportedHeader] = strings.Join(value, ",")
}
if _, ok := replicationToInternalHeaders[supportedHeader]; ok {
continue
}
if value, ok := nv[http.CanonicalHeaderKey(supportedHeader)]; ok {
m[supportedHeader] = strings.Join(value, ",")
}
}
@@ -287,8 +273,7 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
if !stringsHasPrefixFold(key, prefix) {
continue
}
value, ok := nv[http.CanonicalHeaderKey(key)]
if ok {
if value, ok := nv[http.CanonicalHeaderKey(key)]; ok {
m[key] = strings.Join(value, ",")
break
}
@@ -297,6 +282,31 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
return nil
}
// extractReplicationMetadataFromMime restores replication-only metadata after the
// caller has validated that the request is a trusted replication write.
func extractReplicationMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
if v == nil {
bugLogIf(ctx, errInvalidArgument)
return errInvalidArgument
}
nv := make(textproto.MIMEHeader, len(v))
for k, kv := range v {
// Canonicalize all headers, to remove any duplicates.
nv[http.CanonicalHeaderKey(k)] = kv
}
// Ordinary object metadata belongs to the caller. Re-extracting it would
// undo normalization (such as removing aws-chunked) or copy an outer
// Snowball archive's metadata onto its individual entries.
for header, internalHeader := range replicationToInternalHeaders {
if value, ok := nv[http.CanonicalHeaderKey(header)]; ok {
m[internalHeader] = strings.Join(value, ",")
}
}
return nil
}
// Returns access credentials in the request Authorization header.
func getReqAccessCred(r *http.Request, region string) (cred auth.Credentials) {
cred, _, _ = getReqAccessKeyV4(r, region, serviceS3)
+12 -1
View File
@@ -254,6 +254,9 @@ func TestExtractMetadataFromRequestKeepsQueryCompatibility(t *testing.T) {
func TestExtractReplicationMetadataHeaders(t *testing.T) {
header := http.Header{
"Content-Type": []string{"application/wasm"},
"Content-Encoding": []string{"aws-chunked"},
"X-Amz-Meta-Source": []string{"client"},
"X-Minio-Replication-Server-Side-Encryption-Sealed-Key": []string{"sealed-key"},
"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm": []string{"DAREv2-HMAC-SHA256"},
"X-Minio-Replication-Server-Side-Encryption-Iv": []string{"iv"},
@@ -262,12 +265,17 @@ func TestExtractReplicationMetadataHeaders(t *testing.T) {
ReplicationSsecChecksumHeader: []string{"checksum"},
}
metadata := make(map[string]string)
metadata := map[string]string{
"content-type": "application/wasm",
"x-amz-meta-source": "client",
}
if err := extractReplicationMetadataFromMime(t.Context(), textproto.MIMEHeader(header), metadata); err != nil {
t.Fatalf("failed to extract replication metadata: %v", err)
}
expected := map[string]string{
"content-type": "application/wasm",
"x-amz-meta-source": "client",
"X-Minio-Internal-Server-Side-Encryption-Sealed-Key": "sealed-key",
"X-Minio-Internal-Server-Side-Encryption-Seal-Algorithm": "DAREv2-HMAC-SHA256",
"X-Minio-Internal-Server-Side-Encryption-Iv": "iv",
@@ -279,6 +287,9 @@ func TestExtractReplicationMetadataHeaders(t *testing.T) {
if !reflect.DeepEqual(metadata, expected) {
t.Fatalf("unexpected replication metadata: expected %#v, got %#v", expected, metadata)
}
if _, ok := metadata["content-encoding"]; ok {
t.Fatalf("replication metadata restored transport content-encoding: %#v", metadata)
}
}
func TestGetCopyObjectMetadataFromHeaderReplication(t *testing.T) {
+111
View File
@@ -0,0 +1,111 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"errors"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
func TestIAMCredentialRetention(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
secret, err := getTokenSigningKey()
mustIAM(t, err)
parent := "external-idp-parent"
credential := func(exp time.Time) auth.Credentials {
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": exp.Unix(), parentClaim: parent}, secret)
mustIAM(t, err)
cred.ParentUser = parent
return cred
}
// Disablement of an external identity must include cached STS,
// which are kept separately from regular and service accounts.
cred := credential(UTCNow().Add(time.Hour))
_, err = sys.SetTempUser(ctx, cred.AccessKey, cred, "")
mustIAM(t, err)
mustIAM(t, sys.store.DeleteUsers(ctx, []string{parent}))
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(cred.AccessKey, stsUser))
mustIAM(t, err)
if !r.Deleted || !r.ExpiresAt.Equal(cred.Expiration.Add(globalMaxSkewTime)) || r.Credentials.SessionToken != "" || r.Credentials.SecretKey != "" {
t.Fatal("early STS revocation lost its retention boundary or retained a secret")
}
if _, ok := sys.store.GetUser(cred.AccessKey); ok {
t.Fatal("external disablement left the STS cache live")
}
_, err = sys.SetTempUser(withIAMReplicationTime(ctx, UTCNow().Add(time.Minute)), cred.AccessKey, cred, "")
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("same revoked token was reissued by replay: %v", err)
}
var mp MappedPolicy
err = sys.store.loadIAMConfig(ctx, &mp, getMappedPolicyPath(cred.AccessKey, stsUser, false))
if !errors.Is(err, errConfigNotFound) {
t.Fatalf("random STS key produced a permanent mapping: %v", err)
}
// Seed genuinely expired immutable tokens, as an ordinary startup
// loader sees them. Natural expiry leaves no permanent tombstone.
expired := credential(UTCNow().Add(-time.Hour))
path := getUserIdentityPath(expired.AccessKey, stsUser)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: expired, UpdatedAt: UTCNow().Add(-2 * time.Hour)}, path))
_ = sys.store.loadUser(ctx, expired.AccessKey, stsUser, make(map[string]UserIdentity))
var u UserIdentity
if err := sys.store.loadIAMConfig(ctx, &u, path); !errors.Is(err, errConfigNotFound) {
t.Fatalf("natural expiration retained a random key: %v", err)
}
// A retained early-revocation record is collectable only after the
// immutable token's expiration plus the skew allowance.
tomb := UserIdentity{Version: 1, Deleted: true, UpdatedAt: UTCNow().Add(-2 * time.Hour), ExpiresAt: expired.Expiration.Add(globalMaxSkewTime)}
mustIAM(t, sys.store.saveIAMConfig(ctx, &tomb, path))
_ = sys.store.loadUser(ctx, expired.AccessKey, stsUser, make(map[string]UserIdentity))
if err := sys.store.loadIAMConfig(ctx, &u, path); !errors.Is(err, errConfigNotFound) {
t.Fatalf("expired STS revocation not collected: %v", err)
}
if _, ok := sys.store.revisionIndex().snapshot()[path]; ok {
t.Fatal("expired STS retained an index entry")
}
})
}
}
func TestIAMPolicyDeletionRemainsExplicit(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
mustIAM(t, sys.DeletePolicy(ctx, "misspelled-policy", true))
r, err := loadIAMRevision(ctx, sys.store, getPolicyDocPath("misspelled-policy"))
mustIAM(t, err)
if r.Deleted {
t.Fatal("local nonexistent policy created a tombstone")
}
p, err := sys.store.GetPolicy("readwrite")
mustIAM(t, err)
if err := sys.DeletePolicy(ctx, "readwrite", true); err == nil {
t.Fatal("local pristine builtin policy became deletable")
}
_, err = sys.SetPolicy(ctx, "readwrite", p)
mustIAM(t, err)
mustIAM(t, sys.DeletePolicy(ctx, "readwrite", true))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
if _, err := sys.store.GetPolicy("readwrite"); !errors.Is(err, errNoSuchPolicy) {
t.Fatalf("reload restored an explicitly deleted override: %v", err)
}
_, err = sys.SetPolicy(ctx, "readwrite", p)
mustIAM(t, err)
if _, err := sys.store.GetPolicy("readwrite"); err != nil {
t.Fatal("explicit policy recreation failed", err)
}
mustIAM(t, globalSiteReplicationSys.PeerAddPolicyHandler(ctx, "remote-unknown-policy", nil, UTCNow()))
r, err = loadIAMRevision(ctx, sys.store, getPolicyDocPath("remote-unknown-policy"))
mustIAM(t, err)
if !r.Deleted {
t.Fatal("replicated unknown deletion lost its version")
}
})
}
}
+65 -71
View File
@@ -26,7 +26,6 @@ import (
"sync"
"time"
jsoniter "github.com/json-iterator/go"
"github.com/minio/minio-go/v7/pkg/set"
"github.com/minio/minio/internal/config"
"github.com/minio/minio/internal/kms"
@@ -62,6 +61,7 @@ type IAMEtcdStore struct {
sync.RWMutex
*iamCache
index iamRevisionIndex
usersSysType UsersSysType
@@ -69,13 +69,17 @@ type IAMEtcdStore struct {
}
func newIAMEtcdStore(client *etcd.Client, usersSysType UsersSysType) *IAMEtcdStore {
return &IAMEtcdStore{
store := &IAMEtcdStore{
iamCache: newIamCache(),
client: client,
usersSysType: usersSysType,
}
store.revisions = &store.index
return store
}
func (ies *IAMEtcdStore) revisionIndex() *iamRevisionIndex { return &ies.index }
func (ies *IAMEtcdStore) rlock() *iamCache {
ies.RLock()
return ies.iamCache
@@ -103,6 +107,7 @@ func (ies *IAMEtcdStore) saveIAMConfig(ctx context.Context, item any, itemPath s
if err != nil {
return err
}
plain := data
if GlobalKMS != nil {
data, err = config.EncryptBytes(GlobalKMS, data, kms.Context{
minioMetaBucket: path.Join(minioMetaBucket, itemPath),
@@ -111,24 +116,28 @@ func (ies *IAMEtcdStore) saveIAMConfig(ctx context.Context, item any, itemPath s
return err
}
}
return saveKeyEtcd(ctx, ies.client, itemPath, data, opts...)
if err := saveKeyEtcd(ctx, ies.client, itemPath, data, opts...); err != nil {
return err
}
ies.index.observe(itemPath, plain)
return nil
}
func getIAMConfig(item any, data []byte, itemPath string) error {
data, err := decryptData(data, itemPath)
func (ies *IAMEtcdStore) decodeIAMConfig(item any, data []byte, path string) error {
data, err := decryptData(data, path)
if err != nil {
return err
}
json := jsoniter.ConfigCompatibleWithStandardLibrary
ies.index.observe(path, data)
return json.Unmarshal(data, item)
}
func (ies *IAMEtcdStore) loadIAMConfig(ctx context.Context, item any, path string) error {
data, err := readKeyEtcd(ctx, ies.client, path)
data, err := ies.loadIAMConfigBytes(ctx, path)
if err != nil {
return err
}
return getIAMConfig(item, data, path)
return json.Unmarshal(data, item)
}
func (ies *IAMEtcdStore) loadIAMConfigBytes(ctx context.Context, path string) ([]byte, error) {
@@ -136,11 +145,19 @@ func (ies *IAMEtcdStore) loadIAMConfigBytes(ctx context.Context, path string) ([
if err != nil {
return nil, err
}
return decryptData(data, path)
data, err = decryptData(data, path)
if err == nil {
ies.index.observe(path, data)
}
return data, err
}
func (ies *IAMEtcdStore) deleteIAMConfig(ctx context.Context, path string) error {
return deleteKeyEtcd(ctx, ies.client, path)
if err := deleteKeyEtcd(ctx, ies.client, path); err != nil {
return err
}
ies.index.forget(path)
return nil
}
func (ies *IAMEtcdStore) loadPolicyDocWithRetry(ctx context.Context, policy string, m map[string]PolicyDoc, _ int) error {
@@ -162,6 +179,9 @@ func (ies *IAMEtcdStore) loadPolicyDoc(ctx context.Context, policy string, m map
return err
}
if p.Deleted {
return errNoSuchPolicy
}
m[policy] = p
return nil
}
@@ -181,7 +201,11 @@ func (ies *IAMEtcdStore) getPolicyDocKV(ctx context.Context, kvs *mvccpb.KeyValu
return err
}
ies.index.observe(string(kvs.Key), data)
policy := extractPathPrefixAndSuffix(string(kvs.Key), iamConfigPoliciesPrefix, path.Base(string(kvs.Key)))
if p.Deleted {
return errNoSuchPolicy
}
m[policy] = p
return nil
}
@@ -207,7 +231,7 @@ func (ies *IAMEtcdStore) loadPolicyDocs(ctx context.Context, m map[string]Policy
func (ies *IAMEtcdStore) getUserKV(ctx context.Context, userkv *mvccpb.KeyValue, userType IAMUserType, m map[string]UserIdentity, basePrefix string) error {
var u UserIdentity
err := getIAMConfig(&u, userkv.Value, string(userkv.Key))
err := ies.decodeIAMConfig(&u, userkv.Value, string(userkv.Key))
if err != nil {
if err == errConfigNotFound {
return errNoSuchUser
@@ -219,10 +243,14 @@ func (ies *IAMEtcdStore) getUserKV(ctx context.Context, userkv *mvccpb.KeyValue,
}
func (ies *IAMEtcdStore) addUser(ctx context.Context, user string, userType IAMUserType, u UserIdentity, m map[string]UserIdentity) error {
if u.Deleted {
if userType == stsUser && !u.ExpiresAt.IsZero() && UTCNow().After(u.ExpiresAt) {
bestEffortIAMExpiration(ctx, ies, getUserIdentityPath(user, userType))
}
return errNoSuchUser
}
if u.Credentials.IsExpired() {
// Delete expired identity.
deleteKeyEtcd(ctx, ies.client, getUserIdentityPath(user, userType))
deleteKeyEtcd(ctx, ies.client, getMappedPolicyPath(user, userType, false))
bestEffortIAMExpiration(ctx, ies, getUserIdentityPath(user, userType))
return nil
}
if u.Credentials.AccessKey == "" {
@@ -231,16 +259,17 @@ func (ies *IAMEtcdStore) addUser(ctx context.Context, user string, userType IAMU
if u.Credentials.SessionToken != "" {
jwtClaims, err := extractJWTClaims(u)
if err != nil {
if u.Credentials.IsTemp() {
// We should delete such that the client can re-request
// for the expiring credentials.
deleteKeyEtcd(ctx, ies.client, getUserIdentityPath(user, userType))
deleteKeyEtcd(ctx, ies.client, getMappedPolicyPath(user, userType, false))
}
// A temporarily unavailable signing key is not proof of expiration.
return nil
}
u.Credentials.Claims = jwtClaims.Map()
}
if err := checkIAMParentRevision(ctx, ies, u.Credentials); err != nil {
if errors.Is(err, errIAMStaleUpdate) {
return errNoSuchUser
}
return err
}
if u.Credentials.Description == "" {
u.Credentials.Description = u.Credentials.Comment
}
@@ -258,6 +287,9 @@ func (ies *IAMEtcdStore) loadSecretKey(ctx context.Context, user string, userTyp
}
return "", err
}
if u.Deleted {
return "", errNoSuchUser
}
return u.Credentials.SecretKey, nil
}
@@ -274,6 +306,7 @@ func (ies *IAMEtcdStore) loadUser(ctx context.Context, user string, userType IAM
}
func (ies *IAMEtcdStore) loadUsers(ctx context.Context, userType IAMUserType, m map[string]UserIdentity) error {
ctx = withIAMExpirationCleanup(ctx)
var basePrefix string
switch userType {
case svcUser:
@@ -312,6 +345,9 @@ func (ies *IAMEtcdStore) loadGroup(ctx context.Context, group string, m map[stri
}
return err
}
if gi.Deleted {
return errNoSuchGroup
}
m[group] = gi
return nil
}
@@ -349,13 +385,16 @@ func (ies *IAMEtcdStore) loadMappedPolicy(ctx context.Context, name string, user
}
return err
}
if !ies.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
func getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy], basePrefix string) error {
func (ies *IAMEtcdStore) getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy], basePrefix string) error {
var p MappedPolicy
err := getIAMConfig(&p, kv.Value, string(kv.Key))
err := ies.decodeIAMConfig(&p, kv.Value, string(kv.Key))
if err != nil {
if err == errConfigNotFound {
return errNoSuchPolicy
@@ -363,6 +402,9 @@ func getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy],
return err
}
name := extractPathPrefixAndSuffix(string(kv.Key), basePrefix, ".json")
if !ies.index.mappingAllowed(string(kv.Key), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
@@ -392,61 +434,13 @@ func (ies *IAMEtcdStore) loadMappedPolicies(ctx context.Context, userType IAMUse
// Parse all policies mapping to create the proper data model
for _, kv := range r.Kvs {
if err = getMappedPolicy(kv, m, basePrefix); err != nil && !errors.Is(err, errNoSuchPolicy) {
if err = ies.getMappedPolicy(kv, m, basePrefix); err != nil && !errors.Is(err, errNoSuchPolicy) {
return err
}
}
return nil
}
func (ies *IAMEtcdStore) savePolicyDoc(ctx context.Context, policyName string, p PolicyDoc) error {
return ies.saveIAMConfig(ctx, &p, getPolicyDocPath(policyName))
}
func (ies *IAMEtcdStore) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp MappedPolicy, opts ...options) error {
return ies.saveIAMConfig(ctx, mp, getMappedPolicyPath(name, userType, isGroup), opts...)
}
func (ies *IAMEtcdStore) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u UserIdentity, opts ...options) error {
return ies.saveIAMConfig(ctx, u, getUserIdentityPath(name, userType), opts...)
}
func (ies *IAMEtcdStore) saveGroupInfo(ctx context.Context, name string, gi GroupInfo) error {
return ies.saveIAMConfig(ctx, gi, getGroupInfoPath(name))
}
func (ies *IAMEtcdStore) deletePolicyDoc(ctx context.Context, name string) error {
err := ies.deleteIAMConfig(ctx, getPolicyDocPath(name))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (ies *IAMEtcdStore) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
err := ies.deleteIAMConfig(ctx, getMappedPolicyPath(name, userType, isGroup))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (ies *IAMEtcdStore) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
err := ies.deleteIAMConfig(ctx, getUserIdentityPath(name, userType))
if err == errConfigNotFound {
err = errNoSuchUser
}
return err
}
func (ies *IAMEtcdStore) deleteGroupInfo(ctx context.Context, name string) error {
err := ies.deleteIAMConfig(ctx, getGroupInfoPath(name))
if err == errConfigNotFound {
err = errNoSuchGroup
}
return err
}
func (ies *IAMEtcdStore) watch(ctx context.Context, keyPath string) <-chan iamWatchEvent {
ch := make(chan iamWatchEvent)
+164
View File
@@ -0,0 +1,164 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"maps"
"slices"
"time"
"github.com/minio/minio-go/v7/pkg/set"
)
type (
iamGroupGrantsKey struct{}
iamGroupMutationKey struct{}
iamGroupMutation struct {
Members []string
Remove bool
StatusOnly bool
}
)
// Merge the intended mutation with the record read under the distributed
// revision lock, not the older cache used to prepare the request.
func mergeIAMGroupMutation(ctx context.Context, previous GroupInfo, next *GroupInfo) {
op, ok := ctx.Value(iamGroupMutationKey{}).(iamGroupMutation)
if !ok || previous.Deleted || previous.Version == 0 {
return
}
members := set.CreateStringSet(previous.Members...)
grants := maps.Clone(previous.MemberGrants)
if grants == nil {
grants = make(map[string]time.Time)
}
switch {
case op.StatusOnly:
// Only the status changes.
case op.Remove:
for _, member := range op.Members {
members.Remove(member)
delete(grants, member)
}
next.Status = previous.Status
default:
requested := set.CreateStringSet(next.Members...)
for _, member := range op.Members {
if !requested.Contains(member) {
continue
}
at := next.MemberGrants[member]
if at.Before(grants[member]) {
continue
}
members.Add(member)
grants[member] = at
}
next.Status = previous.Status
}
next.Members, next.MemberGrants = members.ToSlice(), grants
slices.Sort(next.Members)
}
// A non-nil map is supplied by the versioned peer envelope, including for
// snapshots. Missing times are unknown, never the snapshot's newer timestamp.
func withIAMGroupGrants(ctx context.Context, grants map[string]time.Time) context.Context {
return context.WithValue(ctx, iamGroupGrantsKey{}, grants)
}
func (c *iamCache) effectiveGroupMembers(gi GroupInfo) []string {
var members []string
for _, member := range gi.Members {
if c.groupMemberAllowed(member, gi.MemberGrants[member], gi.RevokedBefore) {
members = append(members, member)
}
}
return members
}
func (c *iamCache) effectiveUserGroups(user string) []string {
var groups []string
for group := range c.iamUserGroupMemberships[user] {
gi, ok := c.iamGroupsMap[group]
r := c.revisions.get(getGroupInfoPath(group))
if r.RevokedBefore.After(gi.RevokedBefore) {
gi.RevokedBefore = r.RevokedBefore
}
if ok && !r.Deleted && c.groupMemberAllowed(user, gi.MemberGrants[user], gi.RevokedBefore) {
groups = append(groups, group)
}
}
return groups
}
func (c *iamCache) addGroupMembers(ctx context.Context, gi GroupInfo, members []string) (GroupInfo, error) {
grants, versioned := ctx.Value(iamGroupGrantsKey{}).(map[string]time.Time)
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(gi.RevokedBefore) {
gi.RevokedBefore = boundary
}
origin, replicated := iamReplicationTime(ctx)
gi.Members = slices.Clone(gi.Members)
gi.MemberGrants = maps.Clone(gi.MemberGrants)
if gi.MemberGrants == nil {
gi.MemberGrants = make(map[string]time.Time)
}
current := set.CreateStringSet(gi.Members...)
gi.UpdatedAt = UTCNow()
if replicated {
gi.UpdatedAt = origin
}
for _, member := range members {
at := gi.UpdatedAt
r := c.userRevocation(member)
if replicated {
switch {
case versioned:
at = grants[member]
if at.After(origin) {
return gi, errInvalidArgument
}
case !r.RevokedBefore.IsZero() || r.Deleted || !gi.RevokedBefore.IsZero():
// Legacy snapshots cannot prove a post-revocation grant.
continue
case current.Contains(member):
continue
}
if !c.groupMemberAllowed(member, at, gi.RevokedBefore) {
continue
}
} else {
if current.Contains(member) && c.groupMemberAllowed(member, gi.MemberGrants[member], gi.RevokedBefore) {
continue // Editing the group is not reissuing every grant.
}
if !at.After(gi.RevokedBefore) {
at = gi.RevokedBefore.Add(time.Nanosecond)
}
if !at.After(r.RevokedBefore) {
at = r.RevokedBefore.Add(time.Nanosecond)
}
if !at.After(gi.MemberGrants[member]) {
at = gi.MemberGrants[member].Add(time.Nanosecond)
}
if at.After(gi.UpdatedAt) {
gi.UpdatedAt = at
}
}
u, ok := c.iamUsersMap[member]
if !ok {
return gi, errNoSuchUser
}
if u.Credentials.IsTemp() || u.Credentials.IsServiceAccount() {
return gi, errIAMActionNotAllowed
}
if previous := gi.MemberGrants[member]; previous.After(at) {
continue
}
current.Add(member)
gi.MemberGrants[member] = at
}
gi.Members = current.ToSlice()
slices.Sort(gi.Members)
return gi, nil
}
+75
View File
@@ -0,0 +1,75 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"testing"
"time"
)
func TestIAMHealingResumesAfterLeadershipLoss(t *testing.T) {
previous := globalLeaderLock
locks := make(chan LockContext)
globalLeaderLock = &sharedLock{lockContext: locks}
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
c := &SiteReplicationSys{}
go func() { c.startHealRoutine(ctx, nil); close(done) }()
t.Cleanup(func() {
cancel()
select {
case <-done:
case locks <- LockContext{ctx: ctx}:
}
<-done
globalLeaderLock = previous
})
first, loseFirst := context.WithCancel(ctx)
defer loseFirst()
select {
case locks <- LockContext{ctx: first}:
case <-time.After(time.Second):
t.Fatal("healer did not acquire its first leader context")
}
loseFirst() // A transient quorum loss cancels the distributed lease.
select {
case <-done:
t.Fatal("healer permanently exited after temporary leadership loss")
case locks <- LockContext{ctx: ctx}:
case <-time.After(time.Second):
t.Fatal("healer did not wait for reacquired leadership")
}
// Shutdown must also interrupt the wait for leadership after lease loss.
cancel()
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("healer did not stop with its owning context")
}
}
func TestIAMHealingLeadershipWaitCancels(t *testing.T) {
previous := globalLeaderLock
locks := make(chan LockContext)
globalLeaderLock = &sharedLock{lockContext: locks}
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
c := &SiteReplicationSys{}
go func() { c.startHealRoutine(ctx, nil); close(done) }()
cancel()
t.Cleanup(func() {
select {
case <-done:
case locks <- LockContext{ctx: ctx}:
}
<-done
globalLeaderLock = previous
})
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("healer ignored shutdown while waiting for leadership")
}
}
+60
View File
@@ -0,0 +1,60 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
)
// Measures steady-state index traversal, sorting and the capability request.
// The network peer acknowledges real batches but performs no disk I/O; this
// benchmark deliberately does not claim durable catch-up throughput.
func BenchmarkIAMRevisionConvergedHealing(b *testing.B) {
for _, n := range []int{1000, 10000} {
b.Run(fmt.Sprint(n), func(b *testing.B) {
ctx, sys, _ := prepareIAMRevisionFixture(b)
_, err := sys.CreateUser(ctx, "benchmark-sync", madmin.AddOrUpdateUserReq{SecretKey: "valid-sync-password", Status: madmin.AccountEnabled})
mustIAM(b, err)
for i := range n {
at := UTCNow().Add(time.Duration(i) * time.Nanosecond)
data, err := json.Marshal(iamRevision{Deleted: true, UpdatedAt: at, RevokedBefore: at})
mustIAM(b, err)
sys.store.revisionIndex().observe(getUserIdentityPath(fmt.Sprintf("deleted-%06d", i), regUser), data)
}
var puts atomic.Int64
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
w.WriteHeader(http.StatusOK)
return
}
if r.Method == http.MethodPut {
puts.Add(1)
}
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: "benchmark-peer", Digest: "constant"}})
}))
defer server.Close()
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "benchmark-sync", Peers: map[string]madmin.PeerInfo{globalDeploymentID(): {DeploymentID: globalDeploymentID(), Name: "local"}, "remote": {DeploymentID: "remote", Name: "remote", Endpoint: server.URL}}}}
mustIAM(b, c.healIAMDeletions(ctx))
before := puts.Load()
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
mustIAM(b, c.healIAMDeletions(ctx))
}
b.StopTimer()
b.ReportMetric(float64(puts.Load()-before)/float64(b.N), "PUT/op")
if puts.Load() != before {
b.Fatal("steady-state healing replayed acknowledged records")
}
})
}
}
+56 -61
View File
@@ -45,6 +45,7 @@ type IAMObjectStore struct {
sync.RWMutex
*iamCache
index iamRevisionIndex
usersSysType UsersSysType
@@ -52,13 +53,17 @@ type IAMObjectStore struct {
}
func newIAMObjectStore(objAPI ObjectLayer, usersSysType UsersSysType) *IAMObjectStore {
return &IAMObjectStore{
store := &IAMObjectStore{
iamCache: newIamCache(),
objAPI: objAPI,
usersSysType: usersSysType,
}
store.revisions = &store.index
return store
}
func (iamOS *IAMObjectStore) revisionIndex() *iamRevisionIndex { return &iamOS.index }
func (iamOS *IAMObjectStore) rlock() *iamCache {
iamOS.RLock()
return iamOS.iamCache
@@ -87,6 +92,7 @@ func (iamOS *IAMObjectStore) saveIAMConfig(ctx context.Context, item any, objPat
if err != nil {
return err
}
plain := data
if GlobalKMS != nil {
data, err = config.EncryptBytes(GlobalKMS, data, kms.Context{
minioMetaBucket: path.Join(minioMetaBucket, objPath),
@@ -95,7 +101,11 @@ func (iamOS *IAMObjectStore) saveIAMConfig(ctx context.Context, item any, objPat
return err
}
}
return saveConfig(ctx, iamOS.objAPI, objPath, data)
if err := saveConfig(ctx, iamOS.objAPI, objPath, data); err != nil {
return err
}
iamOS.index.observe(objPath, plain)
return nil
}
func decryptData(data []byte, objPath string) ([]byte, error) {
@@ -133,6 +143,7 @@ func (iamOS *IAMObjectStore) loadIAMConfigBytesWithMetadata(ctx context.Context,
if err != nil {
return nil, meta, err
}
iamOS.index.observe(objPath, data)
return data, meta, nil
}
@@ -146,7 +157,11 @@ func (iamOS *IAMObjectStore) loadIAMConfig(ctx context.Context, item any, objPat
}
func (iamOS *IAMObjectStore) deleteIAMConfig(ctx context.Context, path string) error {
return deleteConfig(ctx, iamOS.objAPI, path)
if err := deleteConfig(ctx, iamOS.objAPI, path); err != nil {
return err
}
iamOS.index.forget(path)
return nil
}
func (iamOS *IAMObjectStore) loadPolicyDocWithRetry(ctx context.Context, policy string, m map[string]PolicyDoc, retries int) error {
@@ -171,6 +186,10 @@ func (iamOS *IAMObjectStore) loadPolicyDocWithRetry(ctx context.Context, policy
return err
}
if p.Deleted {
return errNoSuchPolicy
}
if p.Version == 0 {
// This means that policy was in the old version (without any
// timestamp info). We fetch the mod time of the file and save
@@ -200,6 +219,10 @@ func (iamOS *IAMObjectStore) loadPolicy(ctx context.Context, policy string) (Pol
return p, err
}
if p.Deleted {
return PolicyDoc{}, errNoSuchPolicy
}
if p.Version == 0 {
// This means that policy was in the old version (without any
// timestamp info). We fetch the mod time of the file and save
@@ -245,6 +268,9 @@ func (iamOS *IAMObjectStore) loadSecretKey(ctx context.Context, user string, use
}
return "", err
}
if u.Deleted {
return "", errNoSuchUser
}
return u.Credentials.SecretKey, nil
}
@@ -258,10 +284,15 @@ func (iamOS *IAMObjectStore) loadUserIdentity(ctx context.Context, user string,
return u, err
}
if u.Deleted {
if userType == stsUser && !u.ExpiresAt.IsZero() && UTCNow().After(u.ExpiresAt) {
bestEffortIAMExpiration(ctx, iamOS, getUserIdentityPath(user, userType))
}
return UserIdentity{}, errNoSuchUser
}
if u.Credentials.IsExpired() {
// Delete expired identity - ignoring errors here.
iamOS.deleteIAMConfig(ctx, getUserIdentityPath(user, userType))
iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(user, userType, false))
bestEffortIAMExpiration(ctx, iamOS, getUserIdentityPath(user, userType))
return u, errNoSuchUser
}
@@ -272,16 +303,18 @@ func (iamOS *IAMObjectStore) loadUserIdentity(ctx context.Context, user string,
if u.Credentials.SessionToken != "" {
jwtClaims, err := extractJWTClaims(u)
if err != nil {
if u.Credentials.IsTemp() {
// We should delete such that the client can re-request
// for the expiring credentials.
iamOS.deleteIAMConfig(ctx, getUserIdentityPath(user, userType))
iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(user, userType, false))
}
return u, errNoSuchUser
// During startup the site signing key may not be available yet.
// Reject this load without deleting a credential that has not expired.
return UserIdentity{}, errNoSuchUser
}
u.Credentials.Claims = jwtClaims.Map()
}
if err := checkIAMParentRevision(ctx, iamOS, u.Credentials); err != nil {
if errors.Is(err, errIAMStaleUpdate) {
return UserIdentity{}, errNoSuchUser
}
return UserIdentity{}, err
}
if u.Credentials.Description == "" {
u.Credentials.Description = u.Credentials.Comment
@@ -320,6 +353,7 @@ func (iamOS *IAMObjectStore) loadUser(ctx context.Context, user string, userType
}
func (iamOS *IAMObjectStore) loadUsers(ctx context.Context, userType IAMUserType, m map[string]UserIdentity) error {
ctx = withIAMExpirationCleanup(ctx)
var basePrefix string
switch userType {
case svcUser:
@@ -354,6 +388,9 @@ func (iamOS *IAMObjectStore) loadGroup(ctx context.Context, group string, m map[
}
return err
}
if g.Deleted {
return errNoSuchGroup
}
m[group] = g
return nil
}
@@ -391,6 +428,9 @@ func (iamOS *IAMObjectStore) loadMappedPolicyWithRetry(ctx context.Context, name
goto retry
}
if !iamOS.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
@@ -405,6 +445,9 @@ func (iamOS *IAMObjectStore) loadMappedPolicyInternal(ctx context.Context, name
}
return p, err
}
if !iamOS.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return MappedPolicy{}, errNoSuchPolicy
}
return p, nil
}
@@ -824,54 +867,6 @@ func (iamOS *IAMObjectStore) loadAllFromObjStore(ctx context.Context, cache *iam
return nil
}
func (iamOS *IAMObjectStore) savePolicyDoc(ctx context.Context, policyName string, p PolicyDoc) error {
return iamOS.saveIAMConfig(ctx, &p, getPolicyDocPath(policyName))
}
func (iamOS *IAMObjectStore) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp MappedPolicy, opts ...options) error {
return iamOS.saveIAMConfig(ctx, mp, getMappedPolicyPath(name, userType, isGroup), opts...)
}
func (iamOS *IAMObjectStore) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u UserIdentity, opts ...options) error {
return iamOS.saveIAMConfig(ctx, u, getUserIdentityPath(name, userType), opts...)
}
func (iamOS *IAMObjectStore) saveGroupInfo(ctx context.Context, name string, gi GroupInfo) error {
return iamOS.saveIAMConfig(ctx, gi, getGroupInfoPath(name))
}
func (iamOS *IAMObjectStore) deletePolicyDoc(ctx context.Context, name string) error {
err := iamOS.deleteIAMConfig(ctx, getPolicyDocPath(name))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (iamOS *IAMObjectStore) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
err := iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(name, userType, isGroup))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (iamOS *IAMObjectStore) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
err := iamOS.deleteIAMConfig(ctx, getUserIdentityPath(name, userType))
if err == errConfigNotFound {
err = errNoSuchUser
}
return err
}
func (iamOS *IAMObjectStore) deleteGroupInfo(ctx context.Context, name string) error {
err := iamOS.deleteIAMConfig(ctx, getGroupInfoPath(name))
if err == errConfigNotFound {
err = errNoSuchGroup
}
return err
}
// Lists objects in the minioMetaBucket at the given path prefix. All returned
// items have the pathPrefix removed from their names.
func listIAMConfigItems(ctx context.Context, objAPI ObjectLayer, pathPrefix string) <-chan itemOrErr[string] {
+133
View File
@@ -0,0 +1,133 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/grid"
"github.com/pgsty/silo-pkg/v3/policy"
)
// Two independent IAM caches share the same real object backend, as sibling
// nodes do. Deliver the actual peer handler only after the source committed.
func TestIAMPeerDeleteNotificationReloadsCommittedState(t *testing.T) {
for _, name := range []string{"deleted", "recreated", "recreated_without_grant"} {
recreate := name != "deleted"
t.Run(name, func(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
globalObjLayerMutex.Lock()
globalObjectAPI = obj
globalObjLayerMutex.Unlock()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
source := globalIAMSys
const user = "peer-reload-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "original-test-password", Status: madmin.AccountEnabled}
_, err = source.CreateUser(ctx, user, req)
must(err)
_, err = source.PolicyDBSet(ctx, user, "readwrite", regUser, false)
must(err)
_, err = source.AddUsersToGroup(ctx, "peer-reload-group", []string{user})
must(err)
_, err = source.PolicyDBSet(ctx, "peer-reload-group", "readwrite", regUser, true)
must(err)
svc, _, err := source.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{accessKey: "peer-reload-service", secretKey: "service-test-password"})
must(err)
signingKey, err := getTokenSigningKey()
must(err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: user}, signingKey)
must(err)
sts.ParentUser = user
_, err = source.SetTempUser(ctx, sts.AccessKey, sts, "")
must(err)
siblingStore := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
must(siblingStore.LoadIAMCache(ctx, true))
must(siblingStore.UserNotificationHandler(ctx, sts.AccessKey, stsUser))
for _, key := range []string{user, svc.AccessKey, sts.AccessKey} {
if _, ok := siblingStore.GetUser(key); !ok {
t.Fatalf("fixture did not load %s", key)
}
}
must(source.DeleteUser(ctx, user, false))
if recreate {
req.SecretKey = "recreated-test-password"
_, err = source.CreateUser(ctx, user, req)
must(err)
if name == "recreated" {
_, err = source.PolicyDBSet(ctx, user, "readonly", regUser, false)
must(err)
}
}
sibling := &IAMSys{store: siblingStore, usersSysType: MinIOUsersSysType}
globalIAMSys = sibling
defer func() { globalIAMSys = source }()
server := &peerRESTServer{}
for range 2 {
_, remoteErr := server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
}
if recreate {
u, ok := siblingStore.GetUser(user)
if !ok || u.Credentials.SecretKey != req.SecretKey {
t.Fatal("delayed deletion notification removed the recreated user")
}
loaded := make(map[string]UserIdentity)
must(source.store.loadUser(ctx, user, regUser, loaded))
if loaded[user].Credentials.SecretKey != req.SecretKey {
t.Fatal("notification changed the persisted recreated identity")
}
if allowed := sibling.IsAllowed(policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}); allowed != (name == "recreated") {
t.Fatal("notification did not load the recreated user's current grant")
}
}
if sibling.IsAllowed(policy.Args{AccountName: user, Action: policy.PutObjectAction, BucketName: "bucket", ObjectName: "object"}) {
t.Fatal("notification retained an old direct or group grant")
}
for _, key := range []string{svc.AccessKey, sts.AccessKey} {
if _, ok := siblingStore.GetUser(key); ok {
t.Fatalf("notification retained a revoked child: %s", key)
}
}
if !recreate {
for _, key := range []string{user} {
if _, ok := siblingStore.GetUser(key); ok {
t.Fatalf("notification retained a revoked cached identity: %s", key)
}
}
if sibling.IsAllowed(policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}) {
t.Fatal("notification retained the user's old grant")
}
cache := siblingStore.rlock()
member := cache.iamUserGroupMemberships[user].Contains("peer-reload-group")
siblingStore.runlock()
if member {
t.Fatal("notification retained the deleted user's group membership")
}
}
})
}
}
+131
View File
@@ -0,0 +1,131 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"fmt"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
)
// Uses APIs shared with the pre-revision tree so the same benchmark can be
// overlaid on that tree for a comparable local baseline.
func prepareIAMPerformanceFixture(b *testing.B) (context.Context, *IAMSys) {
b.Helper()
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
disks, err := getRandomDisks(1)
if err != nil {
b.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
b.Fatal(err)
}
initAllSubsystems(ctx)
globalIAMSys.initStore(obj, nil)
if err := globalIAMSys.Load(ctx, true); err != nil {
b.Fatal(err)
}
b.Cleanup(func() { cancel(); obj.Shutdown(context.Background()); os.RemoveAll(disks[0]); resetTestGlobals() })
return ctx, globalIAMSys
}
func BenchmarkIAMCachedCredential(b *testing.B) {
for _, kind := range []string{"user", "service", "sts"} {
b.Run(kind, func(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
const parent = "benchmark-parent"
_, err := sys.CreateUser(ctx, parent, madmin.AddOrUpdateUserReq{SecretKey: "benchmark-user-password", Status: madmin.AccountEnabled})
if err != nil {
b.Fatal(err)
}
key := parent
if kind == "service" {
c, _, err := sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "benchmark-service", secretKey: "benchmark-service-password"})
if err != nil {
b.Fatal(err)
}
key = c.AccessKey
}
if kind == "sts" {
secret, err := getTokenSigningKey()
if err != nil {
b.Fatal(err)
}
c, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
if err != nil {
b.Fatal(err)
}
c.ParentUser = parent
if _, err := sys.SetTempUser(ctx, c.AccessKey, c, ""); err != nil {
b.Fatal(err)
}
key = c.AccessKey
}
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
if _, ok := sys.store.GetUser(key); !ok {
b.Fatal("credential missing")
}
}
})
}
}
func BenchmarkIAMSetTempUser(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
const parent = "benchmark-sts-parent"
_, err := sys.CreateUser(ctx, parent, madmin.AddOrUpdateUserReq{SecretKey: "benchmark-user-password", Status: madmin.AccountEnabled})
if err != nil {
b.Fatal(err)
}
secret, err := getTokenSigningKey()
if err != nil {
b.Fatal(err)
}
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
if err != nil {
b.Fatal(err)
}
cred.ParentUser = parent
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
if _, err := sys.SetTempUser(ctx, cred.AccessKey, cred, "readwrite"); err != nil {
b.Fatal(err)
}
}
}
// Run with -benchtime=1x. Preparation is outside the timer; each measured load
// sees a fresh set of expired reusable service-account records.
func BenchmarkIAMColdLoadExpiredServices(b *testing.B) {
for _, count := range []int{100, 1000} {
b.Run(fmt.Sprint(count), func(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
b.ReportAllocs()
for i := 0; i < b.N; i++ {
b.StopTimer()
for j := 0; j < count; j++ {
key := fmt.Sprintf("expired-benchmark-%d-%d", i, j)
u := UserIdentity{Version: 1, UpdatedAt: UTCNow().Add(-2 * time.Hour), Credentials: auth.Credentials{AccessKey: key, SecretKey: "expired-benchmark-password", ParentUser: "absent-idp-parent", Expiration: UTCNow().Add(-time.Hour), Status: auth.AccountOn}}
if err := sys.store.saveIAMConfig(ctx, &u, getUserIdentityPath(key, svcUser)); err != nil {
b.Fatal(err)
}
}
b.StartTimer()
if err := sys.store.LoadIAMCache(ctx, true); err != nil {
b.Fatal(err)
}
}
})
}
}
+168
View File
@@ -0,0 +1,168 @@
package cmd
import (
"context"
"errors"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/pgsty/silo-pkg/v3/policy"
)
func TestReviewIAMRevokedUserReplay(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "review-revoked-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "review-valid-password", Status: madmin.AccountEnabled}
created, err := globalIAMSys.CreateUser(ctx, user, req)
if err != nil {
t.Fatal(err)
}
policyAt, err := globalIAMSys.PolicyDBSet(ctx, user, "readwrite", regUser, false)
if err != nil {
t.Fatal(err)
}
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "review-bucket", ObjectName: "review-object"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("seed must allow object read")
}
if err := globalIAMSys.DeleteUser(ctx, user, false); err != nil {
t.Fatal(err)
}
if err := globalIAMSys.store.LoadIAMCache(ctx, false); err != nil {
t.Fatal(err)
}
if globalIAMSys.IsAllowed(args) {
t.Fatal("deletion did not remove initial permission")
}
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, created); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Errorf("revoked user restored by an older replicated create, GetUserInfo error = %v", err)
}
if err := globalSiteReplicationSys.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: user, UserType: int(regUser), Policy: "readwrite"}, policyAt); err != nil {
t.Fatal(err)
}
if globalIAMSys.IsAllowed(args) {
t.Error("older replicated identity and policy events restored revoked S3 read permission")
}
}
func TestReviewIAMSourceTimestampOrder(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "review-ordered-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "review-valid-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour)
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin); err != nil {
t.Fatal(err)
}
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(time.Minute)); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Fatalf("newer source deletion skipped after delayed creation, GetUserInfo error = %v", err)
}
}
// A user's old group grant must not return after deletion and deliberate recreation.
func TestR3CandidateOldGroupReplayAfterRecreation(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user, group := "r3-group-member", "r3-granting-group"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-r3-user-password", Status: madmin.AccountEnabled}
peer := &globalSiteReplicationSys
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin))
add := &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}
must(peer.PeerGroupInfoChangeHandler(ctx, add, origin.Add(time.Minute)))
must(peer.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: group, IsGroup: true, UserType: int(regUser), Policy: "readwrite"}, origin.Add(time.Minute)))
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "r3-bucket", ObjectName: "probe"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("fixture must grant through group")
}
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin.Add(3*time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if globalIAMSys.IsAllowed(args) {
t.Fatal("recreation must start without deleted group membership")
}
must(peer.PeerGroupInfoChangeHandler(ctx, add, origin.Add(time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if globalIAMSys.IsAllowed(args) {
t.Fatal("old group event restored the deleted user's read grant after recreation and durable reload")
}
}
// The user delete is also a revocation of its earlier group memberships.
func TestR3CandidateLateDeleteRetainsOldGroupGrant(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user, group := "r3-late-member", "r3-late-group"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-r3-user-password", Status: madmin.AccountEnabled}
peer := &globalSiteReplicationSys
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin))
must(peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}, origin.Add(time.Minute)))
must(peer.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: group, IsGroup: true, UserType: int(regUser), Policy: "readwrite"}, origin.Add(time.Minute)))
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "r3-bucket", ObjectName: "probe"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("fixture must grant through group")
}
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin.Add(3*time.Minute)))
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if _, ok := globalIAMSys.GetUser(ctx, user); !ok {
t.Fatal("newer identity must survive")
}
if globalIAMSys.IsAllowed(args) {
t.Fatal("late user deletion retained the older group grant on the recreated identity")
}
}
+321
View File
@@ -0,0 +1,321 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"sort"
"strings"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
xhttp "github.com/minio/minio/internal/http"
"github.com/pgsty/silo-pkg/v3/policy"
)
const (
iamRevisionProtocol = 1
iamRevisionPeerPath = "/v3/site-replication/peer/iam-revisions"
iamUserBoundaryType = "silo-user-revocation"
iamGroupBoundaryType = "silo-group-revocation"
maxIAMRevisionBatch = 128
)
var iamRevisionInstance = mustGetUUID()
type iamUserBoundary struct {
User string `json:"user"`
Before time.Time `json:"before"`
}
type iamGroupBoundary struct {
Group string `json:"group"`
Before time.Time `json:"before"`
}
// The server owns this additive protocol, without changing the client SDK or
// overloading a policy/document field. Old servers reject the dedicated route
// before applying any change that would lose revocation or member metadata.
type iamReplicationItem struct {
madmin.SRIAMItem
GroupGrants map[string]time.Time `json:"groupGrants,omitempty"`
GroupSnapshot bool `json:"groupSnapshot,omitempty"`
UserRevocation *iamUserBoundary `json:"userRevocation,omitempty"`
GroupRevocation *iamGroupBoundary `json:"groupRevocation,omitempty"`
RevokedBefore time.Time `json:"revokedBefore,omitempty"`
}
type iamRevisionBatch struct {
Version int `json:"version"`
Items []iamReplicationItem `json:"items"`
}
type iamRevisionStatus struct {
Version int `json:"version"`
Node string `json:"node"`
Instance string `json:"instance"`
Digest string `json:"digest"`
}
type iamRevisionResponse struct {
iamRevisionStatus
Errors []string `json:"errors,omitempty"`
}
type iamRevisionBatchError struct{ failures []string }
func (e *iamRevisionBatchError) Error() string {
return "IAM revision batch: " + strings.Join(e.failures, "; ")
}
type iamRevisionProgress struct {
Instances map[string]string
Acknowledged map[string]string
}
type iamRevisionMetrics struct {
healFailures atomic.Uint64
healLastSuccess atomic.Int64
healDurationMillis atomic.Int64
}
func iamRevisionDigest(items map[string]iamRevision) string {
paths := make([]string, 0, len(items))
for path := range items {
paths = append(paths, path)
}
sort.Strings(paths)
h := sha256.New()
for _, path := range paths {
r := items[path]
fmt.Fprintf(h, "%q %s %t %s\n", path, r.timestamp().UTC().Format(time.RFC3339Nano), r.Deleted, r.RevokedBefore.UTC().Format(time.RFC3339Nano))
}
return hex.EncodeToString(h.Sum(nil))
}
func (store *IAMStoreSys) iamRevisionStatus() iamRevisionStatus {
node := globalLocalNodeName
if node == "" {
node = "local"
}
return iamRevisionStatus{Version: iamRevisionProtocol, Node: node, Instance: iamRevisionInstance, Digest: store.revisionIndex().digest()}
}
func executeIAMRevisionRequest(ctx context.Context, client *madmin.AdminClient, method string, batch *iamRevisionBatch) (status iamRevisionStatus, err error) {
var content []byte
if batch != nil {
content, err = json.Marshal(batch)
if err != nil {
return status, err
}
}
resp, err := client.ExecuteMethod(ctx, method, madmin.RequestData{RelPath: iamRevisionPeerPath, QueryValues: url.Values{"api-version": {madmin.SiteReplAPIVersion}}, Content: content})
if resp != nil {
defer xhttp.DrainBody(resp.Body)
}
if err != nil {
return status, err
}
if resp.StatusCode != http.StatusOK {
var remote madmin.ErrorResponse
if json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&remote) == nil && remote.Code != "" {
return status, remote
}
return status, fmt.Errorf("IAM revision protocol requires upgraded peers: %s", resp.Status)
}
var response iamRevisionResponse
if err = json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&response); err != nil {
return status, err
}
status = response.iamRevisionStatus
if status.Version != iamRevisionProtocol || status.Node == "" || status.Instance == "" || status.Digest == "" {
return status, errors.New("peer did not acknowledge the IAM revision protocol")
}
if len(response.Errors) != 0 {
return status, &iamRevisionBatchError{failures: response.Errors}
}
return status, nil
}
type (
iamRecordBoundaryKey struct{}
iamGroupSnapshotKey struct{}
)
func (c *SiteReplicationSys) replicationItem(ctx context.Context, item madmin.SRIAMItem) (iamReplicationItem, error) {
out := iamReplicationItem{SRIAMItem: item}
if item.Type == madmin.SRIAMItemSvcAcc && item.SvcAccChange != nil {
var key string
if item.SvcAccChange.Create != nil {
key = item.SvcAccChange.Create.AccessKey
} else if item.SvcAccChange.Update != nil {
key = item.SvcAccChange.Update.AccessKey
}
if key != "" {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getUserIdentityPath(key, svcUser))
if err != nil {
return out, err
}
if r.Deleted || r.timestamp().After(item.UpdatedAt) {
return out, errIAMStaleUpdate
}
out.RevokedBefore = r.RevokedBefore
}
}
if item.Type == madmin.SRIAMItemGroupInfo && item.GroupInfo != nil && !item.GroupInfo.UpdateReq.IsRemove {
out.GroupSnapshot = true
var gi GroupInfo
if err := globalIAMSys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath(item.GroupInfo.UpdateReq.Group)); err != nil {
return out, err
}
// The matching persisted snapshot carries member grant times. If a
// later write won before sending, propagate that whole newer state.
if gi.Deleted {
return out, errIAMStaleUpdate
}
out.UpdatedAt = gi.UpdatedAt
out.RevokedBefore = gi.RevokedBefore
out.GroupInfo = &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: item.GroupInfo.UpdateReq.Group, Status: madmin.GroupStatus(gi.Status)}}
cache := globalIAMSys.store.rlock()
out.GroupInfo.UpdateReq.Members = cache.effectiveGroupMembers(gi)
globalIAMSys.store.runlock()
out.GroupGrants = make(map[string]time.Time, len(out.GroupInfo.UpdateReq.Members))
for _, member := range out.GroupInfo.UpdateReq.Members {
out.GroupGrants[member] = gi.MemberGrants[member]
}
}
if item.Type == madmin.SRIAMItemGroupInfo && item.GroupInfo != nil && item.GroupInfo.UpdateReq.IsRemove && len(item.GroupInfo.UpdateReq.Members) == 0 {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getGroupInfoPath(item.GroupInfo.UpdateReq.Group))
if err != nil {
return out, err
}
if !r.Deleted && !r.RevokedBefore.IsZero() {
out.Type, out.GroupInfo = iamGroupBoundaryType, nil
out.GroupRevocation = &iamGroupBoundary{Group: item.GroupInfo.UpdateReq.Group, Before: r.RevokedBefore}
out.UpdatedAt = r.RevokedBefore
}
}
if item.Type == madmin.SRIAMItemIAMUser && item.IAMUser != nil {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getUserIdentityPath(item.IAMUser.AccessKey, regUser))
if err != nil {
return out, err
}
if item.IAMUser.IsDeleteReq && !r.Deleted && !r.RevokedBefore.IsZero() {
out.Type = iamUserBoundaryType
out.IAMUser = nil
out.UserRevocation = &iamUserBoundary{User: item.IAMUser.AccessKey, Before: r.RevokedBefore}
out.UpdatedAt = r.RevokedBefore
} else if !item.IAMUser.IsDeleteReq {
if r.Deleted || r.timestamp().After(item.UpdatedAt) {
return out, errIAMStaleUpdate
}
out.RevokedBefore = r.RevokedBefore
}
}
return out, nil
}
func applyIAMReplicationItem(ctx context.Context, item iamReplicationItem) error {
if item.GroupInfo != nil {
if item.GroupSnapshot {
ctx = context.WithValue(ctx, iamGroupSnapshotKey{}, true)
}
// A nil map also explicitly denotes unknown legacy grants. Do not
// turn an unrelated group edit into a new grant after a revocation.
ctx = withIAMGroupGrants(ctx, item.GroupGrants)
}
if !item.RevokedBefore.IsZero() {
if item.RevokedBefore.After(item.UpdatedAt) {
return errSRInvalidRequest(errInvalidArgument)
}
ctx = context.WithValue(ctx, iamRecordBoundaryKey{}, item.RevokedBefore)
}
switch item.Type {
case iamUserBoundaryType:
if item.UserRevocation == nil || item.UserRevocation.User == "" || item.UserRevocation.Before.IsZero() {
return errSRInvalidRequest(errInvalidArgument)
}
return iamReplicationError(globalIAMSys.DeleteUser(withIAMReplicationTime(ctx, item.UserRevocation.Before), item.UserRevocation.User, true))
case iamGroupBoundaryType:
if item.GroupRevocation == nil || item.GroupRevocation.Group == "" || item.GroupRevocation.Before.IsZero() {
return errSRInvalidRequest(errInvalidArgument)
}
_, err := globalIAMSys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, item.GroupRevocation.Before), item.GroupRevocation.Group, nil)
return iamReplicationError(err)
case madmin.SRIAMItemPolicy:
if len(item.Policy) == 0 {
return globalSiteReplicationSys.PeerAddPolicyHandler(ctx, item.Name, nil, item.UpdatedAt)
}
p, err := policy.ParseConfig(bytes.NewReader(item.Policy))
if err != nil {
return err
}
if p.IsEmpty() {
p = nil
}
return globalSiteReplicationSys.PeerAddPolicyHandler(ctx, item.Name, p, item.UpdatedAt)
case madmin.SRIAMItemSvcAcc:
return globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, item.SvcAccChange, item.UpdatedAt)
case madmin.SRIAMItemPolicyMapping:
return globalSiteReplicationSys.PeerPolicyMappingHandler(ctx, item.PolicyMapping, item.UpdatedAt)
case madmin.SRIAMItemSTSAcc:
return globalSiteReplicationSys.PeerSTSAccHandler(ctx, item.STSCredential, item.UpdatedAt)
case madmin.SRIAMItemIAMUser:
return globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, item.IAMUser, item.UpdatedAt)
case madmin.SRIAMItemGroupInfo:
return globalSiteReplicationSys.PeerGroupInfoChangeHandler(ctx, item.GroupInfo, item.UpdatedAt)
default:
return errSRInvalidRequest(errInvalidArgument)
}
}
func (a adminAPIHandlers) SRPeerIAMRevisions(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
if obj, _ := validateAdminReq(ctx, w, r, policy.SiteReplicationOperationAction); obj == nil {
return
}
var failures []string
if r.Method == http.MethodPut {
var batch iamRevisionBatch
if err := parseJSONBody(ctx, r.Body, &batch, ""); err != nil {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, err), r.URL)
return
}
if batch.Version != iamRevisionProtocol || len(batch.Items) == 0 || len(batch.Items) > maxIAMRevisionBatch {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, errSRInvalidRequest(errInvalidArgument)), r.URL)
return
}
for i, item := range batch.Items {
if err := applyIAMReplicationItem(ctx, item); err != nil {
failures = append(failures, fmt.Sprintf("item %d (%s): %v", i, item.Type, err))
}
}
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: globalIAMSys.store.iamRevisionStatus(), Errors: failures})
}
// A site endpoint can balance requests across nodes sharing durable IAM state.
// Switching between known node incarnations preserves ACKs; a new incarnation
// conservatively invalidates them so restoring an old backend cannot inherit
// acknowledgements from before the restore.
func (p *iamRevisionProgress) observePeer(status iamRevisionStatus) {
if p.Instances == nil {
p.Instances = make(map[string]string)
}
if p.Instances[status.Node] != status.Instance || p.Acknowledged == nil {
p.Instances[status.Node] = status.Instance
p.Acknowledged = make(map[string]string)
}
}
+161
View File
@@ -0,0 +1,161 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
)
func TestIAMRevisionProtocolDoesNotFallBackToLegacy(t *testing.T) {
var requests atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requests.Add(1)
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" {
t.Errorf("unsafe fallback path: %s", r.URL.Path)
}
w.WriteHeader(http.StatusNotFound)
_, _ = w.Write([]byte(`{"Code":"NotImplemented","Message":"old server"}`))
}))
defer server.Close()
client, err := madmin.New(strings.TrimPrefix(server.URL, "http://"), "test-access", "valid-test-secret", false)
mustIAM(t, err)
_, err = executeIAMRevisionRequest(context.Background(), client, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: []iamReplicationItem{{SRIAMItem: madmin.SRIAMItem{Type: iamUserBoundaryType}, UserRevocation: &iamUserBoundary{User: "recreated", Before: UTCNow()}}}})
if err == nil || requests.Load() != 1 {
t.Fatalf("old peer must reject without fallback, err=%v requests=%d", err, requests.Load())
}
}
type iamNoHealingScanStore struct{ IAMStorageAPI }
func (s *iamNoHealingScanStore) listIAMConfigPaths(context.Context) ([]string, error) {
panic("healing must use the loaded revision index")
}
func TestIAMRevisionHealingAcknowledgements(t *testing.T) {
for _, balanced := range []bool{false, true} {
t.Run(fmt.Sprintf("load_balanced_%t", balanced), func(t *testing.T) { testIAMRevisionHealingAcknowledgements(t, balanced) })
}
}
func testIAMRevisionHealingAcknowledgements(t *testing.T, balanced bool) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "ack-sync", madmin.AddOrUpdateUserReq{SecretKey: "valid-sync-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
for i := range maxIAMRevisionBatch*2 + 1 {
at := UTCNow().Add(time.Duration(i) * time.Nanosecond)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Deleted: true, UpdatedAt: at, RevokedBefore: at}, getUserIdentityPath(fmt.Sprintf("ack-%04d", i), regUser)))
}
sys.store.IAMStorageAPI = &iamNoHealingScanStore{IAMStorageAPI: sys.store.IAMStorageAPI}
var mu sync.Mutex
var puts, gets int
var applied int
instance := "boot-1"
failSecondBatch := true
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
w.WriteHeader(http.StatusOK)
return
}
mu.Lock()
defer mu.Unlock()
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" {
t.Errorf("unexpected request: %s", r.URL.Path)
w.WriteHeader(404)
return
}
var failures []string
if r.Method == http.MethodGet {
gets++
} else {
puts++
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(400)
return
}
if len(batch.Items) > maxIAMRevisionBatch {
t.Error("batch exceeds limit")
}
if failSecondBatch && puts == 2 {
failures = []string{"injected item error"}
} else {
applied += len(batch.Items)
}
}
node := "node-1"
if balanced {
node = fmt.Sprintf("node-%d", (gets+puts)%2+1)
}
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: node, Instance: node + instance, Digest: fmt.Sprintf("%d", applied)}, Errors: failures})
}))
defer server.Close()
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "ack-sync", Peers: map[string]madmin.PeerInfo{globalDeploymentID(): {DeploymentID: globalDeploymentID(), Name: "local"}, "remote": {DeploymentID: "remote", Name: "remote", Endpoint: server.URL}}}}
if err := c.healIAMDeletions(ctx); err == nil {
t.Fatal("item failure was hidden")
}
mu.Lock()
if puts != 3 || applied != maxIAMRevisionBatch+1 {
t.Errorf("failed middle batch blocked later revocations: puts=%d applied=%d", puts, applied)
}
failSecondBatch = false
applied++ // Unrelated remote mutation changes its digest.
mu.Unlock()
at := UTCNow()
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Deleted: true, UpdatedAt: at, RevokedBefore: at}, getUserIdentityPath("ack-new-local", regUser)))
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
if puts != 5 {
t.Errorf("did not resume at unacknowledged batch: puts=%d", puts)
}
mu.Unlock()
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
if puts != 5 || gets != 3 {
t.Errorf("converged records were replayed: puts=%d gets=%d", puts, gets)
}
instance = "boot-2"
mu.Unlock()
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
defer mu.Unlock()
if puts != 8 {
t.Fatalf("peer restart reused an old acknowledgement: puts=%d", puts)
}
}
func TestIAMRevisionIndexRebuildsFromStorage(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
const user = "index-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-parent-password", Status: madmin.AccountEnabled}
_, err := sys.CreateUser(ctx, user, req)
mustIAM(t, err)
mustIAM(t, sys.DeleteUser(ctx, user, false))
before := sys.store.revisionIndex().snapshot()
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, store.LoadIAMCache(ctx, true))
if iamRevisionDigest(before) != iamRevisionDigest(store.revisionIndex().snapshot()) {
t.Fatal("ordinary IAM loading did not restore the deletion index")
}
_, err = store.AddUser(ctx, user, req)
mustIAM(t, err)
r := store.revisionIndex().get(getUserIdentityPath(user, regUser))
if r.Deleted || r.RevokedBefore.IsZero() {
t.Fatal("recreation discarded the retained boundary")
}
if r.Credentials.SecretKey != "" || r.Credentials.SessionToken != "" {
t.Fatal("index retained credentials")
}
}
+393
View File
@@ -0,0 +1,393 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"errors"
"fmt"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
"github.com/pgsty/silo-pkg/v3/policy"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/namespace"
)
func prepareIAMRevisionFixture(t testing.TB, backend ...string) (context.Context, *IAMSys, ObjectLayer) {
t.Helper()
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
disks, err := getRandomDisks(1)
mustIAM(t, err)
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
mustIAM(t, err)
initAllSubsystems(ctx)
// Deliberately omit the periodic refresh goroutine. Fault injection can
// replace this fixture's storage interface without racing initialization.
var client *etcd.Client
if len(backend) != 0 && backend[0] == "etcd" {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if endpoint == "" {
cancel()
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
client, err = etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: 5 * time.Second})
mustIAM(t, err)
prefix := fmt.Sprintf("/silo-boundary-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(client.KV, prefix)
client.Watcher = namespace.NewWatcher(client.Watcher, prefix)
t.Cleanup(func() { client.Delete(context.Background(), "", etcd.WithPrefix()); client.Close() })
}
globalIAMSys.initStore(obj, client)
mustIAM(t, globalIAMSys.Load(ctx, true))
t.Cleanup(func() { cancel(); obj.Shutdown(context.Background()); os.RemoveAll(disks[0]); resetTestGlobals() })
return ctx, globalIAMSys, obj
}
func mustIAM(t testing.TB, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
var errIAMInjectedWrite = errors.New("injected IAM persistence failure")
type iamFailingCleanupStore struct {
IAMStorageAPI
parentPath string
beforeCommit bool
}
func (s *iamFailingCleanupStore) saveIAMConfig(ctx context.Context, item any, path string, opts ...options) error {
if s.beforeCommit || path != s.parentPath {
return errIAMInjectedWrite
}
return s.IAMStorageAPI.saveIAMConfig(ctx, item, path, opts...)
}
func TestIAMRevocationCommitBoundary(t *testing.T) {
for _, before := range []bool{true, false} {
name := "after_identity_commit"
if before {
name = "before_identity_commit"
}
t.Run(name, func(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
const user = "commit-boundary-user"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, req)
mustIAM(t, err)
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin.Add(time.Minute)), user, "readwrite", regUser, false)
mustIAM(t, err)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(time.Minute)), "commit-group", []string{user})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "commit-group", "readwrite", regUser, true)
mustIAM(t, err)
child, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), user, nil, newServiceAccountOpts{accessKey: "commit-child", secretKey: "valid-child-password"})
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("fixture has no grant")
}
siblingStore := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, siblingStore.LoadIAMCache(ctx, true))
sibling := &IAMSys{store: siblingStore, usersSysType: MinIOUsersSysType}
tg, err := grid.SetupTestGrid(2)
mustIAM(t, err)
defer tg.Cleanup()
var notifications atomic.Int32
mustIAM(t, deleteUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
notifications.Add(1)
if err := sibling.LoadUserAfterDelete(ctx, r.Get(peerRESTUser)); err != nil {
return grid.NoPayload{}, grid.NewRemoteErr(err)
}
return grid.NoPayload{}, nil
}))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
mustIAM(t, err)
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{host: host, gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) }}}}
original := sys.store.IAMStorageAPI
sys.store.IAMStorageAPI = &iamFailingCleanupStore{IAMStorageAPI: original, parentPath: getUserIdentityPath(user, regUser), beforeCommit: before}
boundary := origin.Add(2 * time.Minute)
err = sys.DeleteUser(withIAMReplicationTime(ctx, boundary), user, true)
if !errors.Is(err, errIAMInjectedWrite) {
t.Fatalf("expected write failure, got %v", err)
}
sys.store.IAMStorageAPI = original
r, err := loadIAMRevision(ctx, original, getUserIdentityPath(user, regUser))
mustIAM(t, err)
if before {
if r.Deleted || !sys.IsAllowed(args) || !sibling.IsAllowed(args) || notifications.Load() != 0 {
t.Fatal("failure before commit changed the identity or grant")
}
return
}
if !r.Deleted || !r.RevokedBefore.Equal(boundary) {
t.Fatal("cleanup failure lost durable revocation")
}
if sys.IsAllowed(args) || sibling.IsAllowed(args) || notifications.Load() != 1 {
t.Fatal("cleanup failure retained old permission")
}
// Subsequent fixture writes need no additional RPC handlers.
globalNotificationSys = &NotificationSys{}
// Recreate after the partial cleanup. The old mapping, group member
// and child still exist in storage; none may authorize this identity.
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), user, req)
mustIAM(t, err)
reloaded := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, reloaded.LoadIAMCache(ctx, true))
fresh := &IAMSys{store: reloaded, usersSysType: MinIOUsersSysType}
if fresh.IsAllowed(args) {
t.Fatal("cold reload restored partially cleaned-up grants")
}
if _, ok := reloaded.GetUser(child.AccessKey); ok {
t.Fatal("cold reload restored the old child")
}
gd, err := reloaded.GetGroupDescription("commit-group")
mustIAM(t, err)
if len(gd.Members) != 0 {
t.Fatalf("listing exposed a revoked group relation: %v", gd.Members)
}
_, err = sys.AddUsersToGroup(ctx, "commit-group", []string{user})
mustIAM(t, err)
if !sys.IsAllowed(args) {
t.Fatal("explicit new group grant was not accepted")
}
})
}
}
func TestIAMGroupGrantVersionsSurviveSnapshotsAndRecreation(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
for _, user := range []string{"grant-alice", "grant-bob"} {
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, req)
mustIAM(t, err)
}
grant := origin.Add(time.Minute)
_, err := sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), "grant-group", []string{"grant-alice"})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "grant-group", "readwrite", regUser, true)
mustIAM(t, err)
boundary := origin.Add(2 * time.Minute)
mustIAM(t, sys.DeleteUser(withIAMReplicationTime(ctx, boundary), "grant-alice", false))
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), "grant-alice", req)
mustIAM(t, err)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(4*time.Minute)), "grant-group", []string{"grant-bob"})
mustIAM(t, err)
_, err = sys.SetGroupStatus(withIAMReplicationTime(ctx, origin.Add(5*time.Minute)), "grant-group", true)
mustIAM(t, err)
var gi GroupInfo
mustIAM(t, sys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath("grant-group")))
if !gi.MemberGrants["grant-alice"].Equal(grant) {
t.Fatal("unrelated group edits refreshed an old grant")
}
args := policy.Args{AccountName: "grant-alice", Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
for _, stale := range []time.Time{grant, boundary, {}} {
item := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, UpdatedAt: origin.Add(6 * time.Minute), GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: "grant-group", Members: []string{"grant-alice", "grant-bob"}}}}, GroupSnapshot: true, GroupGrants: map[string]time.Time{"grant-alice": stale, "grant-bob": origin.Add(4 * time.Minute)}}
mustIAM(t, applyIAMReplicationItem(ctx, item))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
if sys.IsAllowed(args) {
t.Fatalf("snapshot restored revoked grant %s", stale)
}
gd, err := sys.GetGroupDescription("grant-group")
mustIAM(t, err)
if len(gd.Members) != 1 || gd.Members[0] != "grant-bob" {
t.Fatalf("inconsistent effective members: %v", gd.Members)
}
}
// Only an explicit post-revocation grant restores access.
freshAt, err := sys.AddUsersToGroup(ctx, "grant-group", []string{"grant-alice"})
mustIAM(t, err)
if !sys.IsAllowed(args) {
t.Fatal("explicit regrant rejected")
}
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
mustIAM(t, sys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath("grant-group")))
if !gi.MemberGrants["grant-alice"].Equal(freshAt) {
t.Fatal("new grant version was not persisted")
}
if !gi.MemberGrants["grant-bob"].Equal(origin.Add(4 * time.Minute)) {
t.Fatal("regranting Alice changed Bob's grant")
}
}
func TestIAMGroupRevocationCommitAndRecreation(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) { testIAMGroupRevocationCommitAndRecreation(t, backend) })
}
}
func testIAMGroupRevocationCommitAndRecreation(t *testing.T, backend string) {
ctx, sys, obj := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
user, group := "group-boundary-user", "group-boundary"
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
grant, boundary := origin.Add(time.Minute), origin.Add(2*time.Minute)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), group, []string{user})
mustIAM(t, err)
// A newer mapping must not veto the authoritative group deletion.
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), group, "readwrite", regUser, true)
mustIAM(t, err)
_, err = sys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, boundary), group, nil)
mustIAM(t, err)
r, err := loadIAMRevision(ctx, sys.store, getGroupInfoPath(group))
mustIAM(t, err)
if !r.Deleted || !r.RevokedBefore.Equal(boundary) {
t.Fatal("newer mapping swallowed group deletion")
}
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(4*time.Minute)), group, nil)
mustIAM(t, err)
for _, at := range []time.Time{grant, boundary, {}} {
item := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, UpdatedAt: origin.Add(5 * time.Minute), GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}}, GroupSnapshot: true, GroupGrants: map[string]time.Time{user: at}}
mustIAM(t, applyIAMReplicationItem(ctx, item))
gd, err := sys.GetGroupDescription(group)
mustIAM(t, err)
if len(gd.Members) != 0 {
t.Fatalf("group recreation restored grant %s", at)
}
}
_, err = sys.AddUsersToGroup(ctx, group, []string{user})
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("explicit group regrant was rejected")
}
// The newer live snapshot may arrive before an older group deletion.
lateBoundary := origin.Add(6 * time.Minute)
_, err = sys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, lateBoundary), group, nil)
mustIAM(t, err)
r, err = loadIAMRevision(ctx, sys.store, getGroupInfoPath(group))
mustIAM(t, err)
if r.Deleted || !r.RevokedBefore.Equal(lateBoundary) {
t.Fatal("late deletion lost the live group's revocation boundary")
}
// The old mapping is now revoked; a new explicit mapping restores access.
if sys.IsAllowed(args) {
t.Fatal("late group boundary retained an old mapping")
}
_, err = sys.PolicyDBSet(ctx, group, "readwrite", regUser, true)
mustIAM(t, err)
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
if es, ok := sys.store.IAMStorageAPI.(*IAMEtcdStore); ok {
store.IAMStorageAPI = newIAMEtcdStore(es.client, MinIOUsersSysType)
}
mustIAM(t, store.LoadIAMCache(ctx, true))
fresh := &IAMSys{store: store, usersSysType: MinIOUsersSysType}
if !fresh.IsAllowed(args) {
t.Fatal("reload lost explicit grants after a retained group boundary")
}
item, err := globalSiteReplicationSys.replicationItem(ctx, madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group}}, UpdatedAt: r.timestamp()})
mustIAM(t, err)
if !item.RevokedBefore.Equal(lateBoundary) || !item.GroupGrants[user].After(lateBoundary) {
t.Fatal("group snapshot lost revision metadata")
}
}
// A committed revision is observable before all cached dependents have been
// cleaned up. Every authorization read must apply that boundary in this window.
func TestIAMCachedMappingHonorsCommittedRevision(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
parent := "cached-external-parent"
_, err := sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), parent, "readwrite", stsUser, false)
mustIAM(t, err)
policies, err := sys.PolicyDBGet(parent)
mustIAM(t, err)
if len(policies) == 0 {
t.Fatal("fixture has no STS-parent mapping")
}
mustIAM(t, sys.store.saveIAMConfig(ctx, &MappedPolicy{Version: 1, Deleted: true, UpdatedAt: origin.Add(time.Minute)}, getMappedPolicyPath(parent, stsUser, false)))
policies, err = sys.PolicyDBGet(parent)
mustIAM(t, err)
if len(policies) != 0 {
t.Fatal("cached STS mapping ignored its own namespace tombstone")
}
user, group := "cached-group-user", "cached-group"
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin), user, madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
grant := origin.Add(5 * time.Minute)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), group, []string{user})
mustIAM(t, err)
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), group, "readwrite", regUser, true)
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("fixture has no group grant")
}
// A late deletion preserves the newer member grant but revokes the older
// policy mapping. Simulate the interval before mapping cleanup completes.
gi := GroupInfo{Version: 1, Status: statusEnabled, Members: []string{user}, MemberGrants: map[string]time.Time{user: grant}, UpdatedAt: grant, RevokedBefore: origin.Add(2 * time.Minute)}
mustIAM(t, sys.store.saveIAMConfig(ctx, &gi, getGroupInfoPath(group)))
if sys.IsAllowed(args) {
t.Fatal("cached group mapping ignored the committed group boundary")
}
gd, err := sys.GetGroupDescription(group)
mustIAM(t, err)
if gd.Policy != "" {
t.Fatal("group listing exposed a revoked mapping")
}
}
type (
iamExpiryLockFailure struct {
ObjectLayer
path string
}
iamFailedExpiryLock struct{ RWLocker }
)
func (o *iamExpiryLockFailure) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket == minioMetaBucket && len(objects) == 1 && objects[0] == o.path+".revision-lock" {
return &iamFailedExpiryLock{RWLocker: lock}
}
return lock
}
func (l *iamFailedExpiryLock) GetLock(context.Context, *dynamicTimeout) (LockContext, error) {
return LockContext{}, errIAMInjectedWrite
}
func TestIAMExpiredCredentialCleanupDoesNotBlockLoading(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "healthy-user", madmin.AddOrUpdateUserReq{SecretKey: "healthy-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "healthy-user", "readwrite", regUser, false)
mustIAM(t, err)
c, _, err := sys.NewServiceAccount(ctx, "healthy-user", nil, newServiceAccountOpts{accessKey: "expired-service", secretKey: "expired-service-password"})
mustIAM(t, err)
c.Expiration = UTCNow().Add(-time.Hour)
path := getUserIdentityPath(c.AccessKey, svcUser)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: c, UpdatedAt: UTCNow()}, path))
// A cold loader sees the existing version but cannot acquire the cleanup
// write lock. Healthy users must still load; the expired one stays denied.
fresh := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(&iamExpiryLockFailure{ObjectLayer: obj, path: path}, MinIOUsersSysType)}
mustIAM(t, fresh.LoadIAMCache(ctx, true))
if _, ok := fresh.GetUser("healthy-user"); !ok {
t.Fatal("cleanup failure prevented healthy IAM state from loading")
}
if _, ok := fresh.GetUser(c.AccessKey); ok {
t.Fatal("cleanup failure admitted an expired service account")
}
r, err := loadIAMRevision(ctx, fresh, path)
mustIAM(t, err)
if r.Deleted || !r.Credentials.IsExpired() {
t.Fatal("failed cleanup lost the existing expired revision")
}
}
+220
View File
@@ -0,0 +1,220 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"fmt"
"maps"
"strings"
"sync"
"time"
"github.com/minio/minio/internal/auth"
)
// This index is rebuilt by the existing IAM loaders and updated by successful
// storage operations. It avoids a second full IAM walk during every heal pass.
// It is an optimization of the durable records, never a reason to delete them.
// The index contains no secrets or grants.
type iamParentRevision struct {
deleted bool
before time.Time
}
type iamRevisionIndex struct {
mu sync.RWMutex
items map[string]iamRevision
parents map[string]iamParentRevision
floors map[string]time.Time
generation uint64
}
func (idx *iamRevisionIndex) observe(path string, data []byte) {
if !strings.HasPrefix(path, iamConfigPrefix+"/") {
return
}
var r iamRevision
if json.Unmarshal(data, &r) != nil {
return // The caller reports malformed data using its normal decoder.
}
r.Credentials = auth.Credentials{ParentUser: r.Credentials.ParentUser, Expiration: r.Credentials.Expiration}
idx.mu.Lock()
defer idx.mu.Unlock()
if strings.HasPrefix(path, iamConfigUsersPrefix) {
// Keep a compact name-keyed view for the authentication hot path;
// constructing a config path on every S3 request allocates needlessly.
defer func() {
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile)
if current, ok := idx.items[path]; ok {
if idx.parents == nil {
idx.parents = make(map[string]iamParentRevision)
}
idx.parents[name] = iamParentRevision{deleted: current.Deleted, before: current.RevokedBefore}
} else {
delete(idx.parents, name)
}
}()
}
if floor, ok := idx.floors[path]; ok && r.timestamp().Before(floor) {
return
}
if previous, ok := idx.items[path]; ok {
// A concurrent read that began before a write must not roll it back.
if previous.timestamp().After(r.timestamp()) || (previous.Deleted && !r.Deleted && !r.timestamp().After(previous.timestamp())) {
return
}
if previous.RevokedBefore.After(r.RevokedBefore) {
r.RevokedBefore = previous.RevokedBefore
}
if previous.timestamp().Equal(r.timestamp()) && previous.Deleted == r.Deleted && previous.RevokedBefore.Equal(r.RevokedBefore) {
return
}
}
if r.Deleted && !r.ExpiresAt.IsZero() && UTCNow().After(r.ExpiresAt) {
if _, tracked := idx.items[path]; tracked {
delete(idx.items, path)
idx.generation++
}
delete(idx.floors, path)
return
}
if !r.Deleted && r.RevokedBefore.IsZero() {
_, tracked := idx.items[path]
_, hasFloor := idx.floors[path]
if tracked || hasFloor {
if idx.floors == nil {
idx.floors = make(map[string]time.Time)
}
idx.floors[path] = r.timestamp()
}
if tracked {
delete(idx.items, path)
idx.generation++
}
return
}
if idx.items == nil {
idx.items = make(map[string]iamRevision)
}
idx.items[path] = r
delete(idx.floors, path)
idx.generation++
}
func (idx *iamRevisionIndex) get(path string) iamRevision {
if idx == nil {
return iamRevision{}
}
idx.mu.RLock()
defer idx.mu.RUnlock()
return idx.items[path]
}
func (idx *iamRevisionIndex) snapshot() map[string]iamRevision {
idx.mu.Lock()
defer idx.mu.Unlock()
for path, r := range idx.items {
if r.Deleted && !r.ExpiresAt.IsZero() && UTCNow().After(r.ExpiresAt) {
delete(idx.items, path)
delete(idx.floors, path)
idx.generation++
}
}
return maps.Clone(idx.items)
}
func (idx *iamRevisionIndex) count() int {
idx.mu.RLock()
defer idx.mu.RUnlock()
return len(idx.items)
}
// A process-local generation plus the protocol's instance ID is sufficient
// for acknowledgements. Avoid hashing the entire index on every IAM write.
func (idx *iamRevisionIndex) digest() string {
idx.mu.RLock()
defer idx.mu.RUnlock()
return fmt.Sprintf("%x:%x", idx.generation, len(idx.items))
}
func (idx *iamRevisionIndex) forget(path string) {
idx.mu.Lock()
if _, ok := idx.items[path]; ok {
delete(idx.items, path)
idx.generation++
}
delete(idx.floors, path)
if strings.HasPrefix(path, iamConfigUsersPrefix) {
delete(idx.parents, strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile))
}
idx.mu.Unlock()
}
func (c *iamCache) userRevocation(user string) iamRevision {
r := c.revisions.parentRevision(user)
if u, ok := c.iamUsersMap[user]; ok && u.RevokedBefore.After(r.RevokedBefore) {
r.RevokedBefore = u.RevokedBefore
}
return r
}
func (c *iamCache) groupMemberAllowed(member string, grantedAt, groupBoundary time.Time) bool {
r := c.userRevocation(member)
return !r.Deleted && (r.RevokedBefore.IsZero() || grantedAt.After(r.RevokedBefore)) && (groupBoundary.IsZero() || grantedAt.After(groupBoundary))
}
func iamMappingParentPath(path string) string {
kind, name, ok := strings.Cut(strings.TrimPrefix(path, iamConfigPolicyDBPrefix), "/")
if !ok {
return ""
}
name = strings.TrimSuffix(name, ".json")
switch kind {
case "users", "sts-users":
return getUserIdentityPath(name, regUser)
case "service-accounts":
return getUserIdentityPath(name, svcUser)
case "groups":
return getGroupInfoPath(name)
}
return ""
}
func (idx *iamRevisionIndex) mappingAllowed(path string, mp MappedPolicy) bool {
if mp.Deleted || idx.get(path).Deleted {
return false
}
r := idx.get(iamMappingParentPath(path))
return !r.Deleted && (r.RevokedBefore.IsZero() || mp.UpdatedAt.After(r.RevokedBefore))
}
// Apply the persisted commit boundary even before dependent cache cleanup has
// completed. The map namespace is part of the authorization record's identity.
func (c *iamCache) cachedMappedPolicy(name string, userType IAMUserType, isGroup bool) (MappedPolicy, bool) {
var mp MappedPolicy
var ok bool
switch {
case isGroup:
mp, ok = c.iamGroupPolicyMap.Load(name)
case userType == stsUser:
mp, ok = c.iamSTSPolicyMap.Load(name)
default:
mp, ok = c.iamUserPolicyMap.Load(name)
}
if !ok || !c.revisions.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), mp) {
return MappedPolicy{}, false
}
return mp, true
}
func (idx *iamRevisionIndex) parentRevision(user string) iamRevision {
if idx == nil {
return iamRevision{}
}
idx.mu.RLock()
p := idx.parents[user]
idx.mu.RUnlock()
return iamRevision{Deleted: p.deleted, RevokedBefore: p.before}
}
+305
View File
@@ -0,0 +1,305 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"crypto/sha256"
"fmt"
"os"
"strings"
"sync"
"testing"
"time"
"github.com/minio/madmin-go/v3"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/concurrency"
"go.etcd.io/etcd/client/v3/namespace"
)
type iamRevisionLockObserver struct {
ObjectLayer
path string
waiting chan struct{}
once sync.Once
}
func (o *iamRevisionLockObserver) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket == minioMetaBucket && len(objects) == 1 && objects[0] == o.path {
return &iamRevisionObservedLock{RWLocker: lock, observe: func() { o.once.Do(func() { close(o.waiting) }) }}
}
return lock
}
type iamRevisionObservedLock struct {
RWLocker
observe func()
}
func (l *iamRevisionObservedLock) GetLock(ctx context.Context, timeout *dynamicTimeout) (LockContext, error) {
l.observe()
return l.RWLocker.GetLock(ctx, timeout)
}
type iamRevisionWatchObserver struct {
etcd.Watcher
waiting chan struct{}
once sync.Once
}
func (w *iamRevisionWatchObserver) Watch(ctx context.Context, key string, opts ...etcd.OpOption) etcd.WatchChan {
w.once.Do(func() { close(w.waiting) })
return w.Watcher.Watch(ctx, key, opts...)
}
// Simulate an unavailable cleanup RPC. Mutex.Lock calls Delete after its wait
// is canceled; that RPC must inherit a deadline too, not Client.Ctx() forever.
type iamRevisionCleanupBlocker struct {
etcd.KV
release chan struct{}
}
func (b *iamRevisionCleanupBlocker) Delete(ctx context.Context, key string, opts ...etcd.OpOption) (*etcd.DeleteResponse, error) {
if strings.Contains(key, "/iam-revision-locks/") {
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-b.release:
}
}
return b.KV.Delete(ctx, key, opts...)
}
type iamRevisionReadBlocker struct {
IAMStorageAPI
path string
after int
waiting chan struct{}
}
func (b *iamRevisionReadBlocker) loadIAMConfig(ctx context.Context, item any, path string) error {
if path == b.path {
b.after--
if b.after == 0 {
close(b.waiting)
<-ctx.Done()
return ctx.Err()
}
}
return b.IAMStorageAPI.loadIAMConfig(ctx, item, path)
}
func TestIAMRevisionReadDoesNotBlockAuthentication(t *testing.T) {
for _, stage := range []struct {
name string
offset time.Duration
}{{"deletion", time.Minute}, {"retained_revocation", -time.Minute}} {
t.Run(stage.name, func(t *testing.T) {
resetTestGlobals()
t.Cleanup(resetTestGlobals)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() {
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
})
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
const user = "read-blocked-parent"
created, err := store.AddUser(ctx, user, madmin.AddOrUpdateUserReq{SecretKey: "original-password", Status: madmin.AccountEnabled})
if err != nil {
t.Fatal(err)
}
blocked := &iamRevisionReadBlocker{IAMStorageAPI: store.IAMStorageAPI, path: getUserIdentityPath(user, regUser), after: 1, waiting: make(chan struct{})}
store.IAMStorageAPI = blocked
done := make(chan error, 1)
go func() {
done <- store.DeleteUser(withIAMReplicationTime(ctx, created.Add(stage.offset)), user, regUser)
}()
defer func() { cancel(); <-done }()
select {
case <-blocked.waiting:
case <-time.After(5 * time.Second):
t.Fatal("revision read was not attempted")
}
read := make(chan bool, 1)
go func() {
u, ok := store.GetUser(user)
read <- ok && u.Credentials.SecretKey == "original-password"
}()
select {
case ok := <-read:
if !ok {
t.Fatal("pending revision read changed the cached identity")
}
case <-time.After(time.Second):
t.Fatal("revision read blocked cached authentication")
}
})
}
}
func TestIAMRevisionLockContention(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if backend == "etcd" && endpoint == "" {
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
for _, outcome := range []string{"release", "cancel", "default_timeout"} {
t.Run(outcome, func(t *testing.T) {
resetTestGlobals()
t.Cleanup(resetTestGlobals)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
oldTimeout := defaultContextTimeout
defaultContextTimeout = 2 * time.Second
t.Cleanup(func() { defaultContextTimeout = oldTimeout })
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
const user = "contended-user"
path := getUserIdentityPath(user, regUser)
waiting := make(chan struct{})
var store *IAMStoreSys
var hold func() func()
unblockCleanup := func() {}
if backend == "object" {
disks, err := getRandomDisks(1)
must(err)
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
must(err)
t.Cleanup(func() {
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
})
observed := &iamRevisionLockObserver{ObjectLayer: obj, path: path + ".revision-lock", waiting: waiting}
store = &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
hold = func() func() {
lock := obj.NewNSLock(minioMetaBucket, observed.path)
lc, err := lock.GetLock(ctx, newDynamicTimeout(time.Second, time.Second))
must(err)
store.IAMStorageAPI.(*IAMObjectStore).objAPI = observed
return func() { lock.Unlock(lc) }
}
} else {
client, err := etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: time.Second})
must(err)
t.Cleanup(func() { client.Close() })
prefix := fmt.Sprintf("/silo-lock-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(client.KV, prefix)
client.Watcher = namespace.NewWatcher(client.Watcher, prefix)
store = &IAMStoreSys{IAMStorageAPI: newIAMEtcdStore(client, MinIOUsersSysType)}
hold = func() func() {
session, err := concurrency.NewSession(client, concurrency.WithContext(ctx))
must(err)
lock := concurrency.NewMutex(session, fmt.Sprintf("%s/iam-revision-locks/%x", minioConfigPrefix, sha256.Sum256([]byte(path))))
must(lock.Lock(ctx))
client.Watcher = &iamRevisionWatchObserver{Watcher: client.Watcher, waiting: waiting}
blocker := &iamRevisionCleanupBlocker{KV: client.KV, release: make(chan struct{})}
client.KV = blocker
unblockCleanup = sync.OnceFunc(func() { close(blocker.release) })
t.Cleanup(unblockCleanup)
return func() { session.Close() }
}
}
request := func(secret string) madmin.AddOrUpdateUserReq {
return madmin.AddOrUpdateUserReq{SecretKey: secret, Status: madmin.AccountEnabled}
}
_, err := store.AddUser(ctx, user, request("original-password"))
must(err)
release := sync.OnceFunc(hold())
t.Cleanup(release)
writeCtx, cancelWrite := context.WithCancel(ctx)
defer cancelWrite()
first, second := make(chan error, 1), make(chan error, 1)
var writers sync.WaitGroup
t.Cleanup(func() {
cancelWrite()
unblockCleanup()
release()
writers.Wait()
})
writers.Go(func() {
_, err := store.AddUser(writeCtx, user, request("first-password"))
first <- err
})
select {
case <-waiting:
case <-time.After(5 * time.Second):
t.Fatal("writer did not attempt the held revision lock")
}
// A second writer must queue without taking the cache's RWMutex:
// Go's writer preference would otherwise block every new reader.
writers.Go(func() {
_, err := store.AddUser(ctx, user, request("second-password"))
second <- err
})
select {
case err := <-second:
t.Fatalf("second writer bypassed the first: %v", err)
case <-time.After(50 * time.Millisecond):
}
read := make(chan UserIdentity, 1)
go func() {
u, _ := store.GetUser(user)
read <- u
}()
select {
case u := <-read:
if u.Credentials.SecretKey != "original-password" {
t.Fatal("pending write changed the cached credential")
}
case <-time.After(time.Second):
t.Fatal("distributed lock contention blocked cached authentication")
}
switch outcome {
case "release":
release()
case "cancel":
cancelWrite()
}
select {
case err := <-first:
if outcome == "release" {
must(err)
} else if err == nil {
t.Fatal("canceled or timed-out write succeeded")
}
case <-time.After(5 * time.Second):
t.Fatal("lock wait or cancellation cleanup exceeded its deadline")
}
release()
select {
case err := <-second:
must(err)
case <-time.After(5 * time.Second):
t.Fatal("queued writer did not recover after the first completed")
}
cached, ok := store.GetUser(user)
if !ok || cached.Credentials.SecretKey != "second-password" {
t.Fatal("cached write order was lost")
}
var persisted UserIdentity
must(store.loadIAMConfig(ctx, &persisted, path))
if persisted.Credentials.SecretKey != cached.Credentials.SecretKey || !persisted.UpdatedAt.Equal(cached.UpdatedAt) {
t.Fatal("persistent and cached revisions differ")
}
})
}
})
}
}
+749
View File
@@ -0,0 +1,749 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"crypto/sha256"
"errors"
"fmt"
"net/http"
"sort"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/concurrency"
)
var errIAMStaleUpdate = errors.New("IAM update predates a stored revision or revocation")
// The parent is still live; callers must not broadcast a user deletion when
// only its revocation boundary was retained.
var errIAMRevocationRetained = errors.New("IAM revocation recorded without deleting the record")
// A revocation advances the boundary even when a newer identity already
// exists. Keep this operation distinct from replacing/deleting that identity.
type iamUserRevocation struct {
UserIdentity
retained bool
}
type iamGroupRevocation struct {
GroupInfo
retained bool
requireEmpty bool
}
// Natural expiration is distinct from revoking a live credential. An expired
// immutable STS token can be removed; a reusable service-account key retains
// its revision so an older non-expiring credential cannot return.
type iamExpireIdentity struct{}
// The authoritative revocation is durable even if dependent cleanup fails.
// Callers must publish it to sibling caches before returning the error.
type iamCommittedCleanupError struct {
err error
retained bool
}
func (e *iamCommittedCleanupError) Error() string {
return "IAM revocation committed; cleanup failed: " + e.err.Error()
}
func (e *iamCommittedCleanupError) Unwrap() error { return e.err }
type iamReplicationTimeKey struct{}
func withIAMReplicationTime(ctx context.Context, at time.Time) context.Context {
return context.WithValue(ctx, iamReplicationTimeKey{}, at)
}
func iamReplicationTime(ctx context.Context) (time.Time, bool) {
at, ok := ctx.Value(iamReplicationTimeKey{}).(time.Time)
return at, ok
}
func iamReplicationError(err error) error {
if errors.Is(err, errIAMStaleUpdate) {
// Retrying an obsolete event cannot change the result.
return nil
}
return wrapSRErr(err)
}
// Deletions occupy the original IAM config path. They contain no secret or
// grant and are hidden by the normal loaders, but remain available to heal
// and to timestamp comparisons after a restart. Do not age them out: a peer
// can be offline indefinitely.
type iamRevision struct {
UpdatedAt time.Time `json:"updatedAt"`
UpdateDate time.Time `json:"UpdateDate"`
Deleted bool `json:"deleted"`
RevokedBefore time.Time `json:"revokedBefore"`
ExpiresAt time.Time `json:"expiresAt,omitempty"`
Credentials auth.Credentials `json:"credentials"`
}
func (r iamRevision) timestamp() time.Time {
if r.UpdateDate.After(r.UpdatedAt) {
return r.UpdateDate
}
return r.UpdatedAt
}
func loadIAMRevision(ctx context.Context, store IAMStorageAPI, path string) (iamRevision, error) {
var r iamRevision
err := store.loadIAMConfig(ctx, &r, path)
if errors.Is(err, errConfigNotFound) {
err = nil
}
return r, err
}
func (store *IAMStoreSys) checkIAMRevision(ctx context.Context, path string, deleting bool) error {
at, replicated := iamReplicationTime(ctx)
if !replicated {
return nil
}
return store.withIAMStorage(ctx, func(ctx context.Context) error {
r, err := loadIAMRevision(ctx, store.IAMStorageAPI, path)
if err != nil {
return err
}
if r.timestamp().After(at) || (r.Deleted && !deleting && !at.After(r.timestamp())) {
return errIAMStaleUpdate
}
return nil
})
}
// This signed claim records the parent's revocation boundary at issuance.
// Unlike UpdatedAt, it cannot advance when an offline site edits an old child.
// It travels in the existing service-account Claims and STS SessionToken fields.
const iamParentRevocationClaim = "siloParentRevocation"
func setIAMParentRevocationClaim(ctx context.Context, store IAMStorageAPI, parent string, claims map[string]any) error {
delete(claims, iamParentRevocationClaim)
if parent == "" || parent == globalActiveCred.AccessKey {
return nil
}
r, err := loadIAMRevision(ctx, store, getUserIdentityPath(parent, regUser))
if err != nil {
return err
}
if r.Deleted {
return errIAMStaleUpdate
}
if !r.RevokedBefore.IsZero() {
claims[iamParentRevocationClaim] = r.RevokedBefore.Format(time.RFC3339Nano)
}
return nil
}
func iamCredentialSurvivesRevocation(cred auth.Credentials, at time.Time) bool {
if at.IsZero() {
return true
}
s, _ := cred.Claims[iamParentRevocationClaim].(string)
issuedAfter, err := time.Parse(time.RFC3339Nano, s)
return err == nil && !issuedAfter.Before(at)
}
// Parent revocations delete old children even if an offline peer has edited
// them later. Preserve children that prove issuance after this revocation.
func iamChildDeletionContext(ctx context.Context, child UserIdentity) (context.Context, bool) {
if at, replicated := iamReplicationTime(ctx); replicated {
if !at.IsZero() && iamCredentialSurvivesRevocation(child.Credentials, at) {
return ctx, false
}
if child.UpdatedAt.After(at) {
ctx = withIAMReplicationTime(ctx, child.UpdatedAt)
}
}
return ctx, true
}
// A delayed service account or STS event must not outlive deletion of its
// built-in parent. The caller must populate Claims from the verified token.
func checkIAMParentRevision(ctx context.Context, store IAMStorageAPI, cred auth.Credentials) error {
parent := cred.ParentUser
if parent == "" || parent == globalActiveCred.AccessKey {
return nil
}
r, err := loadIAMRevision(ctx, store, getUserIdentityPath(parent, regUser))
if err != nil {
return err
}
if r.Deleted || !iamCredentialSurvivesRevocation(cred, r.RevokedBefore) {
return errIAMStaleUpdate
}
return nil
}
// Called with the IAM writer mutex and cache lock held. Persistence only
// touches the caller's record, not the cache. Keep writers serialized while
// allowing cached authentication reads throughout storage and lock waits.
func (store *IAMStoreSys) withIAMStorage(ctx context.Context, fn func(context.Context) error) error {
store.IAMStorageAPI.unlock()
defer store.IAMStorageAPI.lock()
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
return fn(ctx)
}
func (store *IAMStoreSys) saveIAMRevision(ctx context.Context, path string, item any, opts ...options) error {
return store.withIAMStorage(ctx, func(ctx context.Context) error {
return saveIAMRevision(ctx, store.IAMStorageAPI, path, item, opts...)
})
}
func (store *IAMStoreSys) checkIAMParentRevision(ctx context.Context, cred auth.Credentials) error {
return store.withIAMStorage(ctx, func(ctx context.Context) error {
return checkIAMParentRevision(ctx, store.IAMStorageAPI, cred)
})
}
// Update the caller's record with the persisted revision before it is cached.
func saveIAMRevision(ctx context.Context, store IAMStorageAPI, path string, item any, opts ...options) error {
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
// Serialize compare-and-write across nodes, as well as goroutines. Use a
// separate lock name so saving the config does not reacquire this lock.
switch s := store.(type) {
case *IAMObjectStore:
lock := s.objAPI.NewNSLock(minioMetaBucket, path+".revision-lock")
lc, err := lock.GetLock(ctx, globalOperationTimeout)
if err != nil {
return err
}
defer lock.Unlock(lc)
ctx = lc.Context()
case *IAMEtcdStore:
// Mutex.Lock also uses Client.Ctx() for cleanup after cancellation.
// Borrow the existing services with the operation's bounded context;
// never close this facade, which does not own those services.
client := etcd.NewCtxClient(ctx, etcd.WithZapLogger(s.client.GetLogger()))
client.KV, client.Lease, client.Watcher = s.client.KV, s.client.Lease, s.client.Watcher
session, err := concurrency.NewSession(client, concurrency.WithContext(ctx))
if err != nil {
return err
}
defer func() {
session.Orphan()
// A canceled operation must still release its lease when etcd is
// reachable. If it is unavailable, stop waiting and let it expire.
cleanupCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), defaultContextTimeout)
defer cancel()
_, _ = s.client.Revoke(cleanupCtx, session.Lease())
}()
lock := concurrency.NewMutex(session, fmt.Sprintf("%s/iam-revision-locks/%x", minioConfigPrefix, sha256.Sum256([]byte(path))))
if err = lock.Lock(ctx); err != nil {
return err
}
// Revoking the session lease releases the lock, including on cancellation.
}
previous, err := loadIAMRevision(ctx, store, path)
if err != nil {
return err
}
if _, expiring := item.(*iamExpireIdentity); expiring {
sts := strings.HasPrefix(path, iamConfigSTSPrefix)
if previous.Deleted {
if sts && !previous.ExpiresAt.IsZero() && UTCNow().After(previous.ExpiresAt) {
return expireIAMSTSConfig(ctx, store, path)
}
return nil
}
if previous.timestamp().IsZero() || !previous.Credentials.IsExpired() {
return nil
}
if sts {
return expireIAMSTSConfig(ctx, store, path)
}
item = &UserIdentity{Version: 1, Deleted: true}
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
var revocation *iamUserRevocation
if op, ok := item.(*iamUserRevocation); ok {
revocation = op
op.UserIdentity = UserIdentity{Version: 1, Deleted: true}
if origin, replicated := iamReplicationTime(ctx); replicated && previous.timestamp().After(origin) {
if previous.Deleted || !origin.After(previous.RevokedBefore) {
return errIAMStaleUpdate
}
op.retained = true
op.UserIdentity = UserIdentity{Version: 1, Credentials: previous.Credentials, UpdatedAt: previous.timestamp(), RevokedBefore: origin}
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
item = &op.UserIdentity
}
var groupRevocation *iamGroupRevocation
if op, ok := item.(*iamGroupRevocation); ok {
groupRevocation = op
var group GroupInfo
if err := store.loadIAMConfig(ctx, &group, path); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
if op.requireEmpty && !group.Deleted {
for _, member := range group.Members {
r := store.revisionIndex().get(getUserIdentityPath(member, regUser))
at := group.MemberGrants[member]
if !r.Deleted && (r.RevokedBefore.IsZero() || at.After(r.RevokedBefore)) && (group.RevokedBefore.IsZero() || at.After(group.RevokedBefore)) {
return errGroupNotEmpty
}
}
}
op.GroupInfo = GroupInfo{Version: 1, Deleted: true}
if origin, replicated := iamReplicationTime(ctx); replicated && previous.timestamp().After(origin) {
if previous.Deleted || !origin.After(previous.RevokedBefore) {
return errIAMStaleUpdate
}
op.retained = true
op.GroupInfo = group
op.RevokedBefore = origin
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
item = &op.GroupInfo
}
var at *time.Time
var deleted bool
switch v := item.(type) {
case *UserIdentity:
at, deleted = &v.UpdatedAt, v.Deleted
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(v.RevokedBefore) {
v.RevokedBefore = boundary
}
if previous.RevokedBefore.After(v.RevokedBefore) {
v.RevokedBefore = previous.RevokedBefore
}
case *GroupInfo:
at, deleted = &v.UpdatedAt, v.Deleted
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(v.RevokedBefore) {
v.RevokedBefore = boundary
}
if !deleted {
var group GroupInfo
if err := store.loadIAMConfig(ctx, &group, path); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
mergeIAMGroupMutation(ctx, group, v)
}
if previous.RevokedBefore.After(v.RevokedBefore) {
v.RevokedBefore = previous.RevokedBefore
}
case *MappedPolicy:
at, deleted = &v.UpdatedAt, v.Deleted
case *PolicyDoc:
at, deleted = &v.UpdateDate, v.Deleted
default:
return errInvalidArgument
}
if strings.HasPrefix(path, iamConfigSTSPrefix) && previous.Deleted && !deleted {
// STS access keys identify immutable tokens, not reusable user names.
return errIAMStaleUpdate
}
if origin, replicated := iamReplicationTime(ctx); replicated {
*at = origin
if previous.timestamp().After(origin) || (previous.Deleted && !deleted && !origin.After(previous.timestamp())) {
return errIAMStaleUpdate
}
if strings.HasPrefix(path, iamConfigServiceAccountsPrefix) && !deleted && previous.Credentials.AccessKey != "" && previous.timestamp().Equal(origin) {
// Duplicate service snapshots are acknowledgements, not new creates
// or edits. Reload the winner without writing, so even a stale
// sibling cache is refreshed by the retry before acknowledging it.
return store.loadIAMConfig(ctx, item, path)
}
if previous.Deleted && deleted && !origin.After(previous.timestamp()) {
// An already-applied tombstone needs no further persistent write.
if v, ok := item.(*UserIdentity); ok {
v.RevokedBefore = previous.RevokedBefore
}
if v, ok := item.(*GroupInfo); ok {
v.RevokedBefore = previous.RevokedBefore
}
return nil
}
} else {
if previous.Deleted && deleted {
// A peer notification without an originating revision must not
// advance a tombstone past a subsequent deliberate recreation.
*at = previous.timestamp()
if v, ok := item.(*UserIdentity); ok {
v.RevokedBefore = previous.RevokedBefore
}
if v, ok := item.(*GroupInfo); ok {
v.RevokedBefore = previous.RevokedBefore
}
return nil
}
if at.IsZero() {
*at = UTCNow()
}
if !at.After(previous.timestamp()) {
*at = previous.timestamp().Add(time.Nanosecond)
}
}
if v, ok := item.(*UserIdentity); ok {
if deleted {
// Retain only the parent name for root-account exclusion during heal.
v.Credentials = auth.Credentials{ParentUser: previous.Credentials.ParentUser}
v.RevokedBefore = *at
if strings.HasPrefix(path, iamConfigSTSPrefix) && !previous.Credentials.Expiration.IsZero() && !previous.Credentials.Expiration.Equal(timeSentinel) {
// The signed STS token cannot authorize beyond this time, even
// if an offline site replays it with a newer event timestamp.
v.ExpiresAt = previous.Credentials.Expiration.Add(globalMaxSkewTime)
opts = []options{{ttl: max(1, int64(time.Until(v.ExpiresAt).Seconds())+1)}}
}
} else {
if v.Credentials.SessionToken != "" && v.Credentials.Claims == nil {
claims, err := extractJWTClaims(*v)
if err != nil {
return err
}
v.Credentials.Claims = claims.Map()
}
if err = checkIAMParentRevision(ctx, store, v.Credentials); err != nil {
return err
}
}
}
if v, ok := item.(*GroupInfo); ok && deleted {
v.RevokedBefore = *at
v.Members, v.MemberGrants = nil, nil
}
if _, ok := item.(*MappedPolicy); ok && !deleted {
if parentPath := iamMappingParentPath(path); parentPath != "" {
parent, err := loadIAMRevision(ctx, store, parentPath)
if err != nil {
return err
}
if parent.Deleted {
return errIAMStaleUpdate
}
if !parent.RevokedBefore.IsZero() && !at.After(parent.RevokedBefore) {
if _, replicated := iamReplicationTime(ctx); replicated {
return errIAMStaleUpdate
}
*at = parent.RevokedBefore.Add(time.Nanosecond)
}
}
}
if err := store.saveIAMConfig(ctx, item, path, opts...); err != nil {
return err
}
if revocation != nil && revocation.retained {
return errIAMRevocationRetained
}
if groupRevocation != nil && groupRevocation.retained {
return errIAMRevocationRetained
}
return nil
}
func (iamOS *IAMObjectStore) listIAMConfigPaths(ctx context.Context) ([]string, error) {
ctx, cancel := context.WithCancel(ctx)
defer cancel()
var paths []string
for item := range listIAMConfigItems(ctx, iamOS.objAPI, iamConfigPrefix+"/") {
if item.Err != nil {
return nil, item.Err
}
paths = append(paths, iamConfigPrefix+"/"+item.Item)
}
return paths, nil
}
func (ies *IAMEtcdStore) listIAMConfigPaths(ctx context.Context) ([]string, error) {
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
r, err := ies.client.Get(ctx, iamConfigPrefix+"/", etcd.WithPrefix(), etcd.WithKeysOnly())
if err != nil {
return nil, err
}
paths := make([]string, 0, len(r.Kvs))
for _, kv := range r.Kvs {
paths = append(paths, string(kv.Key))
}
return paths, nil
}
func iamDeletionItem(path string, r iamRevision) (item madmin.SRIAMItem, ok bool) {
if (strings.HasPrefix(path, iamConfigUsersPrefix) || strings.HasPrefix(path, iamConfigGroupsPrefix)) && !r.RevokedBefore.IsZero() {
// Recreating a parent does not cancel its older revocation of derived
// credentials. Replay this boundary even after the parent is live again.
r.Deleted = true
r.UpdatedAt, r.UpdateDate = r.RevokedBefore, time.Time{}
}
if !r.Deleted {
return item, false
}
item.UpdatedAt = r.timestamp()
switch {
case strings.HasPrefix(path, iamConfigUsersPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile)
item.Type = madmin.SRIAMItemIAMUser
item.IAMUser = &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}
case strings.HasPrefix(path, iamConfigServiceAccountsPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigServiceAccountsPrefix), "/"+iamIdentityFile)
if name == siteReplicatorSvcAcc || r.Credentials.ParentUser == globalActiveCred.AccessKey {
return item, false
}
item.Type = madmin.SRIAMItemSvcAcc
item.SvcAccChange = &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: name}}
case strings.HasPrefix(path, iamConfigGroupsPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigGroupsPrefix), "/"+iamGroupMembersFile)
item.Type = madmin.SRIAMItemGroupInfo
item.GroupInfo = &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: name, IsRemove: true}}
case strings.HasPrefix(path, iamConfigPoliciesPrefix):
item.Type = madmin.SRIAMItemPolicy
item.Name = strings.TrimSuffix(strings.TrimPrefix(path, iamConfigPoliciesPrefix), "/"+iamPolicyFile)
case strings.HasPrefix(path, iamConfigPolicyDBPrefix):
prefix, name, found := strings.Cut(strings.TrimPrefix(path, iamConfigPolicyDBPrefix), "/")
if !found {
return item, false
}
typ := regUser
switch prefix {
case "sts-users":
typ = stsUser
case "service-accounts":
typ = svcUser
}
item.Type = madmin.SRIAMItemPolicyMapping
item.PolicyMapping = &madmin.SRPolicyMapping{UserOrGroup: strings.TrimSuffix(name, ".json"), UserType: int(typ), IsGroup: prefix == "groups"}
default:
// Expired STS credentials are not replayed. Parent revocations and
// their retained timestamp reject delayed copies of derived tokens.
return item, false
}
return item, true
}
func iamDeletionPath(item madmin.SRIAMItem) string {
switch item.Type {
case madmin.SRIAMItemIAMUser:
if item.IAMUser != nil && item.IAMUser.IsDeleteReq {
return getUserIdentityPath(item.IAMUser.AccessKey, regUser)
}
case madmin.SRIAMItemSvcAcc:
if item.SvcAccChange != nil && item.SvcAccChange.Delete != nil {
return getUserIdentityPath(item.SvcAccChange.Delete.AccessKey, svcUser)
}
case madmin.SRIAMItemGroupInfo:
if item.GroupInfo != nil && item.GroupInfo.UpdateReq.IsRemove && len(item.GroupInfo.UpdateReq.Members) == 0 {
return getGroupInfoPath(item.GroupInfo.UpdateReq.Group)
}
case madmin.SRIAMItemPolicy:
if len(item.Policy) == 0 {
return getPolicyDocPath(item.Name)
}
case madmin.SRIAMItemPolicyMapping:
if p := item.PolicyMapping; p != nil && p.Policy == "" {
return getMappedPolicyPath(p.UserOrGroup, IAMUserType(p.UserType), p.IsGroup)
}
}
return ""
}
func (c *SiteReplicationSys) healIAMDeletions(ctx context.Context) (err error) {
started := time.Now()
defer func() {
c.iamRevisionMetrics.healDurationMillis.Store(time.Since(started).Milliseconds())
if err != nil {
c.iamRevisionMetrics.healFailures.Add(1)
} else {
c.iamRevisionMetrics.healLastSuccess.Store(time.Now().Unix())
}
}()
c.iamHealMu.Lock()
defer c.iamHealMu.Unlock()
c.RLock()
defer c.RUnlock()
if !c.enabled {
return nil
}
snapshot := globalIAMSys.store.revisionIndex().snapshot()
paths := make([]string, 0, len(snapshot))
for path := range snapshot {
paths = append(paths, path)
}
sort.Strings(paths)
byType := make(map[string][]iamReplicationItem)
for _, path := range paths {
r := snapshot[path]
item, ok := iamDeletionItem(path, r)
if !ok {
continue
}
out := iamReplicationItem{SRIAMItem: item}
if item.Type == madmin.SRIAMItemIAMUser && !r.Deleted {
out.Type, out.IAMUser = iamUserBoundaryType, nil
out.UserRevocation = &iamUserBoundary{User: item.IAMUser.AccessKey, Before: r.RevokedBefore}
}
if item.Type == madmin.SRIAMItemGroupInfo && !r.Deleted {
out.Type, out.GroupInfo = iamGroupBoundaryType, nil
out.GroupRevocation = &iamGroupBoundary{Group: item.GroupInfo.UpdateReq.Group, Before: r.RevokedBefore}
}
byType[item.Type] = append(byType[item.Type], out)
}
var items []iamReplicationItem
for _, typ := range []string{madmin.SRIAMItemPolicyMapping, madmin.SRIAMItemIAMUser, madmin.SRIAMItemSvcAcc, madmin.SRIAMItemGroupInfo, madmin.SRIAMItemPolicy} {
items = append(items, byType[typ]...)
}
if len(items) == 0 {
return nil
}
if c.iamRevisionProgress == nil {
c.iamRevisionProgress = make(map[string]iamRevisionProgress)
}
for id := range c.iamRevisionProgress {
if _, present := c.state.Peers[id]; !present {
delete(c.iamRevisionProgress, id)
}
}
var progressMu sync.Mutex
cerr := c.concDo(nil, func(id string, p madmin.PeerInfo) error {
// Bound each pass, but retain acknowledgements independently of the
// pass deadline or unrelated changes at either site.
peerCtx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
client, err := c.getAdminClient(peerCtx, id)
if err != nil {
return err
}
remote, err := executeIAMRevisionRequest(peerCtx, client, http.MethodGet, nil)
if err != nil {
return err
}
progressMu.Lock()
progress := c.iamRevisionProgress[id]
progressMu.Unlock()
progress.observePeer(remote)
defer func() {
progressMu.Lock()
c.iamRevisionProgress[id] = progress
progressMu.Unlock()
}()
var pending []iamReplicationItem
for _, item := range items {
path, version := iamReplicationMarker(item)
if progress.Acknowledged[path] != version {
pending = append(pending, item)
}
}
// Acknowledgements are only a replay optimization, never GC proof.
for path := range progress.Acknowledged {
if _, retained := snapshot[path]; !retained {
delete(progress.Acknowledged, path)
}
}
var failures []error
for next := 0; next < len(pending); {
end := min(next+maxIAMRevisionBatch, len(pending))
batch := pending[next:end]
remote, err = executeIAMRevisionRequest(peerCtx, client, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: batch})
if err != nil {
var batchErr *iamRevisionBatchError
if !errors.As(err, &batchErr) {
return errors.Join(append(failures, err)...)
}
failures = append(failures, err)
} else {
progress.observePeer(remote)
for _, item := range batch {
path, version := iamReplicationMarker(item)
progress.Acknowledged[path] = version
}
}
next = end
}
return errors.Join(failures...)
}, "IAM revision convergence")
return errors.Unwrap(cerr)
}
func iamReplicationMarker(item iamReplicationItem) (path, version string) {
path = iamDeletionPath(item.SRIAMItem)
if item.UserRevocation != nil {
path = getUserIdentityPath(item.UserRevocation.User, regUser)
}
if item.GroupRevocation != nil {
path = getGroupInfoPath(item.GroupRevocation.Group)
}
return path, item.Type + ":" + item.UpdatedAt.UTC().Format(time.RFC3339Nano)
}
func (store *IAMStoreSys) savePolicyDoc(ctx context.Context, policyName string, p *PolicyDoc) error {
return store.saveIAMRevision(ctx, getPolicyDocPath(policyName), p)
}
func (store *IAMStoreSys) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp *MappedPolicy, opts ...options) error {
return store.saveIAMRevision(ctx, getMappedPolicyPath(name, userType, isGroup), mp, opts...)
}
func (store *IAMStoreSys) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u *UserIdentity, opts ...options) error {
return store.saveIAMRevision(ctx, getUserIdentityPath(name, userType), u, opts...)
}
func (store *IAMStoreSys) saveGroupInfo(ctx context.Context, name string, gi *GroupInfo) error {
return store.saveIAMRevision(ctx, getGroupInfoPath(name), gi)
}
func (store *IAMStoreSys) deletePolicyDoc(ctx context.Context, name string) error {
return store.saveIAMRevision(ctx, getPolicyDocPath(name), &PolicyDoc{Version: 1, Deleted: true})
}
func (store *IAMStoreSys) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
return store.saveIAMRevision(ctx, getMappedPolicyPath(name, userType, isGroup), &MappedPolicy{Version: 1, Deleted: true})
}
func (store *IAMStoreSys) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
return store.saveIAMRevision(ctx, getUserIdentityPath(name, userType), &UserIdentity{Version: 1, Deleted: true})
}
// Called under the identity's distributed revision lock, after verifying that
// its immutable STS token (or early-revocation retention) has expired. Only the
// old token-key mapping is removed; the reusable parent mapping is unaffected.
func expireIAMSTSConfig(ctx context.Context, store IAMStorageAPI, path string) error {
key := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigSTSPrefix), "/"+iamIdentityFile)
if err := store.deleteIAMConfig(ctx, getMappedPolicyPath(key, stsUser, false)); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
return store.deleteIAMConfig(ctx, path)
}
type (
iamExpirationCleanupKey struct{}
iamExpirationCleanupState struct{ failed atomic.Bool }
)
func withIAMExpirationCleanup(ctx context.Context) context.Context {
if _, ok := ctx.Value(iamExpirationCleanupKey{}).(*iamExpirationCleanupState); ok {
return ctx
}
return context.WithValue(ctx, iamExpirationCleanupKey{}, &iamExpirationCleanupState{})
}
func bestEffortIAMExpiration(ctx context.Context, store IAMStorageAPI, path string) {
state, _ := ctx.Value(iamExpirationCleanupKey{}).(*iamExpirationCleanupState)
if ctx.Err() != nil || (state != nil && state.failed.Load()) {
return
}
ctx, cancel := context.WithTimeout(ctx, time.Second)
defer cancel()
// Failure leaves the expired record and its existing version intact.
// Stop optional reclamation for this load, while still loading healthy
// users. Healthy cleanup has no per-scan quota that could build a backlog.
if err := saveIAMRevision(ctx, store, path, &iamExpireIdentity{}); err != nil {
if state != nil {
state.failed.Store(true)
}
iamLogIf(ctx, err)
}
}
+330
View File
@@ -0,0 +1,330 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
)
// Count physical saves: comparing timestamps alone would miss identical
// tombstones being rewritten on every heal pass.
type iamRevisionWriteCounter struct {
IAMStorageAPI
data []byte
writes int
}
func (s *iamRevisionWriteCounter) loadIAMConfig(_ context.Context, item any, _ string) error {
return json.Unmarshal(s.data, item)
}
func (s *iamRevisionWriteCounter) saveIAMConfig(_ context.Context, item any, _ string, _ ...options) error {
data, err := json.Marshal(item)
if err == nil {
s.data = data
s.writes++
}
return err
}
func TestIAMRevocationTombstoneReplayIsIdempotent(t *testing.T) {
at := time.Date(2026, 9, 14, 12, 0, 0, 0, time.UTC)
for _, record := range []struct {
name string
new func(bool) any
}{
{"user", func(deleted bool) any { return &UserIdentity{Version: 1, Deleted: deleted} }},
{"group", func(deleted bool) any { return &GroupInfo{Version: 1, Deleted: deleted} }},
{"policy", func(deleted bool) any { return &PolicyDoc{Version: 1, Deleted: deleted} }},
{"mapping", func(deleted bool) any { return &MappedPolicy{Version: 1, Deleted: deleted} }},
} {
t.Run(record.name, func(t *testing.T) {
data, err := json.Marshal(iamRevision{Deleted: true, UpdatedAt: at, RevokedBefore: at})
if err != nil {
t.Fatal(err)
}
store := &iamRevisionWriteCounter{data: data}
ctx := context.Background()
for range 3 {
// Site heal carries the original timestamp. Sibling notifications
// have no timestamp; both must leave an applied deletion untouched.
for _, replay := range []context.Context{withIAMReplicationTime(ctx, at), ctx} {
if err := saveIAMRevision(replay, store, record.name, record.new(true)); err != nil {
t.Fatal(err)
}
}
}
if store.writes != 0 || string(store.data) != string(data) {
t.Fatalf("replayed tombstone changed storage: writes=%d, record=%s", store.writes, store.data)
}
for _, deleted := range []bool{false, true} {
err := saveIAMRevision(withIAMReplicationTime(ctx, at.Add(-time.Second)), store, record.name, record.new(deleted))
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("older event accepted, deleted=%t: %v", deleted, err)
}
}
if err := saveIAMRevision(withIAMReplicationTime(ctx, at), store, record.name, record.new(false)); !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("equal-time recreation accepted: %v", err)
}
newer := at.Add(time.Minute)
if err := saveIAMRevision(withIAMReplicationTime(ctx, newer), store, record.name, record.new(true)); err != nil {
t.Fatal(err)
}
r, err := loadIAMRevision(ctx, store, record.name)
if err != nil || store.writes != 1 || !r.timestamp().Equal(newer) || !r.Deleted {
t.Fatalf("newer deletion did not advance storage: writes=%d, revision=%+v, error=%v", store.writes, r, err)
}
if err := saveIAMRevision(withIAMReplicationTime(ctx, newer.Add(time.Minute)), store, record.name, record.new(false)); err != nil {
t.Fatalf("newer recreation rejected: %v", err)
}
})
}
}
// Run with both object storage and etcd through TestIAMRevocation*Lifecycle.
// The object-store case uses the real peer RPC and deletion handler, so a
// spurious notification actually destroys the parent instead of only counting it.
func testIAMRevocationReplayAfterRecreation(ctx context.Context, t *testing.T, sys *IAMSys) {
t.Helper()
peer := &globalSiteReplicationSys
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user := "heal-recreated-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
deleted, recreated := origin.Add(time.Minute), origin.Add(3*time.Minute)
create := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, at))
}
revoke := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, at))
}
create(origin)
revoke(deleted)
create(recreated)
child, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "heal-recreated-child", secretKey: "valid-service-password",
})
must(err)
tg, err := grid.SetupTestGrid(2)
must(err)
t.Cleanup(tg.Cleanup)
var deletes atomic.Int32
server := &peerRESTServer{}
must(deleteUserRPC.Register(tg.Managers[1], func(req *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
deletes.Add(1)
return server.DeleteUserHandler(req)
}))
// Future user updates still use the normal peer reload notification.
must(loadUserRPC.Register(tg.Managers[1], server.LoadUserHandler))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
must(err)
previousNotifications := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection {
return tg.Managers[0].Connection(tg.Hosts[1])
},
}}}
t.Cleanup(func() { globalNotificationSys = previousNotifications })
assertLive := func(key string) {
t.Helper()
if _, ok := sys.GetUser(ctx, key); !ok {
t.Fatalf("live credential %s lost during deletion replay", key)
}
}
assertNoDelete := func() {
t.Helper()
if n := deletes.Load(); n != 0 {
t.Fatalf("retained revocation sent %d destructive sibling notifications", n)
}
}
for range 3 {
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
item, ok := iamDeletionItem(getUserIdentityPath(user, regUser), r)
if !ok || item.IAMUser == nil || !item.UpdatedAt.Equal(deleted) {
t.Fatal("recreated user lost its durable revocation replay")
}
must(peer.PeerIAMUserChangeHandler(ctx, item.IAMUser, item.UpdatedAt))
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
assertLive(child.AccessKey)
assertNoDelete()
}
// A divergent site sends a previously unseen revocation between our old
// boundary and recreation. Retain it and revoke old children, but never
// turn it into an unversioned delete of the recreated parent.
delayed := deleted.Add(time.Minute)
revoke(delayed)
assertNoDelete()
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
if _, ok := sys.GetUser(ctx, child.AccessKey); ok {
t.Fatal("child from before the delayed revocation remains usable")
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
if r.Deleted || !r.RevokedBefore.Equal(delayed) || !r.timestamp().Equal(recreated) {
t.Fatalf("retained revocation damaged the recreated identity: %+v", r)
}
// A genuinely newer deletion must still reach siblings and remove the
// parent plus credentials issued under its latest revocation boundary.
fresh, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "heal-fresh-child", secretKey: "valid-service-password",
})
must(err)
latest := recreated.Add(time.Minute)
revoke(latest)
wantDeletes := int32(1)
if sys.HasWatcher() {
wantDeletes = 0
}
if n := deletes.Load(); n != wantDeletes {
t.Fatalf("new deletion notifications=%d, want %d", n, wantDeletes)
}
for _, key := range []string{user, fresh.AccessKey} {
if _, ok := sys.GetUser(ctx, key); ok {
t.Fatalf("newer deletion left credential %s usable", key)
}
}
// Exercise the actual sibling handler again against the already persisted
// tombstone. Its context has no revision; it must not re-stamp the record.
_, remoteErr := server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
r, err = loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
if !r.Deleted || !r.timestamp().Equal(latest) {
t.Fatalf("sibling re-stamped the tombstone: got %s, want %s", r.timestamp(), latest)
}
create(latest.Add(time.Minute))
_, remoteErr = server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
assertLive(user)
revoke(latest)
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
}
// Counts what a retained revocation actually sends to sibling nodes.
func TestIAMRevocationRetainedReloadsSibling(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
t.Fatal(err)
}
initAllSubsystems(ctx)
globalIAMSys.Init(ctx, obj, nil, 2*time.Second)
defer os.RemoveAll(disks[0])
defer obj.Shutdown(ctx)
defer resetTestGlobals()
sys, peer := globalIAMSys, &globalSiteReplicationSys
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user := "retained-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
deleted, recreated := origin.Add(time.Minute), origin.Add(3*time.Minute)
create := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, at))
}
revoke := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, at))
}
create(origin)
revoke(deleted)
create(recreated)
child, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "retained-child", secretKey: "valid-service-password",
})
must(err)
// A sibling shares persistent state but has an independent IAM cache.
sibling := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, sys.usersSysType)}
must(sibling.LoadIAMCache(ctx, false))
if _, ok := sibling.GetUser(child.AccessKey); !ok {
t.Fatal("sibling fixture did not load child")
}
tg, err := grid.SetupTestGrid(2)
must(err)
t.Cleanup(tg.Cleanup)
var deletes, loads atomic.Int32
server := &peerRESTServer{}
must(deleteUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
deletes.Add(1)
return server.DeleteUserHandler(r)
}))
must(loadUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
loads.Add(1)
// LoadUserHandler delegates to this same cache reload method.
if err := sibling.UserNotificationHandler(ctx, r.Get(peerRESTUser), regUser); err != nil {
return grid.NoPayload{}, grid.NewRemoteErr(err)
}
return grid.NoPayload{}, nil
}))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
must(err)
prev := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) },
}}}
t.Cleanup(func() { globalNotificationSys = prev })
delayed := deleted.Add(time.Minute)
revoke(delayed)
t.Logf("sibling notifications after a retained revocation: destructive=%d reload=%d", deletes.Load(), loads.Load())
if deletes.Load() != 0 {
t.Errorf("destructive sibling delete sent: %d", deletes.Load())
}
if loads.Load() == 0 {
t.Errorf("retained revocation did not notify the sibling")
}
if _, ok := sibling.GetUser(child.AccessKey); ok {
t.Error("sibling still resolves revoked child")
}
if _, ok := sibling.GetUser(user); !ok {
t.Error("sibling lost live parent")
}
if _, ok := sys.store.GetUser(user); !ok {
t.Error("live parent lost")
}
if _, ok := sys.store.GetUser(child.AccessKey); ok {
t.Error("revoked child still resolves on the receiving node")
}
}
+535
View File
@@ -0,0 +1,535 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/namespace"
)
// Exercise the persisted IAM store and the same peer handler used by site heal.
// A delete must survive a cache reload and an older create arriving afterwards.
func TestIAMRevocationRejectsOfflineUser(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "offline-revoked-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "test-password-valid", Status: madmin.AccountEnabled}
created, err := globalIAMSys.CreateUser(ctx, user, req)
if err != nil {
t.Fatal(err)
}
if err = globalIAMSys.DeleteUser(ctx, user, false); err != nil {
t.Fatal(err)
}
if err = globalIAMSys.store.LoadIAMCache(ctx, false); err != nil {
t.Fatal(err)
}
if err = globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, created); err != nil {
t.Fatal(err)
}
if _, err = globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Fatalf("revoked user restored by old peer event: %v", err)
}
}
func TestIAMRevocationHealingContinuesAfterPeerRejectsDelete(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
if _, err := globalIAMSys.CreateUser(ctx, "heal-sync", req); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.CreateUser(ctx, "heal-deleted", req); err != nil {
t.Fatal(err)
}
if err := globalIAMSys.DeleteUser(ctx, "heal-deleted", false); err != nil {
t.Fatal(err)
}
p, err := globalIAMSys.store.GetPolicy("readwrite")
if err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.SetPolicy(ctx, "heal-new-policy", p); err != nil {
t.Fatal(err)
}
var liveUpdates atomic.Int32
peer := func(id string, rejectDelete bool) *httptest.Server {
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
switch {
case strings.HasSuffix(r.URL.Path, "/metainfo"):
_ = json.NewEncoder(w).Encode(madmin.SRInfo{DeploymentID: id})
case r.URL.Path == "/minio/admin/v3/site-replication/peer/iam-revisions":
if r.Method == http.MethodGet {
_ = json.NewEncoder(w).Encode(iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: id, Digest: "fixture"})
return
}
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(http.StatusBadRequest)
return
}
for _, item := range batch.Items {
if rejectDelete && iamDeletionPath(item.SRIAMItem) != "" {
w.WriteHeader(http.StatusForbidden)
_, _ = w.Write([]byte(`{"Code":"AccessDenied","Message":"delete rejected"}`))
return
}
if id == "healthy" && item.Type == madmin.SRIAMItemPolicy && item.Name == "heal-new-policy" && len(item.Policy) > 0 {
liveUpdates.Add(1)
}
}
_ = json.NewEncoder(w).Encode(iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: id, Digest: "fixture"})
default:
t.Errorf("unexpected peer request %s", r.URL.Path)
w.WriteHeader(http.StatusNotFound)
}
}))
}
healthy, rejected := peer("healthy", false), peer("rejected", true)
defer healthy.Close()
defer rejected.Close()
c := &SiteReplicationSys{enabled: true, state: srState{
ServiceAccountAccessKey: "heal-sync",
Peers: map[string]madmin.PeerInfo{
globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()},
"healthy": {Name: "healthy", DeploymentID: "healthy", Endpoint: healthy.URL},
"rejected": {Name: "rejected", DeploymentID: "rejected", Endpoint: rejected.URL},
},
}}
if err := c.healIAMSystem(ctx, obj); err == nil {
t.Fatal("deletion failure was not reported")
}
if liveUpdates.Load() == 0 {
t.Fatal("one peer rejecting a deletion blocked unrelated live IAM healing to a healthy peer")
}
}
func TestIAMRevocationLifecycle(t *testing.T) {
testIAMRevocationLifecycle(t, nil)
}
func TestIAMRevocationEtcdLifecycle(t *testing.T) {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if endpoint == "" {
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
connection, err := etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: 5 * time.Second})
if err != nil {
t.Fatal(err)
}
defer connection.Close()
// The facade borrows the connection's services. Close the owning client,
// not namespace.Watcher while IAM's canceled watch loop is winding down.
ctx, cancel := context.WithCancel(connection.Ctx())
defer cancel()
client := etcd.NewCtxClient(ctx, etcd.WithZapLogger(connection.GetLogger()))
prefix := fmt.Sprintf("/silo-revocation-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(connection.KV, prefix)
client.Watcher = namespace.NewWatcher(connection.Watcher, prefix)
client.Lease = connection.Lease
testIAMRevocationLifecycle(t, client)
}
func testIAMRevocationLifecycle(t *testing.T, client *etcd.Client) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
disk := disks[0]
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err == nil {
initAllSubsystems(ctx)
globalIAMSys.Init(ctx, obj, client, 2*time.Second)
}
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
sys, peer := globalIAMSys, &globalSiteReplicationSys
must := func(t *testing.T, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
reload := func(t *testing.T) { t.Helper(); must(t, sys.store.LoadIAMCache(ctx, false)) }
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
createUser := func(t *testing.T, name string) {
t.Helper()
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, origin))
}
assertAbsent := func(t *testing.T, name string) {
t.Helper()
if _, ok := sys.GetUser(ctx, name); ok {
t.Fatalf("revoked credential %s is usable", name)
}
}
t.Run("replay after recreation", func(t *testing.T) {
testIAMRevocationReplayAfterRecreation(ctx, t, sys)
})
t.Run("origin timestamp and recreation", func(t *testing.T) {
name := "revocation-recreate"
createUser(t, name)
ui, ok := sys.store.GetUser(name)
if !ok || !ui.UpdatedAt.Equal(origin) {
t.Fatalf("origin time changed: %v", ui.UpdatedAt)
}
must(t, sys.DeleteUser(ctx, name, false))
reload(t)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, time.Time{}))
assertAbsent(t, name)
newTime := UTCNow().Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, newTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}, origin.Add(time.Second)))
reload(t)
ui, ok = sys.store.GetUser(name)
if !ok || !ui.UpdatedAt.Equal(newTime) || ui.RevokedBefore.IsZero() {
t.Fatalf("newer recreation lost, or deletion boundary missing: present=%v", ok)
}
})
t.Run("groups policies and mappings", func(t *testing.T) {
user, group, name := "revocation-member", "revocation-group", "revocation-policy"
createUser(t, user)
p, err := sys.store.GetPolicy("readwrite")
must(t, err)
must(t, peer.PeerAddPolicyHandler(ctx, name, &p, origin))
add := &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}
must(t, peer.PeerGroupInfoChangeHandler(ctx, add, origin))
for _, isGroup := range []bool{false, true} {
entity := user
if isGroup {
entity = group
}
mp := &madmin.SRPolicyMapping{UserOrGroup: entity, Policy: name, UserType: int(regUser), IsGroup: isGroup}
must(t, peer.PeerPolicyMappingHandler(ctx, mp, origin))
_, err = sys.PolicyDBSet(ctx, entity, "", regUser, isGroup)
must(t, err)
must(t, peer.PeerPolicyMappingHandler(ctx, mp, origin))
if _, ok := sys.store.GetMappedPolicy(entity, isGroup); ok {
t.Fatal("old grant restored")
}
}
// This receiver never saw the member-removal event preceding deletion.
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, IsRemove: true}}, UTCNow()))
must(t, sys.DeletePolicy(ctx, name, true))
reload(t)
must(t, peer.PeerGroupInfoChangeHandler(ctx, add, origin))
must(t, peer.PeerAddPolicyHandler(ctx, name, &p, origin))
if _, err = sys.GetGroupDescription(group); !errors.Is(err, errNoSuchGroup) {
t.Fatalf("group restored: %v", err)
}
if _, err = sys.store.GetPolicyDoc(name); !errors.Is(err, errNoSuchPolicy) {
t.Fatalf("policy restored: %v", err)
}
paths, err := sys.store.listIAMConfigPaths(ctx)
must(t, err)
found := make(map[string]bool)
for _, path := range paths {
r, err := loadIAMRevision(ctx, sys.store, path)
must(t, err)
if item, ok := iamDeletionItem(path, r); ok {
found[iamDeletionPath(item)] = true
if item.UpdatedAt.IsZero() {
t.Fatal("undated delete replay")
}
}
}
for _, path := range []string{getGroupInfoPath(group), getPolicyDocPath(name), getMappedPolicyPath(user, regUser, false), getMappedPolicyPath(group, regUser, true)} {
if !found[path] {
t.Errorf("deletion missing from heal: %s", path)
}
}
})
t.Run("parent revokes service accounts and STS", func(t *testing.T) {
parent := "revocation-parent"
createUser(t, parent)
svc, svcAt, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: "revocation-service", secretKey: "valid-service-password"})
must(t, err)
secret, err := getTokenSigningKey()
must(t, err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
sts.ParentUser = parent
_, err = sys.SetTempUser(withIAMReplicationTime(ctx, origin), sts.AccessKey, sts, "readwrite")
must(t, err)
must(t, sys.DeleteUser(ctx, parent, false))
reload(t)
assertAbsent(t, parent)
assertAbsent(t, svc.AccessKey)
assertAbsent(t, sts.AccessKey)
// Recreate the parent, then deliver old child events from the offline site.
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, UTCNow()))
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: svc.AccessKey, SecretKey: svc.SecretKey}}, svcAt))
must(t, peer.PeerSTSAccHandler(ctx, &madmin.SRSTSCredential{AccessKey: sts.AccessKey, SecretKey: sts.SecretKey, ParentUser: parent, SessionToken: sts.SessionToken, ParentPolicyMapping: "readwrite"}, origin))
reload(t)
assertAbsent(t, svc.AccessKey)
assertAbsent(t, sts.AccessKey)
// A freshly issued credential is still supported after deliberate recreation.
_, _, err = sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "new-service", secretKey: "valid-service-password"})
must(t, err)
if _, ok := sys.GetUser(ctx, "new-service"); !ok {
t.Fatal("fresh service account rejected")
}
})
t.Run("delete before first create", func(t *testing.T) {
name := "revocation-unseen"
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}, UTCNow()))
createUser(t, name)
assertAbsent(t, name)
svc := "unseen-service"
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: svc}}, UTCNow()))
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "revocation-recreate", AccessKey: svc, SecretKey: "valid-service-password"}}, origin))
assertAbsent(t, svc)
})
t.Run("recreation arrives before revocation", func(t *testing.T) {
parent := "reordered-parent"
createUser(t, parent)
child, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: "reordered-child", secretKey: "valid-service-password"})
must(t, err)
newTime, deleteTime := UTCNow(), origin.Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, newTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
assertAbsent(t, child.AccessKey)
reload(t)
assertAbsent(t, child.AccessKey)
u, ok := sys.GetUser(ctx, parent)
if !ok || !u.UpdatedAt.Equal(newTime) || !u.RevokedBefore.Equal(deleteTime) {
t.Fatal("reordered revocation damaged the new parent or lost its boundary")
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(parent, regUser))
must(t, err)
item, ok := iamDeletionItem(getUserIdentityPath(parent, regUser), r)
if !ok || !item.UpdatedAt.Equal(deleteTime) {
t.Fatal("recreation erased deletion replay")
}
})
t.Run("user cleanup does not supersede group deletion", func(t *testing.T) {
user, group := "cascade-user", "cascade-group"
createUser(t, user)
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}, origin))
// On the origin site the group was removed before the user, but the
// recovering receiver processes those independent events in reverse.
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, IsRemove: true}}, origin.Add(time.Minute)))
reload(t)
if _, err := sys.GetGroupDescription(group); !errors.Is(err, errNoSuchGroup) {
t.Fatalf("deleted group survived reordered cleanup: %v", err)
}
groups, err := sys.ListGroups(ctx)
must(t, err)
for _, name := range groups {
if name == group {
t.Fatal("deleted group listed")
}
}
})
t.Run("parent revocation covers later updates to existing children", func(t *testing.T) {
parent, key := "late-update-parent", "late-update-child"
createUser(t, parent)
_, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: key, secretKey: "valid-service-password"})
must(t, err)
// This site missed the deletion and subsequently edited an old child.
_, err = sys.UpdateServiceAccount(withIAMReplicationTime(ctx, origin.Add(2*time.Minute)), key, updateServiceAccountOpts{description: "edited while the peer was offline"})
must(t, err)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, origin.Add(time.Minute)))
assertAbsent(t, key)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, origin.Add(3*time.Minute)))
reload(t)
assertAbsent(t, key)
})
t.Run("old generation cannot return with a newer event timestamp", func(t *testing.T) {
parent := "generation-parent"
createUser(t, parent)
deleteTime := origin.Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, origin.Add(2*time.Minute)))
// Another offline site issued this child under the original parent,
// after this site's delete/recreate. Wall-clock ordering cannot identify it.
late := origin.Add(3 * time.Minute)
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: "old-gen-service", SecretKey: "valid-service-password"}}, late))
secret, err := getTokenSigningKey()
must(t, err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
must(t, peer.PeerSTSAccHandler(ctx, &madmin.SRSTSCredential{AccessKey: sts.AccessKey, SecretKey: sts.SecretKey, ParentUser: parent, SessionToken: sts.SessionToken}, late))
reload(t)
assertAbsent(t, "old-gen-service")
assertAbsent(t, sts.AccessKey)
// A local issuer knows the new boundary and signs it into both kinds
// of child. Untrusted inherited claims cannot select that boundary.
child, _, err := sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "new-gen-service", secretKey: "valid-service-password", claims: map[string]any{iamParentRevocationClaim: "forged"}})
must(t, err)
newClaims := map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}
must(t, setIAMParentRevocationClaim(ctx, sys.store, parent, newClaims))
fresh, err := auth.GetNewCredentialsWithMetadata(newClaims, secret)
must(t, err)
fresh.ParentUser = parent
_, err = sys.SetTempUser(ctx, fresh.AccessKey, fresh, "")
must(t, err)
reload(t)
// A periodic reload retains the STS cache. Explicitly clear it to
// exercise the cold credential load performed after process restart.
cache := sys.store.lock()
cache.iamSTSAccountsMap = make(map[string]UserIdentity)
sys.store.unlock()
for _, key := range []string{child.AccessKey, fresh.AccessKey} {
u, ok := sys.GetUser(ctx, key)
if !ok || !iamCredentialSurvivesRevocation(u.Credentials, deleteTime) {
t.Fatalf("new-generation credential %s rejected", key)
}
}
})
t.Run("late revocation preserves proven new-generation children", func(t *testing.T) {
for _, recreateFirst := range []bool{false, true} {
parent := fmt.Sprintf("gen-parent-%t", recreateFirst)
createUser(t, parent)
deleteTime, createTime := origin.Add(time.Minute), origin.Add(2*time.Minute)
if recreateFirst {
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, createTime))
}
key := fmt.Sprintf("gen-child-%t", recreateFirst)
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: key, SecretKey: "valid-service-password", Claims: map[string]any{iamParentRevocationClaim: deleteTime.Format(time.RFC3339Nano)}}}, createTime.Add(time.Second)))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
if !recreateFirst {
assertAbsent(t, key)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, createTime))
}
reload(t)
if _, ok := sys.GetUser(ctx, key); !ok {
t.Fatal("late revocation deleted a child issued by the recreated parent")
}
}
})
t.Run("cold loading preserves site-signed STS", func(t *testing.T) {
parent := "cold-sts-parent"
createUser(t, parent)
secret := "site-signing-key-valid"
_, _, err := sys.NewServiceAccount(ctx, globalActiveCred.AccessKey, nil, newServiceAccountOpts{
accessKey: siteReplicatorSvcAcc, secretKey: secret, allowSiteReplicatorAccount: true,
})
must(t, err)
setReplication := func(enabled bool) {
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = enabled
globalSiteReplicationSys.Unlock()
globalSiteReplicatorCred.Set("")
}
setReplication(true)
defer setReplication(false)
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
cred.ParentUser = parent
_, err = sys.SetTempUser(ctx, cred.AccessKey, cred, "")
must(t, err)
// IAM can load before the site replication manager during startup.
// A signing key that is not available yet must not delete live tokens.
setReplication(false)
for range 3 {
unverified := make(map[string]UserIdentity)
_ = sys.store.loadUser(ctx, cred.AccessKey, stsUser, unverified)
if _, ok := unverified[cred.AccessKey]; ok {
t.Fatal("accepted STS before the signing key became available")
}
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(cred.AccessKey, stsUser))
must(t, err)
if r.Credentials.SessionToken == "" {
t.Fatal("cold IAM load physically deleted a non-expired site-signed STS credential")
}
setReplication(true)
loaded := make(map[string]UserIdentity)
must(t, sys.store.loadUser(ctx, cred.AccessKey, stsUser, loaded))
if _, ok := loaded[cred.AccessKey]; !ok {
t.Fatal("STS credential did not recover when the signing key became available")
}
})
t.Run("unverifiable STS stay denied and expired STS are removed", func(t *testing.T) {
parent := "invalid-sts-parent"
createUser(t, parent)
// Keep the etcd watcher from cleaning half of the fixture before the
// second record is seeded; this subtest exercises the loader directly.
sys.store.lock()
defer sys.store.unlock()
for _, expired := range []bool{false, true} {
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, "unavailable-test-signing-key")
must(t, err)
cred.ParentUser = parent
if expired {
cred.Expiration = UTCNow().Add(-time.Minute)
}
identityPath := getUserIdentityPath(cred.AccessKey, stsUser)
mappingPath := getMappedPolicyPath(cred.AccessKey, stsUser, false)
// Seed disk directly to exercise loading, including existing records
// whose key is unknown. The write API should not accept such tokens.
must(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: cred, UpdatedAt: UTCNow()}, identityPath))
must(t, sys.store.saveIAMConfig(ctx, &MappedPolicy{Version: 1, Policies: "readwrite"}, mappingPath))
loaded := make(map[string]UserIdentity)
_ = sys.store.loadUser(ctx, cred.AccessKey, stsUser, loaded)
if _, ok := loaded[cred.AccessKey]; ok {
t.Fatalf("invalid STS accepted, expired=%t", expired)
}
for _, path := range []string{identityPath, mappingPath} {
var record map[string]any
err := sys.store.loadIAMConfig(ctx, &record, path)
if expired {
if !errors.Is(err, errConfigNotFound) {
t.Fatalf("expired STS data not cleaned up at %s: %v", path, err)
}
} else {
must(t, err)
}
}
}
})
}
+239
View File
@@ -0,0 +1,239 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
)
// The receiver missed a deletion and still has the previous service key.
// A later full snapshot must replace it, including when the owner changed.
func TestIAMServiceAccountRecreation(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
for _, parent := range []string{"old-owner", "new-owner"} {
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), parent, madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
}
const key = "reusable-service"
old := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "old-owner", AccessKey: key, SecretKey: "old-service-password"}}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
oldIdentity, _ := sys.store.GetUser(key)
_, err := sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), key, "readwrite", svcUser, false)
mustIAM(t, err)
boundary, newer := origin.Add(time.Minute), origin.Add(2*time.Minute)
fresh := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemSvcAcc, UpdatedAt: newer, SvcAccChange: &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "new-owner", AccessKey: key, SecretKey: "new-service-password", Status: auth.AccountOff}}}, RevokedBefore: boundary}
mustIAM(t, applyIAMReplicationItem(ctx, fresh))
// A sibling may have missed the notification of the committed
// replacement. An equal-version retry must refresh that cache too.
staleCache := sys.store.lock()
staleCache.iamUsersMap[key] = oldIdentity
sys.store.unlock()
mustIAM(t, applyIAMReplicationItem(ctx, fresh)) // duplicate delivery is acknowledged
if current, _ := sys.store.GetUser(key); current.Credentials.SecretKey != fresh.SvcAccChange.Create.SecretKey {
t.Fatal("duplicate snapshot acknowledged without refreshing the stale cache")
}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: key}}, boundary))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
u, ok := sys.store.GetUser(key)
if !ok || u.Credentials.SecretKey != fresh.SvcAccChange.Create.SecretKey || u.Credentials.ParentUser != "new-owner" || u.Credentials.Status != auth.AccountOff || !u.UpdatedAt.Equal(newer) || !u.RevokedBefore.Equal(boundary) {
t.Fatal("recreation did not retain the new identity, disabled status, source version and revocation")
}
if _, ok := sys.GetUser(ctx, key); ok {
t.Fatal("replicated disabled service can authenticate")
}
cache := sys.store.rlock()
_, mapped := cache.cachedMappedPolicy(key, svcUser, false)
sys.store.runlock()
if mapped {
t.Fatal("recreated service inherited an older mapping")
}
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), key, "readwrite", svcUser, false)
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("old service mapping replay was accepted: %v", err)
}
_, _, err = sys.NewServiceAccount(ctx, "new-owner", nil, newServiceAccountOpts{accessKey: key, secretKey: "local-service-password"})
if !errors.Is(err, errIAMServiceAccountNotAllowed) {
t.Fatalf("local duplicate creation must remain rejected: %v", err)
}
// Outbound snapshots must carry the retained service boundary too.
out, err := globalSiteReplicationSys.replicationItem(ctx, fresh.SRIAMItem)
mustIAM(t, err)
if !out.RevokedBefore.Equal(boundary) {
t.Fatal("outbound service snapshot lost its revocation")
}
})
}
}
func TestIAMServiceAccountReplicationRejectsOtherCredentialKinds(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "builtin-collision", madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
secret, err := getTokenSigningKey()
mustIAM(t, err)
token, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: "builtin-collision"}, secret)
mustIAM(t, err)
token.ParentUser = "builtin-collision"
_, err = sys.SetTempUser(ctx, token.AccessKey, token, "")
mustIAM(t, err)
for _, key := range []string{"builtin-collision", token.AccessKey} {
_, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, UTCNow().Add(time.Minute)), "another-owner", nil, newServiceAccountOpts{accessKey: key, secretKey: "valid-service-password"})
if !errors.Is(err, errIAMServiceAccountNotAllowed) {
t.Fatalf("service replication replaced another credential kind: %v", err)
}
}
}
// SR configuration can be temporarily unreadable even though a service token
// is signed with its own valid secret. Do not acknowledge a failed cache load.
func TestIAMServiceAccountRetryReportsClaimLoadFailure(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
globalSiteReplicatorCred.RLock()
previousSigningKey := globalSiteReplicatorCred.secretKey
globalSiteReplicatorCred.RUnlock()
globalSiteReplicatorCred.Set("")
t.Cleanup(func() { globalSiteReplicatorCred.Set(previousSigningKey) })
_, err := sys.CreateUser(ctx, "retry-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
opts := newServiceAccountOpts{accessKey: "retry-service", secretKey: "valid-service-password"}
_, _, err = sys.NewServiceAccount(ctx, "retry-owner", nil, opts)
mustIAM(t, err)
old, _ := sys.store.GetUser(opts.accessKey)
opts.secretKey = "replacement-service-password"
at, err := sys.UpdateServiceAccount(ctx, opts.accessKey, updateServiceAccountOpts{secretKey: opts.secretKey})
mustIAM(t, err)
cache := sys.store.lock()
cache.iamUsersMap[opts.accessKey] = old // Missed sibling notification.
sys.store.unlock()
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = true // No site-replicator credential is installed.
globalSiteReplicationSys.Unlock()
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, at), "retry-owner", nil, opts)
if err == nil {
t.Fatal("acknowledged service retry despite failed claims loading")
}
if _, ok := sys.store.GetUser(opts.accessKey); ok {
t.Fatal("failed cache refresh retained the superseded service secret")
}
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = false
globalSiteReplicationSys.Unlock()
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, at), "retry-owner", nil, opts)
mustIAM(t, err)
}
// A delayed snapshot still has its original absolute expiration. Reapplying
// the local minimum issuance lifetime would leave the old unexpired key alive.
func TestIAMServiceAccountReplicationPreservesExpiration(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
for _, action := range []string{"create", "update"} {
for _, expired := range []bool{false, true} {
name := backend + "/" + action + "/near_expiry"
if expired {
name = backend + "/" + action + "/expired"
}
t.Run(name, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
_, err := sys.CreateUser(ctx, "expiry-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
old := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "expiry-owner", AccessKey: "expiry-service", SecretKey: "old-service-password"}}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
expires := UTCNow().Add(time.Minute)
if expired {
expires = UTCNow().Add(-time.Minute)
}
change := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "expiry-owner", AccessKey: "expiry-service", SecretKey: "new-service-password", Expiration: &expires}}
if action == "update" {
change = &madmin.SRSvcAccChange{Update: &madmin.SRSvcAccUpdate{AccessKey: "expiry-service", SecretKey: "new-service-password", Expiration: &expires}}
}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, change, origin.Add(2*time.Minute)))
u, ok := sys.store.GetUser("expiry-service")
if !ok || u.Credentials.SecretKey != "new-service-password" || !u.Credentials.Expiration.Equal(expires) {
t.Fatal("delayed snapshot lost its new secret or absolute expiration")
}
_, allowed := sys.GetUser(ctx, "expiry-service")
if allowed == expired {
t.Fatal("credential validity disagrees with its absolute expiration")
}
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath("expiry-service", svcUser))
mustIAM(t, err)
if r.Credentials.SecretKey == old.Create.SecretKey || (expired && !r.Deleted) {
t.Fatal("old non-expiring credential returned after reload")
}
_, _, err = sys.NewServiceAccount(ctx, "expiry-owner", nil, newServiceAccountOpts{accessKey: "local-expiry", secretKey: "valid-service-password", expiration: &expires})
if !errors.Is(err, errInvalidSvcAcctExpiration) {
t.Fatalf("local issuance lifetime check changed: %v", err)
}
})
}
}
}
}
// Status-only summaries intentionally omit secrets. Different revisions must
// still trigger live healing, and disabled identities must be eligible sources.
func TestIAMServiceAccountHealingNewerSnapshot(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
_, err := sys.CreateUser(ctx, "heal-svc-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), "heal-svc-owner", nil, newServiceAccountOpts{accessKey: "heal-service", secretKey: "valid-service-password"})
mustIAM(t, err)
at, err := sys.UpdateServiceAccount(ctx, "heal-service", updateServiceAccountOpts{status: auth.AccountOff})
mustIAM(t, err)
var sent atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
return
}
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" || r.Method != http.MethodPut {
t.Errorf("unexpected request %s %s", r.Method, r.URL.Path)
w.WriteHeader(http.StatusNotFound)
return
}
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(http.StatusBadRequest)
return
}
for _, item := range batch.Items {
if item.SvcAccChange != nil && item.SvcAccChange.Create != nil && item.SvcAccChange.Create.AccessKey == "heal-service" && item.SvcAccChange.Create.Status == auth.AccountOff && item.UpdatedAt.Equal(at) {
sent.Add(1)
}
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "node", Instance: "boot", Digest: "fixture"}})
}))
defer server.Close()
peers := map[string]madmin.PeerInfo{globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()}, "remote": {Name: "remote", DeploymentID: "remote", Endpoint: server.URL}}
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "heal-svc-owner", Peers: peers}}
local := madmin.UserInfo{Status: madmin.AccountStatus(auth.AccountOff), UpdatedAt: at}
remote := local
remote.UpdatedAt = origin
if isUserInfoReplicated(2, 2, []madmin.UserInfo{local, remote}) {
t.Fatal("status-only summaries concealed different service revisions")
}
info := srStatusInfo{Sites: peers, UserStats: map[string]map[string]srUserStatsSummary{"heal-service": {globalDeploymentID(): {userInfo: srUserInfo{UserInfo: local}}, "remote": {SRUserStatsSummary: madmin.SRUserStatsSummary{UserInfoMismatch: true}, userInfo: srUserInfo{UserInfo: remote}}}}}
mustIAM(t, c.healUsers(ctx, obj, "heal-service", info))
if sent.Load() == 0 {
t.Fatal("disabled newer service snapshot was not healed")
}
}
+383 -221
View File
File diff suppressed because it is too large Load Diff
+65 -21
View File
@@ -162,6 +162,15 @@ func (sys *IAMSys) LoadUser(ctx context.Context, objAPI ObjectLayer, accessKey s
return sys.store.UserNotificationHandler(ctx, accessKey, userType)
}
// LoadUserAfterDelete reloads a parent's identity and cached dependents after a
// sibling committed a deletion. Each record may already have been recreated.
func (sys *IAMSys) LoadUserAfterDelete(ctx context.Context, accessKey string) error {
if !sys.Initialized() {
return errServerNotInitialized
}
return sys.store.UserDeletionNotificationHandler(ctx, accessKey)
}
// LoadServiceAccount - reloads a specific service account from backend disks or etcd.
func (sys *IAMSys) LoadServiceAccount(ctx context.Context, accessKey string) error {
if !sys.Initialized() {
@@ -596,10 +605,22 @@ func (sys *IAMSys) DeletePolicy(ctx context.Context, policyName string, notifyPe
return errServerNotInitialized
}
for _, v := range policy.DefaultPolicies {
if v.Name == policyName {
if err := checkConfig(ctx, globalObjectAPI, getPolicyDocPath(policyName)); err != nil && err == errConfigNotFound {
return fmt.Errorf("inbuilt policy `%s` not allowed to be deleted", policyName)
if _, replicated := iamReplicationTime(ctx); !replicated && notifyPeers {
for _, v := range policy.DefaultPolicies {
if v.Name == policyName {
var err error
if objectStore, ok := sys.store.IAMStorageAPI.(*IAMObjectStore); ok {
err = checkConfig(ctx, objectStore.objAPI, getPolicyDocPath(policyName))
} else {
var r iamRevision
err = sys.store.loadIAMConfig(ctx, &r, getPolicyDocPath(policyName))
}
if errors.Is(err, errConfigNotFound) {
return fmt.Errorf("inbuilt policy `%s` not allowed to be deleted", policyName)
}
if err != nil {
return err
}
}
}
}
@@ -705,20 +726,30 @@ func (sys *IAMSys) DeleteUser(ctx context.Context, accessKey string, notifyPeers
return errServerNotInitialized
}
if err := sys.store.DeleteUser(ctx, accessKey, regUser); err != nil {
err := sys.store.DeleteUser(ctx, accessKey, regUser)
var cleanupErr *iamCommittedCleanupError
retained := errors.Is(err, errIAMRevocationRetained)
if errors.As(err, &cleanupErr) {
retained = cleanupErr.retained
} else if err != nil && !retained {
return err
}
// Notify all other MinIO peers to delete user.
// Publish the committed state even when dependent cleanup must be retried.
if notifyPeers && !sys.HasWatcher() {
for _, nerr := range globalNotificationSys.DeleteUser(ctx, accessKey) {
if nerr.Err != nil {
logger.GetReqInfo(ctx).SetTags("peerAddress", nerr.Host.String())
iamLogIf(ctx, nerr.Err)
if retained {
sys.notifyForUser(ctx, accessKey, false)
} else {
for _, nerr := range globalNotificationSys.DeleteUser(ctx, accessKey) {
if nerr.Err != nil {
logger.GetReqInfo(ctx).SetTags("peerAddress", nerr.Host.String())
iamLogIf(ctx, nerr.Err)
}
}
}
}
if cleanupErr != nil {
return cleanupErr
}
return nil
}
@@ -1052,6 +1083,7 @@ type newServiceAccountOpts struct {
sessionPolicy *policy.Policy
accessKey string
secretKey string
status string // Used by replication snapshots; local creates default to enabled.
name, description string
expiration *time.Time
allowSiteReplicatorAccount bool // allow creating internal service account for site-replication.
@@ -1116,6 +1148,11 @@ func (sys *IAMSys) NewServiceAccount(ctx context.Context, parentUser string, gro
m[k] = v
}
}
if _, replicated := iamReplicationTime(ctx); !replicated {
if err := setIAMParentRevocationClaim(ctx, sys.store, parentUser, m); err != nil {
return auth.Credentials{}, time.Time{}, err
}
}
var accessKey, secretKey string
var err error
@@ -1134,12 +1171,19 @@ func (sys *IAMSys) NewServiceAccount(ctx context.Context, parentUser string, gro
cred.ParentUser = parentUser
cred.Groups = groups
cred.Status = string(auth.AccountOn)
switch opts.status {
case "", auth.AccountOn, string(madmin.AccountEnabled):
case auth.AccountOff, string(madmin.AccountDisabled):
cred.Status = auth.AccountOff
default:
return auth.Credentials{}, time.Time{}, errInvalidArgument
}
cred.Name = opts.name
cred.Description = opts.description
if opts.expiration != nil {
expirationInUTC := opts.expiration.UTC()
if err := validateSvcExpirationInUTC(expirationInUTC); err != nil {
if err := validateSvcExpirationInUTC(ctx, expirationInUTC); err != nil {
return auth.Credentials{}, time.Time{}, err
}
cred.Expiration = expirationInUTC
@@ -1370,7 +1414,7 @@ func (sys *IAMSys) DeleteServiceAccount(ctx context.Context, accessKey string, n
}
sa, ok := sys.store.GetUser(accessKey)
if !ok || !sa.Credentials.IsServiceAccount() {
if _, replicated := iamReplicationTime(ctx); (!ok || !sa.Credentials.IsServiceAccount()) && !replicated {
return nil
}
@@ -1474,8 +1518,8 @@ func (sys *IAMSys) purgeExpiredCredentialsForExternalSSO(ctx context.Context) {
}
}
// We ignore any errors
_ = sys.store.DeleteUsers(ctx, expiredUsers)
// Keep failed revocations visible so the next purge can retry.
iamLogIf(ctx, sys.store.DeleteUsers(ctx, expiredUsers))
}
// purgeExpiredCredentialsForLDAP - validates if local credentials are still
@@ -1503,8 +1547,8 @@ func (sys *IAMSys) purgeExpiredCredentialsForLDAP(ctx context.Context) {
return
}
// We ignore any errors
_ = sys.store.DeleteUsers(ctx, expiredUsers)
// Keep failed revocations visible so the next purge can retry.
iamLogIf(ctx, sys.store.DeleteUsers(ctx, expiredUsers))
}
// updateGroupMembershipsForLDAP - updates the list of groups associated with the credential.
@@ -1925,12 +1969,12 @@ func (sys *IAMSys) RemoveUsersFromGroup(ctx context.Context, group string, membe
}
updatedAt, err = sys.store.RemoveUsersFromGroup(ctx, group, members)
if err != nil {
var cleanupErr *iamCommittedCleanupError
if err != nil && !errors.As(err, &cleanupErr) {
return updatedAt, err
}
sys.notifyForGroup(ctx, group)
return updatedAt, nil
return updatedAt, err
}
// SetGroupStatus - enable/disabled a group
+24 -2
View File
@@ -992,6 +992,26 @@ func getClusterReplMRFFailedOperationsMD() MetricDescription {
}
}
func getClusterReplMRFDroppedOperationsMD() MetricDescription {
return MetricDescription{
Namespace: nodeMetricNamespace,
Subsystem: replicationSubsystem,
Name: "mrf_dropped_operations_total",
Help: "Total number of replication MRF entries dropped since server start; entries may refer to the same object",
Type: counterMetric,
}
}
func getClusterReplMRFDroppedBytesMD() MetricDescription {
return MetricDescription{
Namespace: nodeMetricNamespace,
Subsystem: replicationSubsystem,
Name: "mrf_dropped_bytes_total",
Help: "Total known bytes of replication MRF entries dropped since server start; delete entries count as zero bytes",
Type: counterMetric,
}
}
func getClusterRepCredentialErrorsMD(namespace MetricNamespace) MetricDescription {
return MetricDescription{
Namespace: namespace,
@@ -2423,6 +2443,8 @@ func getReplicationNodeMetrics(opts MetricsGroupOpts) *MetricsGroupV2 {
avgTransferRate,
maxTransferRate,
mrfCount,
{Description: getClusterReplMRFDroppedOperationsMD(), Value: float64(qs.MRFStats.TotalDroppedCount)},
{Description: getClusterReplMRFDroppedBytesMD(), Value: float64(qs.MRFStats.TotalDroppedBytes)},
}
}
for ep, health := range globalBucketTargetSys.healthStats() {
@@ -3314,10 +3336,10 @@ func getBucketUsageMetrics(opts MetricsGroupOpts) *MetricsGroupV2 {
VariableLabels: map[string]string{"bucket": bucket},
})
if quota != nil && quota.Quota > 0 {
if quotaSize := getBucketQuotaSize(quota); quotaSize > 0 {
metrics = append(metrics, MetricV2{
Description: getBucketUsageQuotaTotalBytesMD(),
Value: float64(quota.Quota),
Value: float64(quotaSize),
VariableLabels: map[string]string{"bucket": bucket},
})
}
+14
View File
@@ -34,6 +34,10 @@ const (
sinceLastSyncMillis = "since_last_sync_millis"
syncFailures = "sync_failures"
syncSuccesses = "sync_successes"
revocationRecords = "revocation_records"
revocationHealFailures = "revocation_heal_failures"
revocationHealDurationMillis = "revocation_heal_duration_millis"
revocationHealLastSuccess = "revocation_heal_last_success_timestamp_seconds"
)
var (
@@ -47,10 +51,20 @@ var (
sinceLastSyncMillisMD = NewCounterMD(sinceLastSyncMillis, "Time (in milliseconds) since last successful IAM data sync.")
syncFailuresMD = NewCounterMD(syncFailures, "Number of failed IAM data syncs since server start.")
syncSuccessesMD = NewCounterMD(syncSuccesses, "Number of successful IAM data syncs since server start.")
revocationRecordsMD = NewGaugeMD(revocationRecords, "Retained IAM deletion records and revocation boundaries in this node's index.")
revocationHealFailuresMD = NewCounterMD(revocationHealFailures, "Failed IAM revocation convergence passes since server start.")
revocationHealDurationMillisMD = NewGaugeMD(revocationHealDurationMillis, "Duration of the last IAM revocation convergence pass in milliseconds.")
revocationHealLastSuccessMD = NewGaugeMD(revocationHealLastSuccess, "Unix timestamp of the last successful IAM revocation convergence pass.")
)
// loadClusterIAMMetrics - `MetricsLoaderFn` for cluster IAM metrics.
func loadClusterIAMMetrics(_ context.Context, m MetricValues, _ *metricsCache) error {
if globalIAMSys.Initialized() {
m.Set(revocationRecords, float64(globalIAMSys.store.revisionIndex().count()))
}
m.Set(revocationHealFailures, float64(globalSiteReplicationSys.iamRevisionMetrics.healFailures.Load()))
m.Set(revocationHealDurationMillis, float64(globalSiteReplicationSys.iamRevisionMetrics.healDurationMillis.Load()))
m.Set(revocationHealLastSuccess, float64(globalSiteReplicationSys.iamRevisionMetrics.healLastSuccess.Load()))
m.Set(lastSyncDurationMillis, float64(atomic.LoadUint64(&globalIAMSys.LastRefreshDurationMilliseconds)))
pluginAuthNMetrics := globalAuthNPlugin.Metrics()
m.Set(pluginAuthnServiceFailedRequestsMinute, float64(pluginAuthNMetrics.FailedRequests))
+2 -2
View File
@@ -167,8 +167,8 @@ func loadClusterUsageBucketMetrics(ctx context.Context, m MetricValues, c *metri
m.Set(usageBucketVersionsCount, float64(usage.VersionsCount), "bucket", bucket)
m.Set(usageBucketDeleteMarkersCount, float64(usage.DeleteMarkersCount), "bucket", bucket)
if quota != nil && quota.Quota > 0 {
m.Set(usageBucketQuotaTotalBytes, float64(quota.Quota), "bucket", bucket)
if quotaSize := getBucketQuotaSize(quota); quotaSize > 0 {
m.Set(usageBucketQuotaTotalBytes, float64(quotaSize), "bucket", bucket)
}
for k, v := range usage.ObjectSizesHistogram {
+8
View File
@@ -35,6 +35,8 @@ const (
replicationMaxQueuedCount = "max_queued_count"
replicationMaxDataTransferRate = "max_data_transfer_rate"
replicationRecentBacklogCount = "recent_backlog_count"
replicationMRFDroppedOperations = "mrf_dropped_operations_total"
replicationMRFDroppedBytes = "mrf_dropped_bytes_total"
)
var (
@@ -64,6 +66,10 @@ var (
"Maximum replication data transfer rate in bytes/sec seen since server start")
replicationRecentBacklogCountMD = NewGaugeMD(replicationRecentBacklogCount,
"Total number of objects seen in replication backlog in the last 5 minutes")
replicationMRFDroppedOperationsMD = NewCounterMD(replicationMRFDroppedOperations,
"Total number of replication MRF entries dropped since server start; entries may refer to the same object")
replicationMRFDroppedBytesMD = NewCounterMD(replicationMRFDroppedBytes,
"Total known bytes of replication MRF entries dropped since server start; delete entries count as zero bytes")
)
// loadClusterReplicationMetrics - `MetricsLoaderFn` for cluster replication metrics
@@ -96,6 +102,8 @@ func loadClusterReplicationMetrics(ctx context.Context, m MetricValues, c *metri
m.Set(replicationMaxDataTransferRate, tots.Peak)
}
m.Set(replicationRecentBacklogCount, float64(qs.MRFStats.LastFailedCount))
m.Set(replicationMRFDroppedOperations, float64(qs.MRFStats.TotalDroppedCount))
m.Set(replicationMRFDroppedBytes, float64(qs.MRFStats.TotalDroppedBytes))
return nil
}
+6
View File
@@ -323,6 +323,10 @@ func newMetricGroups(r *prometheus.Registry) *metricsV3Collection {
sinceLastSyncMillisMD,
syncFailuresMD,
syncSuccessesMD,
revocationRecordsMD,
revocationHealFailuresMD,
revocationHealDurationMillisMD,
revocationHealLastSuccessMD,
},
loadClusterIAMMetrics,
)
@@ -342,6 +346,8 @@ func newMetricGroups(r *prometheus.Registry) *metricsV3Collection {
replicationMaxQueuedCountMD,
replicationMaxDataTransferRateMD,
replicationRecentBacklogCountMD,
replicationMRFDroppedOperationsMD,
replicationMRFDroppedBytesMD,
},
loadClusterReplicationMetrics,
)
+6 -2
View File
@@ -99,9 +99,13 @@ type ObjectOptions struct {
ReplicationSourceTaggingTimestamp time.Time // set if MinIOSourceTaggingTimestamp received
ReplicationSourceLegalholdTimestamp time.Time // set if MinIOSourceObjectLegalholdTimestamp received
ReplicationSourceRetentionTimestamp time.Time // set if MinIOSourceObjectRetentionTimestamp received
ReplicaLockReconcile bool // re-order a trusted replica write against the destination version under the object write lock
DeletePrefix bool // set true to enforce a prefix deletion, only application for DeleteObject API,
DeletePrefixObject bool // set true when object's erasure set is resolvable by object name (using getHashedSetIndex)
// Pools-layer version resolver; the caller holds the shared object lock.
replicaObjectInfo GetObjectInfoFn
Speedtest bool // object call specifically meant for SpeedTest code, set to 'true' when invoked by SpeedtestHandler.
// Use the maximum parity (N/2), used when saving server configuration files
@@ -131,8 +135,8 @@ type ObjectOptions struct {
// when looking up a version by fi.VersionID
InclFreeVersions bool
// SkipFreeVersion skips adding a free version when a tiered version is
// being 'replaced'
// Note: Used only when a tiered object is being expired.
// being replaced. Used when expiring tiered content or retiring a copy
// whose tier reference is still owned by another copy.
SkipFreeVersion bool
MetadataChg bool // is true if it is a metadata update operation.
+113
View File
@@ -0,0 +1,113 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"encoding/base64"
"net/http"
"reflect"
"strconv"
"strings"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
)
func TestPutOptsFromHeadersReplicationTimestamps(t *testing.T) {
stamp := time.Date(2026, 9, 15, 1, 2, 3, 123456789, time.UTC)
context := base64.StdEncoding.EncodeToString([]byte(`{"purpose":"tag-replication"}`))
for _, encryption := range []struct {
name string
headers map[string]string
}{
{name: "none"},
{name: "SSE-S3", headers: map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}},
{name: "SSE-KMS", headers: map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS}},
{name: "SSE-KMS-context", headers: map[string]string{
xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS, xhttp.AmzServerSideEncryptionKmsID: "tag-replication-key",
xhttp.AmzServerSideEncryptionKmsContext: context,
}},
{name: "SSE-C", headers: ssecKeyHeaders([]byte("01234567890123456789012345678901"), false)},
} {
t.Run(encryption.name, func(t *testing.T) {
for _, trusted := range []bool{false, true} {
t.Run("trusted="+strconv.FormatBool(trusted), func(t *testing.T) {
for _, tagging := range []struct {
name, header string
want time.Time
invalid bool
}{
{name: "absent"},
{name: "nanoseconds", header: stamp.Format(time.RFC3339Nano), want: stamp},
{name: "offset-whitespace", header: " " + stamp.In(time.FixedZone("UTC+8", 8*60*60)).Format(time.RFC3339Nano) + " ", want: stamp},
{name: "invalid", header: "not-a-timestamp", invalid: true},
} {
t.Run(tagging.name, func(t *testing.T) {
for _, metadata := range []map[string]string{nil, {"x-amz-meta-test": "kept"}} {
hdr := make(http.Header)
wantEncryption := make(http.Header)
for key, value := range encryption.headers {
hdr.Set(key, value)
wantEncryption.Set(key, value)
}
hdr.Set(xhttp.MinIOSourceTaggingTimestamp, tagging.header)
hdr.Set(xhttp.MinIOSourceMTime, stamp.Add(-time.Hour).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceObjectRetentionTimestamp, stamp.Add(-time.Minute).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceObjectLegalHoldTimestamp, stamp.Add(-time.Second).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceETag, "source-etag")
opts, err := putOptsFromHeaders(t.Context(), hdr, metadata, trusted)
if trusted && tagging.invalid {
if err == nil || !strings.Contains(err.Error(), xhttp.MinIOSourceTaggingTimestamp) {
t.Fatalf("malformed trusted timestamp: got %v", err)
}
continue
}
if err != nil {
t.Fatal(err)
}
wantTag, wantMTime, wantRetention, wantLegalhold, wantETag := time.Time{}, time.Time{}, time.Time{}, time.Time{}, ""
if trusted {
wantTag, wantMTime = tagging.want, stamp.Add(-time.Hour)
wantRetention, wantLegalhold, wantETag = stamp.Add(-time.Minute), stamp.Add(-time.Second), "source-etag"
}
if !opts.ReplicationSourceTaggingTimestamp.Equal(wantTag) {
t.Errorf("tag timestamp=%s, want %s", opts.ReplicationSourceTaggingTimestamp, wantTag)
}
if !opts.MTime.Equal(wantMTime) || !opts.ReplicationSourceRetentionTimestamp.Equal(wantRetention) ||
!opts.ReplicationSourceLegalholdTimestamp.Equal(wantLegalhold) || opts.PreserveETag != wantETag || opts.ReplicationRequest != trusted {
t.Error("other source fields did not preserve the replication trust boundary")
}
if opts.UserDefined == nil || (metadata != nil && !reflect.DeepEqual(opts.UserDefined, metadata)) {
t.Errorf("metadata=%v, want nonnil map preserving %v", opts.UserDefined, metadata)
}
gotEncryption := make(http.Header)
if opts.ServerSideEncryption != nil {
opts.ServerSideEncryption.Marshal(gotEncryption)
}
if !reflect.DeepEqual(gotEncryption, wantEncryption) {
t.Errorf("SSE headers=%v, want %v", gotEncryption, wantEncryption)
}
}
})
}
})
}
})
}
}
+33 -1
View File
@@ -184,6 +184,16 @@ func getAndValidateAttributesOpts(ctx context.Context, w http.ResponseWriter, r
return opts, valid
}
// Reject out-of-range page sizes as ListObjectParts does, instead of
// answering an invalid request with an empty parts listing.
if opts.MaxParts < 0 {
apiErr = errorCodes.ToAPIErr(ErrInvalidMaxParts)
argumentName = strings.ToLower(xhttp.AmzMaxParts)
argumentValue = r.Header.Get(xhttp.AmzMaxParts)
valid = false
return opts, valid
}
if opts.MaxParts == 0 {
opts.MaxParts = maxPartsList
}
@@ -196,6 +206,14 @@ func getAndValidateAttributesOpts(ctx context.Context, w http.ResponseWriter, r
return opts, valid
}
if opts.PartNumberMarker < 0 {
apiErr = errorCodes.ToAPIErr(ErrInvalidPartNumberMarker)
argumentName = strings.ToLower(xhttp.AmzPartNumberMarker)
argumentValue = r.Header.Get(xhttp.AmzPartNumberMarker)
valid = false
return opts, valid
}
opts.ObjectAttributes = parseObjectAttributes(r.Header)
if len(opts.ObjectAttributes) < 1 {
apiErr = errorCodes.ToAPIErr(ErrInvalidAttributeName)
@@ -415,7 +433,16 @@ func putOptsFromHeaders(ctx context.Context, hdr http.Header, metadata map[strin
if err != nil {
return ObjectOptions{}, err
}
sseKms, err := encrypt.NewSSEKMS(keyID, context)
// kms.Context implements encoding.TextMarshaler, so handing it to the
// SDK's interface{} parameter would serialize the context as a JSON
// string, which the receiving ParseHTTP rejects; a nil Context is a
// typed nil there and would be sent as "{}". Pass a plain map, or
// nothing when no context was requested.
var sdkContext any
if context != nil {
sdkContext = map[string]string(context)
}
sseKms, err := encrypt.NewSSEKMS(keyID, sdkContext)
if err != nil {
return ObjectOptions{}, err
}
@@ -425,6 +452,11 @@ func putOptsFromHeaders(ctx context.Context, hdr http.Header, metadata map[strin
MTime: mtime,
PreserveETag: etag,
ReplicationRequest: trustedReplication,
// These timestamps order replicated retention, legal hold and tagging
// updates on an SSE-KMS destination.
ReplicationSourceLegalholdTimestamp: lholdtimestmp,
ReplicationSourceRetentionTimestamp: retaintimestmp,
ReplicationSourceTaggingTimestamp: taggingtimestmp,
}
return op, nil
}
+108
View File
@@ -18,9 +18,11 @@
package cmd
import (
"encoding/xml"
"net/http"
"net/http/httptest"
"reflect"
"strings"
"testing"
xhttp "github.com/minio/minio/internal/http"
@@ -76,3 +78,109 @@ func TestGetAndValidateAttributesOpts(t *testing.T) {
})
}
}
// TestGetAndValidateAttributesOptsPartsRange asserts that GetObjectAttributes
// range checks the ObjectParts pagination headers the way ListObjectParts
// does: a negative value is rejected with the same API error, an absent or
// zero x-amz-max-parts means the default page size, and in-range values are
// passed through unchanged.
func TestGetAndValidateAttributesOptsPartsRange(t *testing.T) {
globalBucketVersioningSys = &BucketVersioningSys{}
bucket := minioMetaBucket
ctx := t.Context()
testCases := []struct {
name string
maxParts string
marker string
wantValid bool
wantErr APIErrorCode
wantArgument string
wantValue string
wantMaxParts int
wantMarker int
}{
{
name: "defaults",
wantValid: true,
wantMaxParts: maxPartsList,
},
{
name: "zero max-parts means the default page size",
maxParts: "0",
wantValid: true,
wantMaxParts: maxPartsList,
},
{
name: "in range values pass through",
maxParts: "10",
marker: "3",
wantValid: true,
wantMaxParts: 10,
wantMarker: 3,
},
{
name: "negative max-parts is rejected",
maxParts: "-1",
wantValid: false,
wantErr: ErrInvalidMaxParts,
wantArgument: strings.ToLower(xhttp.AmzMaxParts),
wantValue: "-1",
},
{
name: "negative part-number-marker is rejected",
marker: "-1",
wantValid: false,
wantErr: ErrInvalidPartNumberMarker,
wantArgument: strings.ToLower(xhttp.AmzPartNumberMarker),
wantValue: "-1",
},
}
for _, testCase := range testCases {
t.Run(testCase.name, func(t *testing.T) {
rec := httptest.NewRecorder()
req := httptest.NewRequest(http.MethodGet, "/testbucket/testobject?attributes", nil)
req.Header.Set(xhttp.AmzObjectAttributes, "ObjectParts")
if testCase.maxParts != "" {
req.Header.Set(xhttp.AmzMaxParts, testCase.maxParts)
}
if testCase.marker != "" {
req.Header.Set(xhttp.AmzPartNumberMarker, testCase.marker)
}
opts, valid := getAndValidateAttributesOpts(ctx, rec, req, bucket, "testobject")
if valid != testCase.wantValid {
t.Fatalf("want valid %v, got %v (%s)", testCase.wantValid, valid, rec.Body.String())
}
if testCase.wantValid {
if opts.MaxParts != testCase.wantMaxParts {
t.Errorf("want MaxParts %d, got %d", testCase.wantMaxParts, opts.MaxParts)
}
if opts.PartNumberMarker != testCase.wantMarker {
t.Errorf("want PartNumberMarker %d, got %d", testCase.wantMarker, opts.PartNumberMarker)
}
return
}
wantErr := errorCodes.ToAPIErr(testCase.wantErr)
if rec.Code != wantErr.HTTPStatusCode {
t.Errorf("want HTTP status %d, got %d", wantErr.HTTPStatusCode, rec.Code)
}
var errResp objectAttributesErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &errResp); err != nil {
t.Fatalf("decode error response: %v (%s)", err, rec.Body.String())
}
if errResp.Code != wantErr.Code || errResp.Message != wantErr.Description {
t.Errorf("want error %s/%q, got %s/%q", wantErr.Code, wantErr.Description, errResp.Code, errResp.Message)
}
if errResp.ArgumentName == nil || *errResp.ArgumentName != testCase.wantArgument {
t.Errorf("want ArgumentName %q, got %v", testCase.wantArgument, errResp.ArgumentName)
}
if errResp.ArgumentValue == nil || *errResp.ArgumentValue != testCase.wantValue {
t.Errorf("want ArgumentValue %q, got %v", testCase.wantValue, errResp.ArgumentValue)
}
})
}
}
+29 -6
View File
@@ -49,6 +49,7 @@ import (
xhttp "github.com/minio/minio/internal/http"
xioutil "github.com/minio/minio/internal/ioutil"
"github.com/minio/minio/internal/logger"
"github.com/minio/sio"
"github.com/pgsty/silo-pkg/v3/trie"
"github.com/pgsty/silo-pkg/v3/wildcard"
"github.com/valyala/bytebufferpool"
@@ -610,7 +611,10 @@ func excludeForCompression(header http.Header, object string, cfg compress.Confi
return true
}
if crypto.Requested(header) && !cfg.AllowEncrypted {
// SSE-C replication sends raw ciphertext without compression metadata.
// Exclude new SSE-C data from compression; other modes follow allow_encryption.
if crypto.SSEC.IsRequested(header) ||
(crypto.Requested(header) && !cfg.AllowEncrypted) {
return true
}
@@ -675,19 +679,35 @@ func getPartFile(entriesTrie *trie.Trie, partNumber int, etag string) (partFile
return partFile
}
func partNumberToRangeSpec(oi ObjectInfo, partNumber int) *HTTPRangeSpec {
func partNumberToRangeSpec(oi ObjectInfo, partNumber int) (*HTTPRangeSpec, error) {
if oi.Size == 0 || len(oi.Parts) == 0 {
return nil
return nil, nil
}
// For an encrypted, uncompressed object derive each part's plaintext length
// from the stored ciphertext length instead of trusting ActualSize: parts
// written before this was normalised record the ciphertext length there.
// The range returned here is consumed in the plaintext domain, where
// GetDecryptedRange and DecryptedSize both use exactly this arithmetic.
_, isEncrypted := crypto.IsEncrypted(oi.UserDefined)
deriveFromSize := isEncrypted && !oi.IsCompressed()
var start int64
end := int64(-1)
for i := 0; i < len(oi.Parts) && i < partNumber; i++ {
partSize := oi.Parts[i].ActualSize
if deriveFromSize {
decrypted, err := sio.DecryptedSize(uint64(oi.Parts[i].Size))
if err != nil {
return nil, errObjectTampered
}
partSize = int64(decrypted)
}
start = end + 1
end = start + oi.Parts[i].ActualSize - 1
end = start + partSize - 1
}
return &HTTPRangeSpec{Start: start, End: end}
return &HTTPRangeSpec{Start: start, End: end}, nil
}
// Returns the compressed offset which should be skipped.
@@ -806,7 +826,10 @@ func NewGetObjectReader(rs *HTTPRangeSpec, oi ObjectInfo, opts ObjectOptions, h
}
if rs == nil && opts.PartNumber > 0 {
rs = partNumberToRangeSpec(oi, opts.PartNumber)
rs, err = partNumberToRangeSpec(oi, opts.PartNumber)
if err != nil {
return nil, 0, 0, err
}
}
_, isEncrypted := crypto.IsEncrypted(oi.UserDefined)

Some files were not shown because too many files have changed in this diff Show More