Compare commits

..

133 Commits

Author SHA1 Message Date
Feng Ruohang f7808a172c deps: pin the coordinated September 13 SILO stack
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 09:54:51 +08:00
Feng Ruohang acc9b514f5 Follow the curl builder rename in delivery verification
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 09:08:38 +08:00
Feng Ruohang 31ba01c5d5 Refresh Go dependencies and build current static curl for both architectures
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-13 09:01:09 +08:00
Feng Ruohang 48ec10312f Merge pull request #180 from pgsty/codex/issue-77-metadata-convergence
fix(replication): preserve bucket metadata source times and converge deletions
2026-09-12 18:07:36 +08:00
Feng Ruohang 114dc10529 docs: archive issue 77 design, adversarial reviews and acceptance evidence
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 17:52:25 +08:00
Feng Ruohang 461e9a7210 test(replication): satisfy diagnostic regression style checks
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 17:17:18 +08:00
Feng Ruohang fcbb93e895 fix(replication): diagnose unusable metadata without a heal source
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 17:09:34 +08:00
Feng Ruohang 62cf066ff5 fix(replication): recover physical creation time and align policy status
An independent adversarial review of the bucket metadata convergence
work found three defects it had introduced.

GetBucketInfo overwrote the physical creation probe with cached
metadata, which a bucket that never held a configuration legitimately
lacks. The new creation-time requirement then failed every policy, tag,
SSE, quota, versioning and Object Lock write on such a bucket, with no
operator recovery path, and initial synchronization skipped it silently.
Return the physical result unchanged when metadata is not requested, as
ListBuckets already does, recover the time during initial
synchronization, and pass it to MakeBucketHook so peers adopt the same
bucket generation.

Replication status compared parsed policies statement by statement while
heal compares the canonical key. An upgraded peer that stored an
equivalent statement order was therefore reported as mismatched forever,
and heal never had anything to write. Compare the key heal compares;
per-site presence counting is unchanged.

Heal diagnostics shared one log key across four conditions, so a real
peer RPC failure could be deduplicated away by an earlier message, and
they were logged at error level for the normal transient of a peer that
does not have the bucket yet. Give each reason its own key at warning
level, report only a field state that exists and still cannot be
ordered, and diagnose nothing when no site holds a state to propagate.

The recovery test now runs against the real ObjectLayer; the stub it
replaced returned the expected time and hid the defect. The policy
status test uses a statement order the canonical encoder reorders, and
adoption coverage is extended past a real field time.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 14:01:35 +08:00
Feng Ruohang 1ee64a8d89 fix(replication): gate metadata tombstone export during rolling upgrades
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 13:50:55 +08:00
Feng Ruohang 01aaef2b50 fix(replication): converge bucket metadata using deterministic source states
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 13:50:55 +08:00
Feng Ruohang bcc62afe3d fix(replication): apply bucket metadata source times under the metadata lock
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 13:50:55 +08:00
Feng Ruohang 5c57658163 Merge pull request #179 from pgsty/codex/federation-copy-fixes
fix(federation): preserve committed copy times and reject raw SSE-C replicas
2026-09-12 02:00:57 +08:00
Feng Ruohang f175e98c34 fix(federation): bind copy timestamps to committed writes
Return the committed object or part time on federation write responses and
capture it for CopyObject and UploadPartCopy without a follow-up read.
Keep successful writes compatible with targets that do not supply a time.

Reject authenticated raw SSE-C replica CopyObject across deployments before
forwarding. Extend the existing SSE fixtures to verify empty objects,
stored checksums, KMS contexts and multipart sources.

Refs: #169, #168, #171
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-12 01:44:22 +08:00
Feng Ruohang 12f631b502 Merge pull request #178 from pgsty/codex/multipool-correctness-20260911
fix(storage): reconcile multi-pool writes and conditional deletes
2026-09-11 20:34:58 +08:00
Feng Ruohang ccb676e60c fix(storage): preserve shared tier references during pool cleanup
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 20:23:40 +08:00
Feng Ruohang 51d41345f7 test: record Linux restart and OIDC release acceptance
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 20:05:56 +08:00
Feng Ruohang e59a3d938e fix(storage): serialize and reconcile multi-pool object updates
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 20:05:56 +08:00
Feng Ruohang b32f2d9dd0 Merge pull request #177 from pgsty/codex/release-consolidation-20260911
Fix signed payloads, Object Lock and federated copies; secure AMQP
2026-09-11 17:08:34 +08:00
Feng Ruohang d63c92e393 build(deps): secure AMQP frames and select merged Console
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:56:33 +08:00
Feng Ruohang 68127c5a63 build(deps): embed the consolidated Console source
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:40:42 +08:00
Feng Ruohang c4b5e1cb45 fix(auth): enforce header-only presigned payload checksums
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:24:42 +08:00
Feng Ruohang 9a303f5096 fix(object): preserve retention and federated copy destination state
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:24:42 +08:00
Feng Ruohang 87d8b5967f fix(auth): align signed request and policy condition semantics
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:24:42 +08:00
Feng Ruohang f760046c44 Merge remote-tracking branch 'origin/main' into codex/release-consolidation-20260911
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-11 16:13:42 +08:00
Feng Ruohang 93e7ef4bcc Merge pull request #175 from pgsty/codex/upstream-sdk-password-20260910
fix(iam)!: split self-service password permissions and update SDK stack
2026-09-10 17:51:08 +08:00
Feng Ruohang a4229b366f build(deps): align final coordinated SILO source pins
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 17:39:01 +08:00
Feng Ruohang 420340bc14 docs(iam): explain breaking password-policy semantics
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 16:58:21 +08:00
Ayush Sharma e7e87402ed fix: forward the legal hold as an Object Lock header on federated CopyObject
A cross-deployment CopyObject that requests
`x-amz-object-lock-legal-hold: ON` answered 200 while the destination
carried no hold. The resolved value reached the remote as ordinary user
metadata, `X-Amz-Meta-X-Amz-Object-Lock-Legal-Hold`, so nothing applied it.
Retention requested on the same copy survived, which is what made the loss
easy to miss.

The federation branch passes the resolved metadata map straight to
`Core.PutObject` as `PutObjectOptions.UserMetadata`. minio-go's `Header()`
writes the typed lock fields first, then prefixes every UserMetadata key it
does not recognise with `x-amz-meta-`; `supportedHeaders` covers
`x-amz-object-lock-mode` and `x-amz-object-lock-retain-until-date` but not
`x-amz-object-lock-legal-hold`, and `isAmzHeader` does not match it either.
Retention therefore arrives as real headers and the hold does not. The
high-level `validate()` that would have rejected the key never runs, because
`Core.PutObject` goes straight to the low-level PUT.

Carry the hold on the typed `LegalHold` option and forward a cloned map with
the raw key removed. The clone matters twice: typed fields are written before
the UserMetadata loop, so a leftover raw key would add a bogus `x-amz-meta-`
entry beside the correct header, and the proxy's own response and event
metadata are rebuilt from the resolved values rather than the forwarding map,
which no longer carries the hold.

Retention stays in the map deliberately. It already passes through as a
standard header, and moving it to the typed `RetainUntilDate` field would
format with `time.RFC3339` and truncate a retain-until date to whole seconds.

The new test asserts the wire: the remote must receive
`X-Amz-Object-Lock-Legal-Hold` and never the `x-amz-meta-` spelling, and the
destination version must actually store the hold. It fails without the change
with "legal hold forwarded as user metadata [ON]".

Fixes #166

Signed-off-by: Ayush Sharma <72848455+Aeirx@users.noreply.github.com>
2026-09-10 13:15:10 +05:30
Feng Ruohang a164e1dda1 fix(iam): separate password changes and refresh coordinated SDK dependencies
Use ChangeMyPassword for the authenticated user and keep CreateUser for other users. Coordinate silo-pkg b3760f56ec23, mcli fa22b40b4eb7, Console 1b95b6cec652 and upstream minio-go 78bfa91607c2. Add legacy-policy and SDK streaming regressions plus upgrade guidance.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 15:16:21 +08:00
Feng Ruohang 2f61325d4a Merge pull request #174 from pgsty/codex/contributor-orenyomtov-20260910
docs: credit @orenyomtov for the SN-2026-011 report
2026-09-10 11:34:52 +08:00
Feng Ruohang a6145e1d2e docs: credit @orenyomtov for the SN-2026-011 report
Add Oren Yomtov (github.com/orenyomtov) to the contributor wall in
README.md, README_ZH.md, and CONTRIBUTORS.md, and to the Issue reports
table, for the private disclosure of the unsigned-header CopyObject
cross-object read fixed as SN-2026-011 (#173). Community contributor
count 40 -> 41.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 11:24:12 +08:00
Feng Ruohang 5232546690 Merge pull request #173 from pgsty/codex/unsigned-amz-header-copy-20260909
fix(auth): reject unsigned x-amz-* headers to close CopyObject confused-deputy (SN-2026-011)
2026-09-10 10:10:49 +08:00
Feng Ruohang 0c46cb641b docs(security): record SN-2026-011 (unsigned x-amz-* header CopyObject)
Ledger entry for the confused-deputy fix in 123325430: an unsigned
x-amz-copy-source header turned a presigned or signed PUT into a
server-side copy of any object the signing key can read. Reported by
Oren Yomtov; inherited from upstream minio/minio; CVE requested.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 09:57:26 +08:00
Feng Ruohang 1233254309 fix(auth): reject unsigned x-amz-* headers to close CopyObject confused-deputy
A presigned or signed PUT authorized for a single object could be turned
into a server-side copy of any object the signing key can read by adding
an unsigned x-amz-copy-source header, executed as the signer. SigV4
verification only walked the signed-headers list, never the headers that
actually arrived; the meta-header check matched only X-Amz-Meta- and ran
only on the presigned path, so an unsigned x-amz-* header outside the
list was never seen while the router still dispatched the PUT to
CopyObjectHandler.

Reject any x-amz-* request header not covered by the signed headers, on
both the presigned (doesPresignedSignatureMatch) and Authorization-header
(doesSignatureMatch) paths, matching AWS S3. The check tests membership
in the signed set rather than value equality, so a header whose first
value is empty (e.g. {"", "/src/secret"}) cannot slip through.
X-Amz-Content-Sha256 is exempt (payload hash: read from the query for
presigned requests and bound into the string-to-sign for signed ones, so
it is self-protected) and X-Amz-Signature-Age is exempt (an internal
scratch header written after verification, so repeated verification of
the same request stays idempotent). The synthesized X-Amz-Tagging header
in PutObjectTagging is now injected after signature verification.

Tests that previously added x-amz-copy-source and friends after signing
(relying on the vulnerable behavior) now re-sign, mirroring real S3
clients. Adds checkUnsignedHeaders unit cases and TestPresignedVerifyIdempotent.

Reported by Oren Yomtov. Inherited unchanged from upstream minio/minio.
Tracked as SN-2026-011.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-10 09:57:01 +08:00
Feng Ruohang 8a2fe9b7a0 Merge pull request #164 from pgsty/codex/main-consolidation-20260909
fix: honor TLS defaults and accept valid bucket metadata reloads
2026-09-09 19:49:08 +08:00
Feng Ruohang f26ee6bf0a test: cover metadata reload with equal maximum timestamps
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 19:35:13 +08:00
Feng Ruohang d69c4ccfe4 Merge TLS default key exchange compatibility fixes
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 19:31:35 +08:00
Feng Ruohang 086e619505 Merge accepted bucket metadata reload correction
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 19:31:10 +08:00
Feng Ruohang 48e1846525 fix(tls): honor Go key exchange defaults across transports
Remove the eight explicit curve overrides so Go 1.27 honors tlsmlkem=0
across Server listeners, node links and outbound transports. Remove the
unused shared curve option and add wire-level regression coverage.

Document CA trust and TLS upgrade behavior, retain the investigation
artifacts, and exclude their synthetic routes from the rebrand guard.
The product compatibility baseline remains unchanged.

Validation: focused race tests, HTTP tests, lint, compatibility guard
positive/negative controls, and a fresh Linux build with three isolated
OIDC integration scenarios all pass.

Adversarial review: Claude Code Fable 5.1, max effort.
Final verdict: APPROVE FOR COMMIT.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 18:50:12 +08:00
Feng Ruohang bcc8871b1d Merge pull request #163 from pgsty/fix/issue-158-federated-copy-sse
fix: forward plaintext on federated CopyObject of SSE objects (#158)
2026-09-09 18:02:45 +08:00
Feng Ruohang 25cb3511c9 test: cover multipart SSE sources on federated CopyObject
A multipart SSE-S3 source is encrypted per part, so its logical size is
the sum of the parts' decrypted sizes and the decrypting reader crosses
a part boundary. Copy such a source across the federation to a plain
and to an SSE-S3 destination and check the destination plaintext and
single-encryption size.

The test router registers routes in endpoint order and the plain
PutObject route has no query matcher, so the multipart endpoints are
listed first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FodsDpa6VkghaeRE6WjmEe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 17:51:32 +08:00
Feng Ruohang cfefc049c1 fix: send the SSE-KMS context as a JSON object on federated copies
putOptsFromReq handed the parsed kms.Context straight to
encrypt.NewSSEKMS. kms.Context implements encoding.TextMarshaler, so the
SDK serialized it as a JSON string, and a request without a context
still produced one because the nil Context is a typed nil inside the
interface value and marshals to "{}". The receiving ParseHTTP rejects
both forms, so every federated CopyObject to an SSE-KMS destination
failed with InvalidArgument once the forwarded stream was correct.

Pass a plain map, or nothing when no context was requested, and cover
SSE-KMS destinations with and without an explicit context.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FodsDpa6VkghaeRE6WjmEe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 17:51:32 +08:00
Feng Ruohang af56d17630 fix: forward plaintext on federated CopyObject of SSE objects (#158)
The legacy etcd bucket-federation branch of CopyObjectHandler reads its
source through getObjectNInfo, which yields the decrypted and
decompressed bytes, but it also ran the destination encryption locally
and then forwarded that stream to the remote PutObject with the source's
stored size and the destination SSE option. SSE to plain and plain to
SSE therefore failed on a Content-Length mismatch, while SSE to SSE
matched by coincidence: the remote encrypted the ciphertext a second
time and stored an unreadable object, and a destination GET returned
the inner ciphertext with HTTP 200.

The remote write owns the destination's storage transformations, so
hand it the logical bytes at their logical size and let it encrypt
exactly once. Compression was already excluded on this branch; apply
the same rule to encryption, size the forwarded reader by actualSize,
and declare that size on the forwarded PutObject.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FodsDpa6VkghaeRE6WjmEe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 17:51:32 +08:00
Feng Ruohang d1105bbb3d Merge pull request #162 from pgsty/codex/replication-reliability-20260909
fix(replication): complete purges, expose MRF drops, and cancel resyncs reliably
2026-09-09 14:55:35 +08:00
Feng Ruohang 66fe61ff65 test: reuse replication compatibility fixtures
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:44:02 +08:00
Feng Ruohang a1141a43f2 style: format replication regression fixture
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:24 +08:00
Feng Ruohang 702f113f51 fix(replication): scope resync cancellation and drain worker lifecycle
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:02 +08:00
Feng Ruohang 63aace4099 fix(replication): expose bounded MRF queue drops
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:02 +08:00
Feng Ruohang 9d7094b770 fix(replication): complete single-object delete marker purges
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 14:37:02 +08:00
Feng Ruohang 450dcb8484 Merge pull request #161 from pgsty/codex/server-dependency-refresh-20260909
build: refresh maintained SILO stack dependencies
2026-09-09 12:04:16 +08:00
Feng Ruohang 4074d00b96 build: refresh maintained SILO stack dependencies
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-09 11:49:13 +08:00
Feng Ruohang 75ba0ce402 fix: drop the broken bucket-metadata reload publication guard (issue #105 T3)
The #105 T3 change (PR #156) tried to keep the resident metadata cache
monotonic by guarding peer-reload publication on lastUpdate(). But
lastUpdate() is the max of per-config timestamps and cannot order whole
records: a node caching {policy@20, CORS@10} that receives a newer CORS@15
still has lastUpdate()==20, so the guard rejects the legitimately-newer
record and the periodic refresh (same comparator) cannot repair it. A
paused reload could also resurrect deleted resident state.

Per the maintainer decision, revert the reload publication to its original
unconditional (acceptable-until-refresh) behavior:
- remove setReloaded and restore the plain Set plus notification/target
  registry updates in LoadBucketMetadataHandler;
- restore refreshBucketsMetadataLoop's own lastUpdate() staleness check and
  globalEventNotifier.set / globalBucketTargetSys.set publication;
- restore the unconditional GetConfig cache-miss publication;
- document the known freshness limitation at the reload site (the periodic
  refresh is best-effort and cannot repair an equal-maximum-timestamp
  divergence).

The T1 lifecycle merge-under-lock (UpdateExpiryLCConfig) and both T2 fixes
(DeleteBucket takes metadata.lock before deleting; saveMetadata and
loadBucketMetadataParseUnderLock recheck physical bucket existence) are
kept fully intact.

Tests:
- drop the T3 reproductions (overlapping-reload resident-cache test and the
  peer-reload-preserves-current-targets publication test);
- add lockBucketMetadataAcquireHook, a nil-in-production atomic test hook in
  the shared metadata.lock path, so tests can deterministically observe a
  caller (notably DeleteBucket, whose lock is taken through its
  erasureServerPools receiver and is invisible to an injected object layer)
  reaching the lock;
- rewrite the T2 delete-race ghost test to hold metadata.lock MID-SAVE (past
  saveMetadata's existence recheck) and synchronize on the delete's actual
  lock attempt via the hook, so it isolates the lock-before-delete fix:
  removing only DeleteBucket's metadata.lock (recheck kept) now fails it;
- rewrite the cancellation test to observe the delete's actual lock attempt,
  then cancel and await its error while still holding the lock, so a
  scheduling-delayed delete stopped by the canceled context can no longer
  pass on a broken tree.

Refs #105. Follows #156.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 16:13:54 +08:00
Feng Ruohang da142327f2 Merge pull request #159 from pgsty/fix/issue-99-100-followups
fix: repair federated CopyObject checksum edge cases (empty body, inherited, multipart-suffix)
2026-09-08 15:58:22 +08:00
Feng Ruohang a3df317ae0 Merge pull request #60 from mrjavadseydi/feat/access-based-ilm
[ILM] Relocate hot objects across server pools by GET frequency
2026-09-08 15:42:15 +08:00
Feng Ruohang 2c50d11f72 fix: correct federated CopyObject checksum edge cases (#99 follow-ups)
Three residual checksum defects in the legacy etcd federation branch of
CopyObjectHandler, found by post-merge review of #157.

1. Empty-source 500 regression. A checksum-less object gains the S3 default
   CRC-64NVME (WantServerSideChecksumType is set), but minio-go streams no
   trailing checksum for a 0-byte body (contentLength == 0), so the remote
   computed none, federatedChecksumValue was empty, hash.NewChecksumWithType
   returned nil, and the handler returned 500 -- so every empty-object
   federated copy failed. For a 0-byte source, forward the empty-content digest
   as an ordinary checksum request header instead, so the remote validates,
   persists and returns it, matching the local path (e.g. CRC32 "AAAAAA==").

2. Inherited full-object checksum dropped. When the source already carries a
   full-object checksum, the local path sets dstOpts.WantChecksum, not
   WantServerSideChecksumType (only multipart-composite sources are promoted).
   The federated branch inspected only WantServerSideChecksumType, so a
   checksum-bearing source's checksum was silently discarded on a federated
   copy that requested no algorithm. Forward WantChecksum.Encoded (always a
   plain digest) as a checksum header so the remote validates and persists it,
   and bind the returned value, matching local persistence.

3. Multipart-suffixed remote value accepted. The bind accepted a value like
   "NSRBwg==-0": NewChecksumWithType parses the "-N" as ChecksumMultipart with
   WantParts 0 and the length-only validator passes, so the destination was
   returned as COMPOSITE. A single forwarded PutObject must yield a full-object
   digest, so reject a multipart-marked parsed value in addition to the
   existing nil (missing/malformed) rejection.

A forwarded checksum request header is stripped from objInfo.UserDefined so it
is not mistaken for object metadata.

Out of scope: the SSE federated-copy corruption (srcInfo.Reader/Size mismatch
for encrypted sources) predates this work and is filed separately.

New federated regressions cover empty source with requested and default
checksum (200 + correct value + persisted), an inherited full-object checksum
preserved without a requested algorithm, and a multipart-suffixed remote value
rejected. Red/green verified for each against the merged code.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:38:11 +08:00
Feng Ruohang 5ac33e1583 test: make access move failure recovery deterministic
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:28:33 +08:00
Feng Ruohang d57c4e8407 Merge remote-tracking branch 'origin/main' into codex/access-tiering-ci-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:08:34 +08:00
Feng Ruohang 374de0fa32 fix: make access tier moves preserve versions and isolate writes
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 15:08:34 +08:00
Feng Ruohang 80e23dc9f2 Merge pull request #156 from pgsty/codex/bucket-metadata-merge-20260908
fix: preserve bucket metadata across concurrent updates and reloads
2026-09-08 15:03:52 +08:00
Feng Ruohang 2cc0e3c6ed test: construct notification fixtures with the event ARN type
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:52:48 +08:00
Feng Ruohang f9da3b919d fix: order bucket deletion and metadata publication safely
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:43:15 +08:00
Feng Ruohang 1309853f57 Merge remote-tracking branch 'origin/main' into codex/access-tiering-ci-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:34:51 +08:00
Feng Ruohang 39b8e6c30a Merge remote-tracking branch 'origin/main' into codex/bucket-metadata-merge-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 14:24:22 +08:00
Feng Ruohang 9a6e1477f4 ci: record access tiering compatibility identifiers
Record the ten documented MINIO_ILM_ACCESS settings, the internal object
metadata stamp, and the tracker storage-path suffix introduced by this PR.
The guard places the /ilm/access string in its routes set, but the value
is a component of the tracker object prefix, not a public HTTP endpoint.

Keep all existing compatibility entries. The guard, delivery rebrand check,
and Docker entrypoint compatibility tests pass with the refreshed manifest.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 13:18:50 +08:00
Feng Ruohang 49375ed2d3 Merge pull request #157 from pgsty/codex/federated-copy-merge-20260908
fix: preserve metadata and checksums in federated object copies
2026-09-08 13:16:30 +08:00
Feng Ruohang f817b5261c Merge pull request #151 from nikitapogromsky/fix/idempotent-add-system-target
logger: make AddSystemTarget idempotent
2026-09-08 13:14:30 +08:00
Feng Ruohang 885bd2c20a fix: reject missing remote checksums on federated copies
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 13:05:45 +08:00
Feng Ruohang 89b75e7913 Merge branch 'main' into feat/access-based-ilm 2026-09-08 13:04:25 +08:00
Feng Ruohang 079ebb1926 fix: serialize logger initialization and publish the console target safely
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:58:21 +08:00
Feng Ruohang 04c29aac11 Merge branch 'audit/issue-105-bucketmeta-races' into codex/bucket-metadata-merge-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:55:51 +08:00
Feng Ruohang 95e7a190b3 Merge branch 'fix/issue-99-100-federated-copyobject' into codex/federated-copy-merge-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:55:51 +08:00
Feng Ruohang 8e2392e48f Merge remote-tracking branch 'origin/main' into codex/logger-idempotency-20260908
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 12:55:50 +08:00
Feng Ruohang 3abe0d95a5 Merge pull request #149 from pgsty/codex/embedded-console-compat
fix: restore embedded Console proxy and WebSocket configuration
2026-09-08 10:09:21 +08:00
Feng Ruohang 9c6c9805de fix: select the validated Console mainline for embedding
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-08 10:00:45 +08:00
nikitapogromsky 5cb900bfad logger: make AddSystemTarget idempotent
Subscribe re-registered the console target on every console-log subscription, producing duplicate minio_logger_webhook_* series on each /minio/metrics/v3 scrape. Fixes #150

Signed-off-by: nikitapogromsky <129324283+nikitapogromsky@users.noreply.github.com>
2026-09-07 12:03:39 +03:00
Feng Ruohang 0af5d22286 fix: restore embedded Console proxy and WebSocket configuration
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 13:36:24 +08:00
Feng Ruohang f1687f402b fix: scope rebrand checks to the retired repository
Match the exact retired repository while preserving links to distinct repositories and historical issue titles in contributor credits. Continue rejecting live links to the retired repository in those credits.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 11:49:23 +08:00
Feng Ruohang ce606df2c4 fix: align contribution tooling with SILO AGPL policy
Clarify SILO contribution ownership and preserve prior copyright notices. Consolidate issue templates and route the legacy credits command through the maintained generator.

Validation: make rebrand-guard; bash -n update-credits.sh; regenerated credits match CREDITS; template and link checks.
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 11:33:21 +08:00
Feng Ruohang fd44dc4e9b docs: include the latest merged community contribution
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 11:01:04 +08:00
Feng Ruohang 479745e764 docs: credit contributors across the SILO repositories
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:58:47 +08:00
Feng Ruohang cc1c54475f Merge pull request #146 from pgsty/fix/issue-108-console-loopback-tls
fix: keep embedded Console login working over loopback TLS (#108)
2026-09-07 10:53:45 +08:00
Feng Ruohang e7654d470c fix: close residual bucket-metadata races (issue #105 audit)
Audit of the three deferred #105 follow-ups. Each reproduces with a
deterministic red test in cmd/bucket-metadata-race_test.go, and each fix is
the minimal change that turns its test green while preserving the
<bucket>.lck -> metadata.lock -> .metadata.bin lock order established by #103.

1. Lifecycle expiry merge lost update (persistent). PeerBucketLCConfigHandler
   and healBucketILMExpiry read the current lifecycle with an unlocked
   GetConfigFromDisk, merged the replicated expiry rules with the local
   transition rules, then wrote the pre-computed blob via Update. Any lifecycle
   transition change committed between the merge read and the merge write was
   silently lost. New BucketMetadataSys.UpdateExpiryLCConfig performs the read,
   merge, and save under one metadata.lock; mergeExpiryWithLCConfig now takes
   the locked snapshot and validates object-lock retention from it instead of
   re-reading (avoids a re-entrant metadata load under the lock).

2. DeleteBucket ghost .metadata.bin (persistent). DeleteBucket took only
   <bucket>.lck while config writers take only metadata.lock, so a writer that
   was mid-save could re-create .metadata.bin after the prefix purge. The purge
   now runs under metadata.lock, with a best-effort unlocked fallback so a
   delete is never blocked from completing.

3. Overlapping peer reloads publishing a stale resident cache (freshness only;
   the persisted record stays correct). LoadBucketMetadataHandler and the
   GetConfig cache-miss path published with an unconditional Set, so a reload
   that read an older revision could overwrite a newer resident record until the
   next refresh. New BucketMetadataSys.setReloaded (and a matching GetConfig
   guard) refuses to regress a newer resident record, mirroring
   refreshBucketsMetadataLoop.

Verification: go build -tags kqueue,dev ./...; go vet ./cmd; gofmt clean;
rebrand-guard baseline unchanged; go test -tags kqueue,dev ./cmd (207s) green;
new tests plus the #103 metadata suite green under -race.

Refs #105. Parent #102. Foundation #103.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:33:11 +08:00
Feng Ruohang 4b25f7e819 fix: return and persist checksum on federated CopyObject (#99)
The legacy etcd federation branch of CopyObjectHandler forwards the copied
bytes with minio-go Core.PutObject but never asked the remote for a checksum
and discarded any it returned, so a cross-deployment whole-object copy that
requested a checksum returned 200 with an empty checksum, and a checksum-less
source did not gain the S3 default CRC-64NVME that the local path assigns. The
request was neither honored nor rejected. This is the whole-object counterpart
of #72, which repaired the same class of defect for federated UploadPartCopy.

When a server-side checksum is wanted -- explicitly requested, inherited from a
multipart source, or the CRC-64NVME default for a checksum-less object, all
already captured in dstOpts.WantServerSideChecksumType -- the forwarded
PutObject now streams a trailing checksum of that type, so the remote computes
and persists it and echoes it in the response. The value the remote reports for
that exact write is bound into objInfo.Checksum, matching how the local
CopyObject path carries checksums into the CopyObjectResult. Reading the value
from the same UploadInfo that produced the ETag keeps the pair bound to one
write.

Only the requested algorithm is returned; a malformed or absent remote value
leaves objInfo.Checksum unset, so an ordinary copy that wanted no checksum
still returns none. Two small mapping helpers convert between the server's
hash.ChecksumType and the minio-go request type and response field.

New end-to-end tests drive the real federation branch through
getRemoteInstanceClient and minio-go into a second in-process deployment and
assert that CRC32/CRC32C/SHA256/CRC64NVME and the no-algorithm default are all
returned in the CopyObjectResult and persisted on the destination, and that a
requested algorithm never leaks other algorithms into the response.

Fixes #99

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:32:59 +08:00
Feng Ruohang 711b092f86 fix: keep embedded Console login working over loopback TLS (#108)
The embedded Console reaches the S3/STS API at https://127.0.0.1:<port>
(minioConfigToConsoleFeatures), a loopback endpoint whose TLS certificate is
not expected to carry a 127.0.0.1 SAN. silo-console v2.3.x began verifying
every outbound TLS peer, so the Console's STS AssumeRole handshake to that
loopback endpoint now fails certificate validation and BOTH local and LDAP
logins fail with a generic "invalid login". The failure happens in the Console
HTTP client before any request reaches a server auth/STS/LDAP handler, so no
server-side auth error is logged, matching the report.

Restore the documented loopback bypass by opting the embedded Console into its
endpoint-scoped CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY switch whenever the server
falls back to the 127.0.0.1 endpoint under TLS. The exemption is scoped to that
single loopback origin inside Console; every other HTTPS peer (IdP, Prometheus,
webhooks) stays verified, preserving the v2.3.x hardening. An explicitly
configured endpoint is reached under its own verified name and is never
exempted. initConsoleServer unsets CONSOLE_* before re-deriving them, so the
switch cannot be supplied by the operator on the embedded path; the server must
assert it.

Fixes #108

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:32:59 +08:00
Feng Ruohang 2bc103b80c fix: strip all reserved metadata on federated CopyObject (#100)
The legacy etcd federation branch of CopyObjectHandler forwards the copied
source metadata to the remote deployment with minio-go Core.PutObject, after
removing only two reserved keys (compression and actual-size). Every small
object is stored inline, so its stored metadata also carries
X-Minio-Internal-inline-data; the remote's setRequestLimitMiddleware rejects
any request bearing a reserved-prefix header (containsReservedMetadata), so
the forwarded write failed with 400 InvalidArgument "Your metadata headers
are not supported." for the default COPY metadata directive.

A plain federated PutObject must not carry any internal storage metadata, so
strip the whole reserved-prefix class before forwarding instead of an
enumerated subset. Enumerating a third key would only defer the next leak:
besides inline-data, replication bookkeeping (replica/replication status and
timestamps) is added to the same map earlier in the handler and would be
rejected just the same. None of these keys is required by the remote for a
correct plain PutObject; they are internal storage details the remote sets
for itself. The stripping uses stringsHasPrefixFold, matching the remote's
own case-insensitive detection. Ordinary user metadata (x-amz-meta-*) is
untouched and still copied.

The pre-existing UUID ETag on the federated write (no Content-MD5 is sent) is
out of scope and left unchanged, as recorded in the issue.

A new end-to-end test drives the real federation branch through
getRemoteInstanceClient and minio-go into a second in-process deployment,
copying an inline source with the default COPY directive. It asserts the copy
now succeeds, that no reserved-prefix header reaches the remote on any
forwarded request, and that copied user metadata survives.

Fixes #100

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 10:32:59 +08:00
Feng Ruohang ad873c7357 Merge pull request #132 from mrjavadseydi/fix/issue-106-bucket-quota-metrics
fix: report effective bucket quotas in metrics
2026-09-07 09:54:01 +08:00
Feng Ruohang 0af0907eff Merge pull request #134 from pgsty/fix/issue-120-ssec-replica-retransmit
fix: retransmit and re-order Object Lock for SSE-C replicas (single erasure set)
2026-09-07 00:24:45 +08:00
Feng Ruohang e27ba2bc14 Merge pull request #131 from pgsty/fix/issue-117-lock-resend-compare
fix: stop re-replicating an object whose retention was removed
2026-09-07 00:22:21 +08:00
Feng Ruohang 236e163c0b fix: repair an undecodable SSE-C replica on retransmit
PutObjectHandler's precondition callback ran DecryptObjectInfo on the
stored object before checkPreconditionsPUT, so an authenticated raw
SSE-C replica overwrite was rejected when the stored version could not
decrypt. A replica a pre-fix destination (issue #109) left as
compress(ciphertext) or a re-encrypted body has an invalid decrypted
length, so DecryptObjectInfo returned errObjectTampered and the
retransmission that repairs it never ran -- the version stayed damaged
through resync. #134's raw-replica exemption only covered the
version/ETag duplicate check inside checkPreconditionsPUT, one step too
late.

Skip the stored object's decryption precondition only for a PURE raw
SSE-C replica overwrite (a trusted SSE-C replica write with no public
precondition), keyed on the incoming request's restored SSE-C metadata,
the same predicate checkPreconditionsPUT uses. Such a write fully
replaces the object, so requiring the damaged stored version to decrypt
is both wrong and unnecessary. A conditional request keeps the check:
DecryptObjectInfo also normalizes the stored sealed ETag to the
client-visible one, and If-Match/If-None-Match must compare against
that, not the sealed ETag -- skipping it for every replica inverted both
conditions. Ordinary writes and non-SSE-C replicas are unchanged.

Adds red/green regressions: a raw retransmit over a version staged as an
undecodable body returns 500 XMinioObjectTampered before this change and
200 with full customer-key recovery after; and a conditional replica PUT
(If-Match / If-None-Match) on the client-visible ETag is honoured rather
than inverted. Fixes the single-PUT compression-damage recovery gap in

Signed-off-by: Feng Ruohang <rh@vonng.com>
#120.
2026-09-07 00:11:04 +08:00
Feng Ruohang 7220210e8d test: reconcile #134 fixtures with #119 and refresh the rebrand baseline
Rebased onto current main. #119 made PutObjectPart derive an encrypted
part's plaintext length and reject a part that cannot be a valid sio
stream, so TestReplicaLockReconcileNullVersion's completeNullMPU fixture
(a 4-byte plaintext part under SSE-C metadata) no longer stores; build it
with sio.Encrypt like the other encrypted-part fixtures. Also regenerate
the rebrand-guard baseline for the replication SSE header the retransmit
path reintroduces (headers 84 -> 85). Mechanical integration only.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 34cbca97ea docs: point the multi-pool lock follow-up at pgsty/silo#133
Fill the tracked-issue number into the scope comments of the single
erasure set Object Lock reconcile added for SSE-C replica retransmit.
No behaviour change.

Refs pgsty/silo#120
Refs pgsty/silo#133

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 109d824e5f fix: retransmit and re-order Object Lock for SSE-C replicas (single erasure set)
Issue #120 routes an existing SSE-C replica through PutObjectHandler and
NewMultipartUploadHandler. On main those handlers assigned the incoming retention
and legal hold directly, without the source-timestamp ordering #111 added to
CopyObjectHandler and without persisting the ordering timestamps, so in
active-active replication a retransmit carrying an older value could overwrite a
destination version's newer lock state.

Share #111's ordering decision as applyReplicatedObjectLock in
cmd/bucket-object-lock.go and call it from CopyObject, PUT and multipart
initiation. A request that is not an actual trusted replica keeps ordinary write
semantics (a validated value is applied and stamped now); only a real replica
update is ordered against the stored version, so a marker-only peer write no
longer drops a validated hold or default retention. CopyObject keeps its SSE-C
key-rotation encMetadata reconciliation inline. putReplicationOpts now emits a
stored retention ordering timestamp even when the value keys are absent, so a
removal recorded on the retransmit PUT path still replicates onward. replicateAll
marks Failed and carries the error when putReplicationOpts fails.

The handler decision is made against the version as it stands then, which a
concurrent lock update can outrun before the write commits, and for multipart
across the whole initiation-to-completion span. Close that window under the
object write lock the receiving erasure set holds: a trusted SSE-C replica full
write sets ObjectOptions.ReplicaLockReconcile, and erasureObjects.PutObject and
CompleteMultipartUpload re-run the ordering (reconcileStoredObjectLock, which
orders retention and legal hold independently by their reserved timestamps)
against the destination version read on that set before committing. Persisted
upload metadata records the null version as an empty VersionID, so completion
looks that up as the null version rather than the latest. The reconcile runs only
against an existing version; a not-found destination keeps the write's own
accepted lock, including a pre-upgrade upload that persisted values without
ordering timestamps, and a non-not-found read error fails the write. Scoped to
the SSE-C paths this issue enables; CopyObject is left as #111 wrote it.

Scope: this orders Object Lock against the destination version under the write
lock and is correct for a single erasure set. A multi-pool deployment -- where a
version can have duplicate copies across pools, object ModTime ties do not track
per-field lock timestamps, and the object namespace lock is per-pool -- needs a
cross-pool lock-safe reconcile and is deliberately out of scope here, tracked in
pgsty/silo#TBD-multipool-lock.

Tests: TestAPISSECReplicaRetransmitObjectLockOrdering and its multipart sibling;
TestAPIReplicaMultipartNewerHoldSurvivesCompletion and
TestReplicaPutObjectLockReconcileUnderWriteLock (a hold or retention reaching the
version after the handler decision, or after multipart initiation, survives the
commit; a pre-upgrade upload on an absent version keeps its lock);
TestReplicaLockReconcileNullVersion (a null-version completion reconciles the null
version, not a coexisting UUID version, and an absent null version keeps its
accepted lock); TestAPIReplicaMarkerOnlyAppliesObjectLock; TestReplicaStoredLock;
the timestamp-only putReplicationOpts round trip; and the retransmit, exemption
and target-head tests. The #111 CopyObject replica suite and the existing #120
suite stay green, as do the PUT/multipart handler and object-layer regression
suites. Compatibility: the shared helper preserves #111's CopyObject behavior; a
non-replica PUT or multipart initiation that sets Object Lock now also stamps the
reserved ordering timestamp, matching CopyObject since #111; only trusted SSE-C
replica writes take the in-lock reconcile.

Refs pgsty/silo#120

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 87746913fc fix: retransmit existing SSE-C replicas instead of metadata-copying them
The replication sender's target HEAD carries no SSE-C customer key, so for
an SSE-C object the target answers 400 and replicateAll fell into a
metadata-only CopyObject that fails on any non-empty SSE-C object (the
undecryptable source checksum makes the target recompute one and rewrite
the data with a plaintext-sized reader). Once a non-empty SSE-C replica
existed, tag, retention and legal-hold changes never reached it, a heal
never retransmitted, and a resync neither repaired the replica nor
counted it correctly. Forcing a full retransmit alone was not enough:
checkPreconditionsPUT rejects a write whose PreserveETag and VersionID
match the stored version, only the single-part sealed ETag is truncated
before that comparison, so a multipart SSE-C retransmit answered 412,
which the sender turns into success. Inherited from upstream ad04afe38.

Select replicateAll when the SSE-C HEAD cannot answer (the two previous
assignments were dead: rAction still forced the metadata path), exempt an
authenticated replica write that carries an SSE-C seal from the duplicate
version and ETag rejection (the predicate is the incoming write's
restored SSE-C metadata, not what the destination holds), and send the
internal replication marker on the resync accounting HEAD for SSE-C
objects so a peer answers with the replica metadata instead of 400.

Tests: TestAPISSECReplicaRetransmitOverExistingVersion (multipart replica
initiation over the same version and ETag answered 412 on main, now 200
with parts sent and plaintext readback; single-part and zero-byte writes
unchanged), TestAPISSECReplicaWriteExemptionIsKeyedOnTheIncomingWrite
(plaintext replica over an SSE-C version still 412; SSE-C replica over a
plaintext version exempted and readable), and
TestAPISSECReplicationTargetHead (keyless HEAD 400, missing key 404,
marked HEAD 200 with metadata, metadata CopyObject ExcessData on a
non-empty object) on ErasureSD and Erasure. Compatibility: every update
of an SSE-C object now retransmits its bytes; a peer that rejects the
internal marker fails the accounting HEAD as before; the #109 destination
fix must be deployed first or a retransmitted replica is transformed
again.

Fixes pgsty/silo#120

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:11:04 +08:00
Feng Ruohang 7935c84f9a fix: recognize timestamp-only retention-removal tombstone in resend compare
retentionRemovedAtSource only recognized representation (1) of a removed
retention: the object lock key present with an empty value. But a removal
that arrived by replication persists representation (2): restoreRetention
(and the receiver's replica update path) writes only the retention ordering
timestamp when the mode is empty, leaving the mode and retain-until-date keys
absent. For that shape the helper returned false, so replicationActionForTarget
skipped the GetObjectRetention confirmation and let getReplicationAction's
replicateNone stand, silently dropping a needed removal when the destination
HEAD hides retention behind a permission-filtered credential.

Recognize representation (2) as well: a present retention ordering timestamp
with the mode value absent or empty is a removal. A present timestamp paired
with a non-empty mode is a retention that was set, not removed, and still
returns false.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-07 00:06:47 +08:00
Feng Ruohang c185635b43 fix: treat empty object lock values as absent when comparing
getReplicationAction builds its source map from oi1.UserDefined, where a
removed retention is a present key with an empty value, and its target map
from the destination's HEAD headers, which can never carry those keys because
setObjectHeaders skips empty lock values and FilterObjectLockMetadata drops
both keys when the mode is invalid. The comparison then always reports a
difference, the replicateNone fast path is dead for such versions, and an
otherwise matching version re-copies its metadata on every evaluation.

Skip an entry whose value is empty and whose key is x-amz-object-lock-mode or
x-amz-object-lock-retain-until-date, case-insensitively, in both comparison
loops, using the joined value on the target side. Normalizing only the source
would regress the case where both sides hold the empty pair.

HEAD also omits a real retention from a credential without
s3:GetObjectRetention, which the documented target policy does not grant, so
that normalization alone would read a destination hiding a retention as in
sync and drop the removal. replicationActionForTarget therefore confirms with
the destination before skipping the resend: only an explicit answer, no
retention on the version, clears it. Everything else keeps today's metadata
resend, including a denied or unreachable destination, a mode the SDK does not
recognize, and InvalidRequest, which names a bucket without Object Lock but is
also what a destination answers when its own read of that configuration fails.
The null version an existing object resync excludes is never reopened.

Tests: TestGetReplicationActionEmptyObjectLockValues (eight cases, red on 2
and 3 before this change), TestRetentionRemovedAtSource,
TestTargetRetentionConfirmedAbsent,
TestReplicationActionForTargetRetentionRemoval,
TestReplicationActionForTargetNullVersionResync and
TestEmptyRetentionValuesAreOmittedFromObjectResponseHeaders.
Compatibility: sender side only, no wire or storage change, so a fixed source
converges against any destination version.

Fixes pgsty/silo#117

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-07 00:06:47 +08:00
Feng Ruohang c201148738 Merge pull request #145 from pgsty/fix/issue-10-conditional-delete
fix: support conditional DeleteObject (If-Match) with atomic precondition
2026-09-06 23:58:42 +08:00
Feng Ruohang 53adb21c52 Merge pull request #142 from pgsty/fix/issue-141-resync-dispatch-scope
fix: scope resync worker dispatch to the target being resynced
2026-09-06 23:57:17 +08:00
Feng Ruohang 58b0ee36ca Merge pull request #140 from pgsty/fix/issue-139-resync-classification
fix: count resync success by replication outcome, not target existence
2026-09-06 23:56:57 +08:00
Feng Ruohang 8a1f594add Merge pull request #138 from pgsty/fix/issue-136-resync-counter-flush
fix: persist an honest resync terminal status about object counts
2026-09-06 23:55:59 +08:00
Feng Ruohang 5b9959617e Merge pull request #130 from pgsty/docs/issue-116-startup-readiness
docs: describe the startup readiness window of the health probes
2026-09-06 23:55:37 +08:00
Feng Ruohang da19d91b64 Merge pull request #143 from pgsty/fix/issue-107-chunked-checksum
fix: honor a header-delivered checksum advertised as a chunked trailer
2026-09-06 23:55:33 +08:00
Feng Ruohang 40bee4b7ba fix(delete): honor If-Match precondition on DeleteObject (#10)
DeleteObject ignored the If-Match request header and always deleted the
object (204). AWS S3 conditional deletes require that when If-Match is
provided and does not match the object's current ETag, the delete is
refused with 412 Precondition Failed and the object is left intact.

The precondition is evaluated in erasureServerPools.DeleteObject, while
the server-pool delete lock is held, before the delete-marker short-circuit
and before any version is removed. It runs against the version that will
actually be deleted: pinfo.ObjInfo for a normal delete, or the specifically
addressed version (read under the held lock) for a version-scoped delete,
since getPoolInfoExistingWithOpts strips VersionID. The check is a pure
function (no ResponseWriter writes) and returns PreConditionFailed, which
toAPIError maps to 412; CheckPrecondFn is cleared before lower layers run
so the precondition is evaluated exactly once.

Semantics:
- If-Match mismatch on a live object -> 412, object preserved.
- If-Match "*" requires a live object; a delete-marker-latest -> 412, and
  an explicitly addressed delete-marker version -> 412 (getObjectInfo
  returns the marker with MethodNotAllowed; the marker is the precondition
  target, not a 405).
- SSE-C/SSE-KMS: compared against the public ETag derived without the
  customer key, so a satisfiable condition is never falsely rejected.
- Explicit versionId -> evaluated against that version; a missing addressed
  version -> NoSuchVersion whether or not the key exists; a missing object
  (no versionId) -> NoSuchKey; no If-Match -> unchanged (including the
  unconditional version-scoped delete's error behavior).

Scope: atomicity is guaranteed for a single erasure set (the default
deployment). Multi-pool conditional-delete atomicity (concurrent writers
across pools, cross-pool version selection) is tracked as a follow-up.

Tests: pure-helper unit test (delete marker, "*", SSE-C without key);
object-layer tests (unversioned match/mismatch/missing, versioned
delete-marker-latest and addressed delete-marker version, explicit-version
selection, missing version on present and absent keys, read-quorum loss);
handler tests (412/204/wildcard/404) across both backends, with red/green
demonstrated per guard.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 23:05:27 +08:00
Feng Ruohang 6a9b5d6763 fix: honor header checksum when x-amz-trailer is advertised on non-trailer chunked PUT (#107)
The AWS Java SDK v2, with chunked encoding enabled (its default), sends a
PutObject as a non-trailer signed aws-chunked stream
(x-amz-content-sha256: STREAMING-AWS4-HMAC-SHA256-PAYLOAD). When a checksum
algorithm is set it puts the precomputed value in the x-amz-checksum-crc32
header, yet still advertises the checksum in x-amz-trailer even though no
trailer chunk is ever sent.

GetContentChecksum treated any x-amz-trailer-advertised checksum as trailing
with an empty value, deferring it to a trailer. For the non-trailer auth type
the handler sets req.Trailer = nil, so at EOF the hash.Reader looked the value
up in a nil trailer, got "", and returned XAmzContentChecksumMismatch (HTTP
400) even though the correct value sat in the request header. Real S3 accepts
the request, and disabling chunked encoding removed the trailer advertisement,
matching the reported symptom.

Honor the header value directly when a trailer-advertised checksum is already
present in the request headers; fall back to trailing delivery only when the
header is absent. When the header carries the checksum but it does not parse,
reject the request with ErrInvalidChecksum instead of falling through to a
no-validation path, so a malformed client-supplied checksum is never silently
dropped. This also restores the checksum echo on the response and the stored
value, while keeping genuine trailer uploads and wrong-checksum rejection
intact.

Fixes #107.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 22:48:29 +08:00
Feng Ruohang 0720ed477e fix(replication): scope resync dispatch to the target being resynced
The resync worker pool runs for a single target (opts.arn), but the dispatch
loop admitted any object whose ExistingObjResync.mustResync() was true for ANY
target. On a bucket with per-target rules (A and B), a resync of A would pull
in objects that only qualify for B - even with a single active resync, since
qualification is any-target. After the outcome-based classification (previous
change) such a cross-target object leaves A absent from its per-object result
and is counted as an A failure - an object A was never responsible for.

Scope admission to the resync's own target: dispatch an object only if it must
resync for opts.arn specifically (mustResyncTarget), via a small pure helper
objectNeedsResyncForARN. Only opts.arn carries this resync's ResetID, and that
reset is already folded into its per-target decision, so the per-target check
both scopes dispatch and honors the reset. Each target has its own resyncBucket,
so no cross-target object is dropped - it is handled by that target's resync.
The classifier's absent-ARN failure path is now unreachable for normally
dispatched objects and remains only as defense-in-depth (e.g. a config/lock
error before replication is attempted).

Delete-marker/version-purge handling, the null-version exclusion, the
finalization ordering and the outcome-based classification are unchanged.

Adds a table-driven regression for the predicate: with A/B rules and a resync
of A, an object qualifying only for B is not admitted; any-target scoping
admits it and fails the test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 21:57:52 +08:00
Feng Ruohang 46e82eb54d fix(replication): count resync success by outcome, not target existence
The resync worker classified each object by whether the target version
merely existed (a tgt.StatObject HEAD), ignoring the outcome of the
replicateObject/replicateDelete call it had just made. A quota-rejected
update leaves the old version in place, so StatObject succeeded and the
resync recorded a false success - reported as Completed / N success /
0 failed and persisted across restart (issue #139). #134's SSE-C HEAD
marker made StatObject succeed for SSE-C too, exposing it there. The delete
path had the mirror flaw (a failed delete leaves the object, so the HEAD
succeeded), and FailedSize was never incremented (a failed 196,608-byte
object counted as 1 failed / 0 bytes).

replicateObject and replicateDelete already build the per-target
replicatedInfos (each replicatedTargetInfo carries Arn, ReplicationStatus
and Err) but discarded it. Return it (callers that only trigger replication
ignore the value - a Go call statement discards it, so the queue paths are
unchanged) and classify the resync from the target whose Arn == opts.arn via
a small pure helper:

- Completed without error -> replicated (+ that target's size, falling back
  to the object size).
- Failed or errored -> failed (+ the object size, fixing FailedSize).
- opts.arn absent from the result (not attempted) -> failed; a resync that
  cannot confirm the object reached the target is not a success.

The StatObject-existence block (including the delete-marker/MethodNotAllowed
special case, now subsumed by the delete outcome) is removed. #134's SSE-C
HEAD marker is left intact - it is needed for genuine SSE-C success.

Adds a table-driven regression for the classifier covering a completed
update, a failed update over an existing version, an errored-but-Completed
result, a delete failure, a delete-marker success (zero bytes) and an
un-attempted ARN. Classifying by existence makes the failed cases count
success and fails the test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 18:48:15 +08:00
Feng Ruohang f8ca4a8656 fix(replication): keep resync Completed status honest about object counts
resyncBucket could publish and persist a Completed resync status that did
not actually cover every object, in two ways:

1. It joined only the producer workers before the deferred markStatus ran,
   not the goroutine that folds each worker result into the status, so a
   Completed status could omit the last object (or a failed object) until the
   periodic ~1m flush (issue #136). The same finalization also closed the
   result channel on early-return paths while workers were still in flight,
   risking a send-on-closed-channel panic and a lost result.
2. markStatus persists under its own background context, so if the parent
   context was cancelled during the drain - workers then return without
   sending their computed result - or a worker dropped a result on the
   resync-cancel signal, a bare Completed was still recorded with counts that
   no longer matched the objects seen.

Fixes (count integrity only; the inherited cancellation deadlock, walker leak,
and single-token routing are tracked as separate follow-ups):

- Centralize shutdown in a resyncResults helper whose finish() stops the
  workers (closes inputs, waits for them to exit) before closing the result
  channel and waiting for the consumer to drain, then lets the deferred
  markStatus persist the final counts. finish() now runs on every exit path.
- Record a dropped result via sendResyncResult (a worker consuming the
  resync-cancel token returns without sending), and in the finalizer downgrade
  a Completed status to Failed via finalResyncStatus when the parent context
  was cancelled or a worker aborted - so a persisted Completed never
  misrepresents an incomplete resync.

Deterministic tests: an on-disk round-trip of the terminal status (complete
counts stay Completed; parent-cancel-during-drain and worker-abort each
downgrade to Failed), and testing/synctest drain/worker-order assertions that
fail deterministically if a finish() wait is removed. The inherited
cancellation structure (inline Walk, the dispatch send, the worker cancel
branches) is left unchanged for the follow-ups.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 15:28:55 +08:00
Feng Ruohang e7d0e62f16 Merge pull request #135 from pgsty/fix/attributes-encrypted-parts-test-reconcile
test: reconcile encrypted-parts attributes test with #119 write validation
2026-09-06 09:59:54 +08:00
Feng Ruohang aee290fc34 test: reconcile encrypted-parts attributes test with #119 write validation
TestAPIGetObjectAttributesEncryptedPartLengths (from #128) built its
fixtures by PutObjectPart-ing plaintext bodies under encrypted-object
metadata with per-part sizes 5245473 and 1. Since #119, PutObjectPart
always derives an encrypted part's plaintext length from the bytes
written and rejects a part that cannot be a valid sio stream, so those
fixtures can no longer be created through a normal write and both
variants failed at write time.

Such an on-disk shape now only exists as pre-#119 data or from an old
peer, which is exactly the state the GetObjectAttributes per-part
tamper check (#128) defends. Inject that ObjectInfo directly through a
stub object layer (the setObjectLayer pattern used by the #110 tamper
test) and exercise the handler, which is what this test pins. The
handler path, the crafted part sizes, and both assertions
(separately-encrypted-parts -> ErrObjectTampered; legacy-single-stream
-> stored fragment sizes) are unchanged. Test-only; reconciles two
already-merged correct changes (#119 and #128).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 09:51:09 +08:00
Feng Ruohang 62cce2b152 Merge pull request #121 from pgsty/fix/issue-110-tampered-status
fix: return 500 for unreadable objects instead of 206
2026-09-06 09:26:28 +08:00
Feng Ruohang 65d4806a7b Merge pull request #127 from pgsty/fix/issue-112-bucket-metadata-cors
fix: include per-bucket CORS in bucket metadata export and import
2026-09-06 09:26:21 +08:00
Feng Ruohang 765757473a Merge pull request #126 from pgsty/fix/issue-118-no-compressed-ssec
fix: exclude SSE-C objects from compression
2026-09-06 09:26:12 +08:00
Feng Ruohang 425bd7fff1 Merge pull request #124 from pgsty/fix/issue-119-ssec-part-actual-size
fix: record plaintext part sizes for replicated SSE-C multipart parts
2026-09-06 09:26:04 +08:00
Feng Ruohang e12e739a53 Merge pull request #128 from pgsty/fix/issue-114-115-object-attributes-parts
fix: report logical part sizes and end pagination correctly in GetObjectAttributes
2026-09-06 09:25:28 +08:00
Feng Ruohang 8b736dee34 Merge pull request #123 from pgsty/fix/issue-113-rotation-checksum-algorithm
fix: honor a requested checksum algorithm on SSE-C key rotation
2026-09-06 09:25:02 +08:00
Feng Ruohang 32e75c27bf Merge pull request #122 from pgsty/fix/issue-109-raw-ssec-replica
fix: store raw SSE-C replicas verbatim on the destination
2026-09-06 09:24:26 +08:00
Feng Ruohang 885ca604a1 Merge pull request #129 from pgsty/fix/issue-111-object-lock-replica-ordering
fix: order value-less replicated Object Lock updates by timestamp
2026-09-06 09:23:34 +08:00
Feng Ruohang d10382d0dc build: refresh rebrand compatibility baseline for the SSE-C replica helper
The raw SSE-C seal headers moved from an inline slice in
cmd/object-multipart-handlers.go to the shared isRawSSECReplica helper
in cmd/replication-trust.go (already recorded), so regenerate the
compat-baseline allowlist to drop the three stale object-multipart
entries. No behaviour change; make rebrand-guard is green.

Refs pgsty/silo#109

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-06 09:08:22 +08:00
mr javad seydi 0db4bf3b00 fix: report effective bucket quotas in metrics
Signed-off-by: mr javad seydi <seydi.birjand@gmail.com>
2026-09-05 21:23:46 +03:30
Feng Ruohang 87c621965d docs: describe the startup readiness window of the health probes
After a restart a node's remote erasure drives that could not be connected
during startup stay uninstalled until the next connectDisks pass, about 15
seconds later. During that window the liveness, readiness and both cluster
probes answer 200, admin info shows every drive online and mcli ready agrees,
because the cluster probes aggregate each peer's report of its own local
drives rather than the drives this node has installed. A PUT through that
node can still fail with 503 SlowDownWrite and a cross-node GET can answer
404 NoSuchKey until the window closes. The behaviour is inherited from
upstream and reproduced on the 0806 and 0903 releases alike.

Document the window and the bounded data-path check (PUT through each node,
read each object through every node, fixed deadline, re-read acknowledged
objects) that automation should use instead of the probes, and correct the
readiness probe description, which also fails on request-queue overload and
an unreachable KMS. No product change: the probes keep their documented
purpose, and changing them or the reconnect cadence was judged unproven
tuning in the agreed plan.

Refs pgsty/silo#116

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 17:17:05 +08:00
Feng Ruohang b2dca43fda fix: order value-less replicated Object Lock updates by timestamp
CopyObjectHandler rebuilt the destination metadata with the public Object Lock keys stripped (cmd/object-handlers.go:1708) and then restored a value only inside retentionMode.Valid() and legalHold.Status.Valid() (cmd/object-handlers.go:1715 and :1732), so a replica update that carried no retention or legal-hold value never reached the ordering comparison and silently erased whatever the destination held, however new it was; a retention removal that did win recorded no ordering timestamp either, so cmd/bucket-object-lock.go:370 later read an unparseable stored timestamp and let an older retained value back in. Each replica field is now decided on its source timestamp first and its incoming value second, and both restore helpers write the stored timestamp back before returning early on an empty stored value, which is the only way a removal timestamp survives the REPLACE metadata directive. Legal hold stays deliberately asymmetric: S3 has no legal-hold removal, an explicitly empty status is already rejected as invalid, and an absent status conveys no change even when an orphaned timestamp arrives with it, so only a valid ON or OFF can win.

Three inherited defects would have defeated that ordering, so they are fixed here too. The SSE-KMS branch of putOptsFromHeaders built its own ObjectOptions and dropped the parsed lock timestamps, leaving every replicated lock update unordered on a bucket with default KMS encryption; it now carries them. The in-place SSE-C key rotation snapshots the stored reserved metadata into encMetadata before the lock decision exists and merges it back afterwards to preserve the encryption headers, reinstating the ordering timestamp the decision had just replaced; the snapshot is now reconciled with the decision for a trusted replica. Finally, the value-less handling applies only to an actual replica: a trusted peer that sends the replication marker without REPLICA status keeps the previous behaviour, so a REPLACE copy carrying no lock headers still writes a version with no retention and no hold.

Tests: TestAPICopyObjectReplicaAbsentLockFieldsPreserveNewerState, TestAPICopyObjectReplicaRetentionRemovalKeepsOrderingTimestamp, TestAPICopyObjectReplicaObjectLockOrdering, TestAPICopyObjectReplicaRetentionRemovalUnderBucketKMS, TestAPICopyObjectReplicaLockTimestampSurvivesSSECKeyRotation and TestAPICopyObjectMarkerOnlyLeavesObjectLockUnchanged, all on ErasureSD and Erasure. Compatibility: no API, wire or stored-field change, and a field arriving with no source timestamp is unordered and now preserves destination state, so an un-upgraded 0806 peer keeps replicating safely while it still runs the old erasing receiver.

Fixes pgsty/silo#111

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 16:58:11 +08:00
Feng Ruohang 33a91d972f fix: carry per-bucket CORS through bucket metadata export/import
The admin bucket-metadata handlers enumerate every bucket config by name, and
per-bucket CORS was never added to that enumeration: export omitted cors.xml
(cmd/admin-bucket-handlers.go:414 cfgFiles) and import ignored the entry
outright, with no case in applyImportedBucketMetadata (:598) or SetStatus
(:629), so a CORS-only archive reported 0/0 buckets imported and a restored
bucket silently lost its configuration. Export now writes the stored document
verbatim and import validates it with the same parser and validator as
PutBucketCorsHandler, merging it under the existing bucket metadata lock and
announcing it through the dedicated SRBucketMetaTypeCorsConfig event; the local
CORS timestamp rule is extracted into localCORSUpdatedAt and reused so an
imported document always lands strictly above bucket creation, which matters
because the import stamps its fields before creating any missing bucket and a
CORS event below Created is dropped as an older bucket incarnation.

The import reads one byte past the declared entry size so archive/zip reaches
EOF and verifies the entry checksum, otherwise a corrupt or over-long entry
carrying well formed XML would overwrite the stored document; and the CORS
event is sent even when the shared bucket metadata hook failed, so an
unreachable peer cannot withhold an already committed CORS document from the
reachable ones.

Tests: TestAdminBucketMetadataCORSRoundTrip and
TestAdminBucketMetadataCORSImportReplicatesPastPeerFailure (new, ErasureSD and
Erasure).
Compatibility: the ZIP gains one entry, older archives stay importable and
leave CORS untouched; no mcli or madmin-go change is needed because
madmin.BucketStatus already carries Cors and mcli copies the export ZIP
verbatim.

Fixes pgsty/silo#112

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 16:21:12 +08:00
Feng Ruohang 2cbd48a3c3 fix: end GetObjectAttributes part pagination correctly
GetObjectAttributes decided truncation by comparing the last returned
part number with the part count (cmd/object-handlers.go:720), which is
only a coincidence of contiguous numbering. Sparse parts 1/3 reported a
complete page as truncated with a marker that loops, and parts 1/3/5
with max-parts=1 stopped after part 3 and silently dropped part 5. Set
IsTruncated in the break that proves an eligible part was left
unreturned, and zero NextPartNumberMarker when the listing is complete,
as ListObjectParts already does. Also reject negative x-amz-max-parts
and x-amz-part-number-marker in getAndValidateAttributesOpts with the
same API errors ListObjectParts uses, instead of answering an invalid
request with an empty parts listing; an absent or zero max-parts still
means the default page size.

Tests: TestAPIGetObjectAttributesPartsPagination (sparse 1/3/5 and
contiguous 1/2 walks on ErasureSD and Erasure),
TestGetAndValidateAttributesOptsPartsRange, and the sparse variants of
TestAPIGetObjectAttributesMultipartLogicalPartSize. Compatibility: no
field is added or removed; IsTruncated and NextPartNumberMarker change
only where they were wrong, and negative pagination values that no SDK
sends now fail fast.

Fixes pgsty/silo#115

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 16:17:39 +08:00
Feng Ruohang b5409ca112 fix: report logical part sizes in GetObjectAttributes
GetObjectAttributes filled ObjectPart.Size from the on-disk part length
(cmd/object-handlers.go:715), so every compressed or encrypted multipart
object reported transformed sizes that do not sum to the logical
ObjectSize the same response returns from objInfo.GetActualSize().
Report each part's uploaded plaintext length instead: a compressed part
uses its recorded ActualSize, and a separately encrypted part derives the
plaintext length with sio.DecryptedSize, because ActualSize is the
ciphertext length for a replicated SSE-C part and zero for parts written
before actualSize existed.

Only parts of an encrypted multipart object are streams of their own. A
legacy encrypted object carries no multipart marker and is one continuous
stream that the erasure writer split into storage fragments, so those
fragments keep their stored size. Where a part is a stream, one whose
length cannot be a valid encrypted stream has no logical length, and the
request now fails with XMinioObjectTampered rather than reporting the
ciphertext length; DecryptObjectInfo does not catch that case, because
ObjectInfo.isMultipart gives up on the first bad part and only the object
total is then validated.

Tests: TestAPIGetObjectAttributesMultipartLogicalPartSize (plain,
compressed, SSE-C and compressed+SSE-C, consecutive and sparse part
numbers), TestAPIGetObjectAttributesCompressedEmptyTrailingPart,
TestAPIGetObjectAttributesEncryptedPartLengths, and a part-size assertion
added to TestAPISSECMultipartReplicationTrust. Compatibility: the XML
shape is unchanged and nothing is written to disk, only the value of the
existing Size element is corrected.

Fixes pgsty/silo#114

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 16:17:39 +08:00
Feng Ruohang 35bd75948a fix: exclude SSE-C objects from compression
With compression allow_encryption=on an SSE-C object is stored as
encrypt(s2(plaintext)), while replication reads it raw (NoDecryption at
cmd/erasure-object.go:257) and putReplicationOpts drops the internal
compression and actual-size headers (cmd/bucket-replication.go:786). The
replica keeps the source seal with no compression marker, so a GET with the
correct customer key returns HTTP 200 and the raw S2 stream instead of the
object, and the source records the transfer as COMPLETED.

Widen the one condition in excludeForCompression (cmd/object-api-utils.go:613)
so SSE-C is never compressed, whatever allow_encryption says. This covers all
four producers at once, PutObject, NewMultipartUpload, CopyObject and
PutObjectExtract, plus any future caller of isCompressible.
crypto.SSEC.IsRequested ignores copy-source headers, so a copy is judged on its
destination key only, and a raw SSE-C replica write is unaffected because it
carries no public SSE-C headers. allow_encryption keeps its meaning for SSE-S3
and SSE-KMS, where the server owns the key and decompresses before replicating.

Tests: TestAPISSECCompressionReplicaStaysReadable (single PUT and multipart),
TestAPISSECCompressionProducerMatrix, TestAPISSECCompressionSkippedOnCopyObject,
TestAPISSECCompressionSkippedOnSnowballExtract and the control
TestSSECBatchReplicationCannotRead in cmd/compression-ssec_test.go. Two existing
expectations pinned the removed shape and are updated:
TestAPICopyObjectSSECKeyRotationNullVersionCompressesRewrite is renamed
TestAPICopyObjectSSECKeyRotationNullVersionSkipsCompression and now expects an
uncompressed rewrite, keeping its body, checksum, version and ETag assertions;
the SSE-C compressed-encrypted variant of
TestAPICopyObjectServerSideChecksumEncryption becomes compressible-extension and
expects an uncompressed destination, its SSE-S3 sibling keeping the compressed
coverage.

Compatibility: a deliberate behaviour change. Deployments with
allow_encryption=on no longer store new or rewritten SSE-C data compressed, so
those writes cost more space; objects already stored compressed keep working on
the source and are the concern of pgsty/silo#109, which rejects them at
replication time. Multipart uploads initiated before this change keep
compressing their parts from the metadata saved at initiation. Upstream
468a9fae8 refused this combination at PUT time and a2cab0255 removed the guard;
upstream master is still unguarded, so this is a deliberate divergence.

Fixes pgsty/silo#118

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 15:51:24 +08:00
Feng Ruohang 0c8d74205b fix: record plaintext part sizes for replicated SSE-C multipart parts
Trusted SSE-C replication uploads parts as raw ciphertext with the
ciphertext length as Content-Length, and erasureObjects.PutObjectPart only
derived the plaintext length when the caller passed a negative size, so
each replicated part persisted the ciphertext length as ActualSize (the
field defined as the uploaded size without encryption bytes). On the
replica, partNumberToRangeSpec turned those lengths into a plaintext range,
so GET/HEAD ?partNumber=N returned the wrong bytes and shifted
Content-Range (2560, 2560 and 5120 bytes for 5 MiB, 5 MiB and 1 MiB
parts), and a later decommission or rebalance re-uploaded the parts with
the stale value and recomputed the object-level actual-size from their sum,
after which a whole-object GET advertised a Content-Length larger than the
body it wrote.

Derive the plaintext length of an encrypted, uncompressed part from the
bytes actually written (sio.DecryptedSize) in PutObjectPart, the single
place a part is persisted, rejecting a length that cannot be a valid
stream before the part is committed; and derive part lengths from
part.Size in partNumberToRangeSpec for encrypted, uncompressed objects,
returning an error instead of a nil range, so replicas already on disk
read correctly without a resync. Compressed parts keep ActualSize.

Tests: TestAPISSECReplicaPartNumberReads (three-part SSE-C replica,
?partNumber=N bytes, Content-Length and Content-Range equal the source)
and TestSSECReplicaPartActualSizeDataMovement (replay through the data
movement path leaves part and object sizes at plaintext values) fail on
main and pass with the fix on ErasureSD and Erasure;
TestAPIGetObjectWithPartNumberHandler, TestAPISSECMultipartReplicationTrust
and TestAPIListObjectPartsHandler stay green. Compatibility: no wire or
API change; objects an unfixed server already moved carry a poisoned
object-level actual-size and need a rewrite or resync.

Fixes pgsty/silo#119

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 15:33:23 +08:00
Feng Ruohang fcc4d77895 fix: honor a requested checksum algorithm on SSE-C key rotation
An in-place SSE-C key rotation takes the fast path at cmd/object-handlers.go:1523
that only rewraps the object key, while every line that turns
x-amz-checksum-algorithm into a stored checksum lives in the re-encrypting else
branch at 1571-1609, so a requested algorithm was silently dropped and the stale
source checksum was kept and reported. Extend the canRotateKeyInPlace guard so a
client request carrying the header falls through to the copy that recomputes,
stores and reports it. Replica-trusted requests keep the fast path: getOpts
leaves their source reader encrypted, so a rewrite would hash ciphertext, and a
replica has to keep the checksum its source assigned.

Tests: TestAPICopyObjectSSECKeyRotationChecksumAlgorithm (new, red before the
guard), TestAPICopyObjectSSECKeyRotationKeepsChecksumAbsence (new, pins the
accepted limitation that a headerless rotation preserves the stored checksum
state including absence, gaining no default CRC64NVME) and
TestAPICopyObjectSSECKeyRotationReplicaKeepsFastPath (new, pins the replica
carve-out on a non-empty and on a zero byte source).
Compatibility: no API or wire change; a rotation without the header and every
replica-trusted rotation are unchanged, while a client rotation carrying the
header now rewrites the object data, so the ETag changes, a multipart source
collapses to a single part object, the copy replicates as an object rather than
as metadata, and the rewritten bytes are compressed if compression is enabled for
that object, as AWS CopyObject documents. Upstream MinIO carries the same
defect from 2718d9a43 (minio/minio#21399); this is a deliberate divergence.

Fixes pgsty/silo#113

Signed-off-by: Feng Ruohang <rh@vonng.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
2026-09-05 15:25:08 +08:00
Feng Ruohang c52acc1a5d fix: store raw SSE-C replicas verbatim on the destination
A raw SSE-C replica write carries the source ciphertext and the source seal
in X-Minio-Replication-Server-Side-Encryption-* headers but no public SSE-C
request headers, so crypto.Requested() was false and PutObjectHandler and
NewMultipartUploadHandler applied the destination's default encryption and
compression to bytes that were already ciphertext (upstream 468a9fae8,
"Enable replication of SSE-C objects", never exempted the raw path). With
destination default SSE-S3 the replica's IV and seal were overwritten and
GET returned 400; with destination compression the replica stored
compress(ciphertext) and GET failed, while the source reported COMPLETED.

Recognize a validated raw SSE-C replica (replicaTrusted plus a seal header,
shared helper isRawSSECReplica) and skip bucket default encryption,
compression and the encryption branch on the single PUT path, and default
encryption plus compression on the multipart initiation path, which already
skipped key generation. On the sender, reject replication of an object that
is both compressed and SSE-C, since the wire carries no compression state
and the destination would otherwise store an undetectable S2 stream, and
make replicateObject/replicateAll report a putReplicationOpts failure as
Failed instead of Completed.

Tests: TestAPISSECReplicaSkipsDestinationTransforms (single PUT and
multipart under destination default SSE-S3, compression and an explicit SSE
header, plus an untrusted control), TestAPISSECMultipartReplicaRoundTripWith
Compression, and TestPutReplicationOptsRejectsCompressedSSEC fail on main
and pass with the fix on ErasureSD and Erasure; the replication-trust,
multipart and PutObject suites stay green. Compatibility: no wire, API or
metadata change; the destination change applies only to trusted replica
writes carrying a source seal; replicas already transformed must be
rewritten from an intact source (see pgsty/silo#120 for why a resync does
not do that yet).

Fixes pgsty/silo#109

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 15:07:22 +08:00
Feng Ruohang 5703426b3c fix: return 500 for unreadable objects instead of 206
ErrObjectTampered was mapped to http.StatusPartialContent since upstream
ca6b4773e (2017), so a GET or HEAD of an object the server cannot decode
(invalid encrypted size, malformed actual-size, bad multipart ETag shape)
answered with a success status and an XML error document that SDKs handed
back as object content; boto3 returned the XML as Body and a zero-length
HEAD as success. Every origin of errObjectTampered is a stored-state
defect, not caller input, so map the entry to 500 Internal Server Error
and keep the XMinioObjectTampered code and message. The comment records
the deliberate divergence from upstream.

Tests: TestObjectTamperedGETHEADStatus (signed GET and HEAD, ErasureSD and
Erasure) fails with 206 on main and passes with 500; TestAPIErrCode,
TestAPIErrCodeDefinition, TestAPIHeadObjectHandler,
TestAPIHeadObjectHandlerWithEncryption and TestAPIGetObjectHandler stay
green. Compatibility: only the status line of one MinIO-specific error
changes; clients now retry damaged-object reads per their 5xx policy.

Fixes pgsty/silo#110

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7qJqWwy8oFA6aCXWRzXQe
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-05 14:50:00 +08:00
Feng Ruohang f0bd164b92 docs: refresh community contributors
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 19:21:26 +08:00
Feng Ruohang 9936a69d89 ci: recover container publication from verified component pins
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 15:10:43 +08:00
Feng Ruohang ce2326c946 build: pin the published mcli 20260903 archives
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 14:38:54 +08:00
Feng Ruohang 4c164907f5 build: align Helm defaults with the published release tag
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-04 14:36:12 +08:00
mr javad seydi 7a060cab1e feat(ilm): relocate hot objects across server pools by GET frequency
Keep NVMe/HDD pool pairs useful without remote tiering: promote objects that
are read often, demote previously moved objects once they go idle, and leave
the feature off until operators set a two-pool topology.

Signed-off-by: mr javad seydi <seydi.birjand@gmail.com>
2026-08-15 13:48:22 +03:30
244 changed files with 38326 additions and 1877 deletions
-54
View File
@@ -1,54 +0,0 @@
---
name: Bug report
about: Create a report to help us improve
title: ''
labels: community, triage
assignees: ''
---
## NOTE
Silo issues are handled by community maintainers on a best-effort basis. There
is no SLA, SLO, or emergency production-support channel. Follow the local
[Code of Conduct](../code_of_conduct.md) when participating. Report suspected
vulnerabilities through the private process in [SECURITY.md](../SECURITY.md),
not in a public issue.
<!--- Provide a general summary of the issue in the Title above -->
## Expected Behavior
<!--- If you're describing a bug, tell us what should happen -->
<!--- If you're suggesting a change/improvement, tell us how it should work -->
## Current Behavior
<!--- If describing a bug, tell us what happens instead of the expected behavior -->
<!--- If suggesting a change/improvement, explain the difference from current behavior -->
## Possible Solution
<!--- Not obligatory, but suggest a fix/reason for the bug, -->
<!--- or ideas how to implement the addition or change -->
## Steps to Reproduce (for bugs)
<!--- Provide a link to a live example, or an unambiguous set of steps to -->
<!--- reproduce this bug. Include code to reproduce, if relevant -->
<!--- and include relevant Silo logs with secrets and credentials removed -->
1.
2.
3.
4.
## Context
<!--- How has this issue affected you? What are you trying to accomplish? -->
<!--- Providing context helps us come up with a solution that is most useful in the real world -->
## Regression
<!-- Is this issue a regression? (Yes / No) -->
<!-- If Yes, optionally include the Silo version, commit id, or PR that caused this regression. -->
## Your Environment
<!--- Include as many relevant details about the environment you experienced the bug in -->
* Version used (`silo --version`):
* Server setup and configuration:
* Operating System and version (`uname -a`):
+8
View File
@@ -7,6 +7,14 @@ assignees: ''
---
Report bugs in the PGSTY SILO server (`pgsty/silo`) here. Community maintainers
handle reports on a best-effort basis. There is no SLA, SLO, or emergency
production-support channel. Follow the
[Code of Conduct](https://github.com/pgsty/silo/blob/main/code_of_conduct.md).
Report suspected vulnerabilities privately through
[SECURITY.md](https://github.com/pgsty/silo/blob/main/SECURITY.md).
For patches, see the [contribution guide](https://github.com/pgsty/silo/blob/main/CONTRIBUTING.md).
<!--- Provide a general summary of the issue in the Title above -->
## Expected Behavior
@@ -7,6 +7,9 @@ assignees: ''
---
Suggest improvements to the PGSTY SILO server (`pgsty/silo`) here.
For patches, see the [contribution guide](https://github.com/pgsty/silo/blob/main/CONTRIBUTING.md).
**Is your feature request related to a problem? Please describe.**
A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]
+7 -3
View File
@@ -1,10 +1,14 @@
## Contribution Licensing (no CLA, inbound=outbound, DCO required)
This project does not use a CLA; contributions are accepted inbound=outbound.
This pull request contributes to PGSTY SILO (`pgsty/silo`). Code contributions
are accepted under AGPL-3.0-or-later, the same license as the server.
This project does not use a CLA or require a separate Apache-2.0 license grant.
By submitting this pull request I represent that I have the right to contribute
the changes, which are licensed under this repository's
the code changes under this repository's
[GNU Affero General Public License v3.0 or later](https://www.gnu.org/licenses/agpl-3.0.html)
and remain my copyright. Every commit must carry a DCO `Signed-off-by` trailer
and retain copyright in my original work. Existing copyright and license
notices remain intact; separately licensed material keeps its applicable terms.
Every commit must carry a DCO `Signed-off-by` trailer
(`git commit -s`) certifying the
[Developer Certificate of Origin](https://developercertificate.org/) — see
[CONTRIBUTING.md](https://github.com/pgsty/silo/blob/main/CONTRIBUTING.md).
+60 -3
View File
@@ -7,6 +7,11 @@ on:
description: "Published RELEASE.* tag to package as pgsty/silo"
required: true
type: string
recovery:
description: "Run the current main workflow against an already-published tag"
required: false
default: false
type: boolean
permissions:
contents: read
@@ -75,12 +80,18 @@ jobs:
fetch-depth: 0
- name: Verify workflow identity matches release source
env:
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
RECOVERY: ${{ inputs.recovery }}
run: |
set -euo pipefail
CHECKED_OUT_REVISION="$(git rev-parse HEAD)"
if [ "${CHECKED_OUT_REVISION}" != "${GITHUB_SHA}" ]; then
echo "Checked out ${CHECKED_OUT_REVISION}, but workflow identity is ${GITHUB_SHA}. Dispatch this workflow from ${RELEASE_TAG}." >&2
exit 1
if [ "${RECOVERY}" != "true" ] || [ "${GITHUB_REF}" != "refs/heads/${DEFAULT_BRANCH}" ]; then
echo "Checked out ${CHECKED_OUT_REVISION}, but workflow identity is ${GITHUB_SHA}. Dispatch from ${RELEASE_TAG}, or use recovery from ${DEFAULT_BRANCH}." >&2
exit 1
fi
echo "Recovery workflow ${GITHUB_SHA} is packaging published source ${CHECKED_OUT_REVISION}."
fi
- name: Prepare verified Docker contexts
@@ -129,10 +140,50 @@ jobs:
mkdir -p "${context}/dockerscripts"
tar -xzf "${archive}" -C "${context}" silo
cp Dockerfile.goreleaser Dockerfile.distroless LICENSE NOTICE CREDITS "${context}/"
cp dockerscripts/docker-entrypoint.sh dockerscripts/download-static-curl.sh \
cp dockerscripts/docker-entrypoint.sh dockerscripts/build-static-curl.sh \
"${context}/dockerscripts/"
done
# The classic image bundles mcli. Resolve its two archive digests
# from the immutable published release instead of trusting defaults
# copied into an older Server tag. This also gives a recovery run a
# narrow override when a tag selected the right mcli release but
# accidentally retained stale archive pins.
MC_REPO="$(awk -F= '/^ARG MC_REPO=/{print $2; exit}' Dockerfile.goreleaser)"
MC_VERSION="$(awk -F= '/^ARG MC_VERSION=/{print $2; exit}' Dockerfile.goreleaser)"
test -n "${MC_REPO}"
test -n "${MC_VERSION}"
MC_VERSION_HYPHEN="${MC_VERSION#RELEASE.}"
MC_PKG_VERSION="$(echo "${MC_VERSION_HYPHEN}" | sed -E 's/^([0-9]{4})-([0-9]{2})-([0-9]{2})T([0-9]{2})-([0-9]{2})-([0-9]{2})Z$/\1\2\3\4\5\6.0.0/')"
if [ "${MC_PKG_VERSION}" = "${MC_VERSION_HYPHEN}" ]; then
echo "Invalid bundled mcli tag: ${MC_VERSION}" >&2
exit 1
fi
if [ "$(gh release view "${MC_VERSION}" --repo "${MC_REPO}" --json isDraft --jq .isDraft)" != false ] || \
[ "$(gh release view "${MC_VERSION}" --repo "${MC_REPO}" --json isPrerelease --jq .isPrerelease)" != false ] || \
[ "$(gh release view "${MC_VERSION}" --repo "${MC_REPO}" --json isImmutable --jq .isImmutable)" != true ]; then
echo "Bundled mcli ${MC_REPO}@${MC_VERSION} must be a published immutable release" >&2
exit 1
fi
mc_checksums="mcli_${MC_PKG_VERSION}_checksums.txt"
gh release download "${MC_VERSION}" --repo "${MC_REPO}" \
--dir "${assets_dir}" --pattern "${mc_checksums}"
gh attestation verify "${assets_dir}/${mc_checksums}" \
--repo "${MC_REPO}" \
--signer-workflow "${MC_REPO}/.github/workflows/release.yml" \
--source-ref "refs/tags/${MC_VERSION}" >/dev/null
MC_AMD64_SHA256="$(awk -v name="mcli_${MC_PKG_VERSION}_linux_amd64.tar.gz" '$2 == name {print $1}' "${assets_dir}/${mc_checksums}")"
MC_ARM64_SHA256="$(awk -v name="mcli_${MC_PKG_VERSION}_linux_arm64.tar.gz" '$2 == name {print $1}' "${assets_dir}/${mc_checksums}")"
[[ "${MC_AMD64_SHA256}" =~ ^[0-9a-f]{64}$ ]]
[[ "${MC_ARM64_SHA256}" =~ ^[0-9a-f]{64}$ ]]
{
echo "MC_AMD64_SHA256=${MC_AMD64_SHA256}"
echo "MC_ARM64_SHA256=${MC_ARM64_SHA256}"
} >> "${GITHUB_ENV}"
echo "RELEASE_REVISION=$(git rev-parse HEAD)" >> "${GITHUB_ENV}"
- name: Set up QEMU
@@ -162,6 +213,9 @@ jobs:
file: docker-release/amd64/Dockerfile.goreleaser
platforms: linux/amd64
push: true
build-args: |
MC_AMD64_SHA256=${{ env.MC_AMD64_SHA256 }}
MC_ARM64_SHA256=${{ env.MC_ARM64_SHA256 }}
tags: |
pgsty/silo:${{ env.RELEASE_TAG }}-amd64
pgsty/silo:latest-amd64
@@ -178,6 +232,9 @@ jobs:
file: docker-release/arm64/Dockerfile.goreleaser
platforms: linux/arm64
push: true
build-args: |
MC_AMD64_SHA256=${{ env.MC_AMD64_SHA256 }}
MC_ARM64_SHA256=${{ env.MC_ARM64_SHA256 }}
tags: |
pgsty/silo:${{ env.RELEASE_TAG }}-arm64
pgsty/silo:latest-arm64
+24 -2
View File
@@ -10,7 +10,7 @@ on:
- "Dockerfile.distroless"
- "cmd/healthcheck-main.go"
- "cmd/main.go"
- "dockerscripts/download-static-curl.sh"
- "dockerscripts/build-static-curl.sh"
- "dockerscripts/docker-entrypoint.sh"
- "dockerscripts/docker-entrypoint_test.sh"
- "silo.service"
@@ -41,6 +41,28 @@ permissions:
contents: read
jobs:
curl:
name: Static curl (${{ matrix.arch }})
runs-on: ${{ matrix.runner }}
strategy:
matrix:
include:
- arch: amd64
runner: ubuntu-latest
- arch: arm64
runner: ubuntu-24.04-arm
steps:
- uses: actions/checkout@v7
- name: Build and exercise curl in an empty runtime
run: |
docker build --build-arg TARGETARCH=${{ matrix.arch }} \
--target curl-runtime -f Dockerfile.goreleaser -t silo-curl-test .
docker run --rm silo-curl-test --version | tee curl-version.txt
grep -F 'curl 8.22.0 ' curl-version.txt
grep -F 'HTTP2' curl-version.txt
docker run --rm silo-curl-test --fail --silent --show-error \
--connect-timeout 15 --max-time 60 https://curl.se/robots.txt
validate:
runs-on: ubuntu-latest
steps:
@@ -490,7 +512,7 @@ jobs:
bash -n buildscripts/package/lifecycle_test.sh
buildscripts/package/lifecycle_test.sh
bash -n dockerscripts/docker-entrypoint_test.sh
bash -n dockerscripts/download-static-curl.sh
bash -n dockerscripts/build-static-curl.sh
dockerscripts/docker-entrypoint_test.sh
go run ./buildscripts/rebrand-guard
buildscripts/verify-rebrand.sh
+1 -1
View File
@@ -29,7 +29,7 @@ jobs:
- name: Install govulncheck
run: |
go install golang.org/x/vuln/cmd/govulncheck@v1.7.0
go install golang.org/x/vuln/cmd/govulncheck@v1.8.0
echo "$(go env GOPATH)/bin" >> "${GITHUB_PATH}"
- name: Run govulncheck
+14 -8
View File
@@ -79,9 +79,10 @@ documentation is owned by the separate
## Licensing of Contributions
Silo is licensed under the [GNU AGPL v3.0 or later](LICENSE). Its core is
Copyright (c) MinIO, Inc.; the combined work can never be relicensed, and this
fork does not try to.
Code contributions to PGSTY SILO (`pgsty/silo`) are accepted under the
[GNU AGPL v3.0 or later](LICENSE), the same license as the server. Submit issues
and pull requests to this repository's maintainers. No separate Apache-2.0
license grant to SILO or upstream MinIO maintainers is required.
* **No CLA.** We do not ask you to sign a Contributor License Agreement and we
do not take your copyright. Contributions are accepted inbound=outbound: you
@@ -112,15 +113,20 @@ fork does not try to.
`Signed-off-by` trailers) and add your own sign-off as the person passing it
along. Never import code from a proprietary distribution.
* **File headers.** Files derived from upstream keep the original MinIO
copyright header unchanged. New files added by this fork use the dual
header, followed by the standard AGPL boilerplate:
* **File headers.** Preserve existing copyright and license notices in inherited
and third-party files. New original files name their actual copyright holders
and use AGPL-3.0-or-later. Use a header such as the following, then append the
standard AGPL boilerplate:
```
// Copyright (c) 2015-2025 MinIO, Inc.
// Copyright (c) 2025-2026 PGSTY
// Copyright (c) 2026 Your Name
```
* **Separately licensed material.** Documentation contributions in `docs/`
follow its existing [CC BY 4.0 license](docs/LICENSE). Third-party components
and earlier Apache-2.0 contributions retain their original licenses and
attribution; this policy does not relicense earlier work.
* **Squash merges** must keep the `Signed-off-by:` trailers in the resulting
commit message.
+170 -65
View File
File diff suppressed because one or more lines are too long
+19 -11
View File
@@ -1,4 +1,15 @@
FROM golang:1.27.1-alpine AS build
FROM golang:1.27.1-alpine@sha256:cf6fca6641884b8433441b2b0652976f975e1d0fdd26d177eaaf8596087f3125 AS curl-build
ARG TARGETARCH
COPY dockerscripts/build-static-curl.sh /build/build-static-curl
RUN /bin/sh /build/build-static-curl
# Exercise the exact shipped curl without a dynamic loader or shared libraries.
FROM scratch AS curl-runtime
COPY --from=curl-build /go/bin/curl /curl
COPY --from=curl-build /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/ca-certificates.crt
ENTRYPOINT ["/curl"]
FROM golang:1.27.1-alpine@sha256:cf6fca6641884b8433441b2b0652976f975e1d0fdd26d177eaaf8596087f3125 AS build
ARG TARGETARCH
@@ -6,9 +17,9 @@ ENV GOPATH=/go
ENV CGO_ENABLED=0
ARG MC_REPO=pgsty/mc
ARG MC_VERSION=RELEASE.2026-09-03T07-13-05Z
ARG MC_AMD64_SHA256=6387cbeebb17c4bd52b447ee332ba22777e129034befa6598308c3aa04f02d09
ARG MC_ARM64_SHA256=5fc434c7e416bb4787e92306a8165d55284aac639a29744160b2a198fdd7573a
ARG MC_VERSION=RELEASE.2026-09-13T00-00-00Z
ARG MC_AMD64_SHA256=9d2a92de9c7b887d9b944fe9ddce68d23f1b6df3415092e737594e56593f5e2b
ARG MC_ARM64_SHA256=3d82e9ea6c601c4cb44fe5dd5f2ad1b7d7d64369378110f9ada9c524688a452a
RUN apk add -U --no-cache \
ca-certificates \
@@ -56,18 +67,14 @@ RUN apk add -U --no-cache \
chmod +x /go/bin/mcli && \
ln -sf mcli /go/bin/mc
COPY dockerscripts/download-static-curl.sh /build/download-static-curl
RUN chmod +x /build/download-static-curl && \
/build/download-static-curl
FROM registry.access.redhat.com/ubi9/ubi:latest AS certs
FROM registry.access.redhat.com/ubi9/ubi:latest@sha256:206b65b8ee0f04b992818c9a51b29081b14974630d4850bc358097d0c44ea156 AS certs
RUN dnf -y install ca-certificates && \
update-ca-trust && \
cp /etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem /tmp/ca-certificates.crt && \
dnf clean all && \
rm -rf /var/cache/dnf
FROM registry.access.redhat.com/ubi9/ubi-micro:latest
FROM registry.access.redhat.com/ubi9/ubi-micro:latest@sha256:f332c99eb8f798a8486821c91937f10ad64ee83d7e739303be2df051040918f6
LABEL org.opencontainers.image.title="Silo" \
org.opencontainers.image.description="S3-Interface Libre Object Storage" \
@@ -88,7 +95,8 @@ ENV MINIO_ACCESS_KEY_FILE=access_key \
COPY --from=certs /tmp/ca-certificates.crt /etc/ssl/certs/ca-certificates.crt
COPY silo /usr/bin/silo
COPY --from=build /go/bin/mcli /usr/bin/mcli
COPY --from=build /go/bin/curl* /usr/bin/
COPY --from=curl-build /go/bin/curl /usr/bin/curl
COPY --from=curl-build /go/share/curl /licenses/curl
COPY dockerscripts/docker-entrypoint.sh /usr/bin/docker-entrypoint.sh
COPY LICENSE /licenses/LICENSE
COPY NOTICE /licenses/NOTICE
+1 -1
View File
@@ -216,7 +216,7 @@ docker: checks build-debugging ## builds the local Linux Silo container image
--ldflags "$(LDFLAGS)" -o "$$context/silo"; \
mkdir -p "$$context/dockerscripts"; \
cp Dockerfile.goreleaser LICENSE NOTICE CREDITS "$$context/"; \
cp dockerscripts/docker-entrypoint.sh dockerscripts/download-static-curl.sh \
cp dockerscripts/docker-entrypoint.sh dockerscripts/build-static-curl.sh \
"$$context/dockerscripts/"; \
docker build -q --no-cache --platform linux/$(GOARCH) -t $(TAG) --build-arg TARGETARCH=$(GOARCH) \
-f "$$context/Dockerfile.goreleaser" "$$context"
+65 -45
View File
@@ -89,59 +89,79 @@ The S3 API, `MINIO_*` variables, `minio_*` metrics, `x-minio-*` headers, `/minio
Every divergence from upstream is listed in the code-verified [compatibility audit](https://silo.pgsty.com/compatibility/server/). Treat each release as a downstream upgrade: pin versions, read the [release notes](https://silo.pgsty.com/tags/silo/), and keep a rollback path.
### TLS and Go upgrades
TLS key exchange follows Go's defaults across the S3 listener, node links,
replication, identity providers, etcd, and external HTTP services. If an endpoint
cannot accept ML-KEM, `GODEBUG=tlsmlkem=0` disables the default hybrid exchanges
for the process; certificate verification remains enabled. This option does not
disable ML-DSA signatures or resolve every TLS reset. Prefer updating the
incompatible endpoint before removing the temporary setting.
If only the new SecP hybrids cause problems, `GODEBUG=tlssecpmlkem=0` disables
those groups while retaining X25519MLKEM768.
For builds targeting Go 1.27, setting either `SSL_CERT_FILE` or `SSL_CERT_DIR`
on macOS replaces Keychain trust with on-disk roots and Go's verifier. Stale or
incomplete CA paths can break previously trusted connections; unset inherited
values to restore Keychain trust. Explicit certificates in the configured `CAs`
directory remain additive to the selected root pool.
Go 1.27 binaries require macOS 13 or later. See the
[Go release notes](https://go.dev/doc/go1.27) and the
[SILO stack investigation](docs/investigations/go127-stack.md).
## Security & Contributing
Report vulnerabilities privately as described in [`SECURITY.md`](SECURITY.md); every fix ships with a public [advisory](https://silo.pgsty.com/blog/security/). Contributions are accepted inbound=outbound under AGPL-3.0-or-later with no CLA — only DCO sign-off (`git commit -s`) is required; see [`CONTRIBUTING.md`](CONTRIBUTING.md).
## Contributors
<table>
<tr>
<td align="center" width="150">
<a href="https://github.com/ZouhairCharef"><img src="https://github.com/ZouhairCharef.png?size=100" width="72" alt="ZouhairCharef"><br><sub><b>@ZouhairCharef</b></sub></a><br><sub>CVE-2026-34986</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/mfredenhagen"><img src="https://github.com/mfredenhagen.png?size=100" width="72" alt="mfredenhagen"><br><sub><b>@mfredenhagen</b></sub></a><br><sub>CVE-2026-39883</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/pinginfo"><img src="https://github.com/pinginfo.png?size=100" width="72" alt="pinginfo"><br><sub><b>@pinginfo</b></sub></a><br><sub>Notification streaming</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/waterkip"><img src="https://github.com/waterkip.png?size=100" width="72" alt="waterkip"><br><sub><b>@waterkip</b></sub></a><br><sub>Documentation links</sub>
</td>
</tr>
</table>
<p>
<a href="https://github.com/magicxor"><img src="https://github.com/magicxor.png?size=64" width="44" alt="magicxor" title="@magicxor"></a>
<a href="https://github.com/ycjlin"><img src="https://github.com/ycjlin.png?size=64" width="44" alt="ycjlin" title="@ycjlin"></a>
<a href="https://github.com/h5vx"><img src="https://github.com/h5vx.png?size=64" width="44" alt="h5vx" title="@h5vx"></a>
<a href="https://github.com/Dansyuqri"><img src="https://github.com/Dansyuqri.png?size=64" width="44" alt="Dansyuqri" title="@Dansyuqri"></a>
<a href="https://github.com/davinkevin"><img src="https://github.com/davinkevin.png?size=64" width="44" alt="davinkevin" title="@davinkevin"></a>
<a href="https://github.com/lem21h"><img src="https://github.com/lem21h.png?size=64" width="44" alt="lem21h" title="@lem21h"></a>
<a href="https://github.com/sulin37392"><img src="https://github.com/sulin37392.png?size=64" width="44" alt="sulin37392" title="@sulin37392"></a>
<a href="https://github.com/mosesdd"><img src="https://github.com/mosesdd.png?size=64" width="44" alt="mosesdd" title="@mosesdd"></a>
<a href="https://github.com/Xavier-777"><img src="https://github.com/Xavier-777.png?size=64" width="44" alt="Xavier-777" title="@Xavier-777"></a>
<a href="https://github.com/jiadzh"><img src="https://github.com/jiadzh.png?size=64" width="44" alt="jiadzh" title="@jiadzh"></a>
<a href="https://github.com/TLINDEN"><img src="https://github.com/TLINDEN.png?size=64" width="44" alt="TLINDEN" title="@TLINDEN"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://github.com/AntonOfTheWoods.png?size=64" width="44" alt="AntonOfTheWoods" title="@AntonOfTheWoods"></a>
<a href="https://github.com/zylpsrs"><img src="https://github.com/zylpsrs.png?size=64" width="44" alt="zylpsrs" title="@zylpsrs"></a>
<a href="https://github.com/nsanitate"><img src="https://github.com/nsanitate.png?size=64" width="44" alt="nsanitate" title="@nsanitate"></a>
<a href="https://github.com/makinikm"><img src="https://github.com/makinikm.png?size=64" width="44" alt="makinikm" title="@makinikm"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://github.com/spaceg00se-r.png?size=64" width="44" alt="spaceg00se-r" title="@spaceg00se-r"></a>
<a href="https://github.com/heroes1412"><img src="https://github.com/heroes1412.png?size=64" width="44" alt="heroes1412" title="@heroes1412"></a>
<a href="https://github.com/vampywiz17"><img src="https://github.com/vampywiz17.png?size=64" width="44" alt="vampywiz17" title="@vampywiz17"></a>
<a href="https://github.com/chalukyaj"><img src="https://github.com/chalukyaj.png?size=64" width="44" alt="chalukyaj" title="@chalukyaj"></a>
<a href="https://github.com/cbornet"><img src="https://github.com/cbornet.png?size=64" width="44" alt="cbornet" title="@cbornet"></a>
<a href="https://github.com/jvasile"><img src="https://github.com/jvasile.png?size=64" width="44" alt="jvasile" title="@jvasile"></a>
<a href="https://github.com/Kesavaambati"><img src="https://github.com/Kesavaambati.png?size=64" width="44" alt="Kesavaambati" title="@Kesavaambati"></a>
<a href="https://github.com/redfoxfox"><img src="https://github.com/redfoxfox.png?size=64" width="44" alt="redfoxfox" title="@redfoxfox"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://github.com/kuldeep-link11.png?size=64" width="44" alt="kuldeep-link11" title="@kuldeep-link11"></a>
<a href="https://github.com/meesudzu"><img src="https://github.com/meesudzu.png?size=64" width="44" alt="meesudzu" title="@meesudzu"></a>
<a href="https://github.com/pmezhuev"><img src="https://github.com/pmezhuev.png?size=64" width="44" alt="pmezhuev" title="@pmezhuev"></a>
<a href="https://github.com/kh0mka"><img src="https://github.com/kh0mka.png?size=64" width="44" alt="kh0mka" title="@kh0mka"></a>
**41 community contributors** build SILO, Console, mcli, shared packages, and related projects. The list includes maintainers and every human Issue or PR author, ordered by merged PRs, other PRs, then issue reports. Gold rings highlight significant contributions.
<p align="center">
<a href="https://github.com/Vonng"><img src="https://silo.pgsty.com/images/contributors/Vonng.svg" width="60" height="60" alt="@Vonng" title="@Vonng — Maintains SILO, Console, mcli, shared packages, releases, and documentation"></a>
<a href="https://github.com/h5vx"><img src="https://silo.pgsty.com/images/contributors/h5vx.svg" width="60" height="60" alt="@h5vx" title="@h5vx — Implemented per-bucket CORS configuration and enforcement"></a>
<a href="https://github.com/mrjavadseydi"><img src="https://silo.pgsty.com/images/contributors/mrjavadseydi.svg" width="60" height="60" alt="@mrjavadseydi" title="@mrjavadseydi — Fixed effective bucket quota metrics; proposed access-frequency ILM"></a>
<a href="https://github.com/Dansyuqri"><img src="https://silo.pgsty.com/images/contributors/Dansyuqri.svg" width="60" height="60" alt="@Dansyuqri" title="@Dansyuqri — Added ChecksumType to multipart completion responses"></a>
<a href="https://github.com/ycjlin"><img src="https://silo.pgsty.com/images/contributors/ycjlin.svg" width="60" height="60" alt="@ycjlin" title="@ycjlin — Fixed missing-bucket ListObjects semantics"></a>
<a href="https://github.com/pinginfo"><img src="https://silo.pgsty.com/images/contributors/pinginfo.svg" width="60" height="60" alt="@pinginfo" title="@pinginfo — Repaired bucket notification streaming"></a>
<a href="https://github.com/ZouhairCharef"><img src="https://silo.pgsty.com/images/contributors/ZouhairCharef.svg" width="60" height="60" alt="@ZouhairCharef" title="@ZouhairCharef — Patched CVE-2026-34986 in go-jose"></a>
<a href="https://github.com/mfredenhagen"><img src="https://silo.pgsty.com/images/contributors/mfredenhagen.svg" width="60" height="60" alt="@mfredenhagen" title="@mfredenhagen — Patched CVE-2026-39883 in OpenTelemetry"></a>
<a href="https://github.com/waterkip"><img src="https://silo.pgsty.com/images/contributors/waterkip.svg" width="60" height="60" alt="@waterkip" title="@waterkip — Repointed documentation links to the SILO portal"></a>
<a href="https://github.com/mikemikimike"><img src="https://silo.pgsty.com/images/contributors/mikemikimike.svg" width="60" height="60" alt="@mikemikimike" title="@mikemikimike — Contributed the replicated SSE-C plaintext part-size fix"></a>
<a href="https://github.com/metaneutrons"><img src="https://silo.pgsty.com/images/contributors/metaneutrons.svg" width="60" height="60" alt="@metaneutrons" title="@metaneutrons — Reported and proposed explicit-version delete authorization"></a>
<a href="https://github.com/magicxor"><img src="https://silo.pgsty.com/images/contributors/magicxor.svg" width="60" height="60" alt="@magicxor" title="@magicxor — Reported and proposed conditional DELETE support for If-Match"></a>
<a href="https://github.com/davinkevin"><img src="https://silo.pgsty.com/images/contributors/davinkevin.svg" width="60" height="60" alt="@davinkevin" title="@davinkevin — Proposed the distroless container image and dependency automation"></a>
<a href="https://github.com/lem21h"><img src="https://silo.pgsty.com/images/contributors/lem21h.svg" width="48" height="48" alt="@lem21h" title="@lem21h — Proposed robustness and goroutine improvements"></a>
<a href="https://github.com/sulin37392"><img src="https://silo.pgsty.com/images/contributors/sulin37392.svg" width="48" height="48" alt="@sulin37392" title="@sulin37392 — Proposed dependency updates"></a>
<a href="https://github.com/cbornet"><img src="https://silo.pgsty.com/images/contributors/cbornet.svg" width="60" height="60" alt="@cbornet" title="@cbornet — Reported multipart and streaming checksum defects and missing-bucket semantics"></a>
<a href="https://github.com/vampywiz17"><img src="https://silo.pgsty.com/images/contributors/vampywiz17.svg" width="60" height="60" alt="@vampywiz17" title="@vampywiz17 — Reported LDAP TLS and Console login regressions"></a>
<a href="https://github.com/orenyomtov"><img src="https://silo.pgsty.com/images/contributors/orenyomtov.svg" width="60" height="60" alt="@orenyomtov" title="@orenyomtov — Reported the unsigned-header CopyObject cross-object read (SN-2026-011)"></a>
<a href="https://github.com/mumu-lab"><img src="https://silo.pgsty.com/images/contributors/mumu-lab.svg" width="48" height="48" alt="@mumu-lab" title="@mumu-lab — Reported bucket quota metrics reading a deprecated field"></a>
<a href="https://github.com/jvasile"><img src="https://silo.pgsty.com/images/contributors/jvasile.svg" width="48" height="48" alt="@jvasile" title="@jvasile — Reported missing user, group, and defaults in Debian packages"></a>
<a href="https://github.com/pmezhuev"><img src="https://silo.pgsty.com/images/contributors/pmezhuev.svg" width="48" height="48" alt="@pmezhuev" title="@pmezhuev — Reported missing RPM package signatures"></a>
<a href="https://github.com/TLINDEN"><img src="https://silo.pgsty.com/images/contributors/TLINDEN.svg" width="48" height="48" alt="@TLINDEN" title="@TLINDEN — Reported the missing client in release tarballs"></a>
<a href="https://github.com/makinikm"><img src="https://silo.pgsty.com/images/contributors/makinikm.svg" width="48" height="48" alt="@makinikm" title="@makinikm — Reported the missing client in the container image"></a>
<a href="https://github.com/meesudzu"><img src="https://silo.pgsty.com/images/contributors/meesudzu.svg" width="48" height="48" alt="@meesudzu" title="@meesudzu — Requested the migration guide from upstream MinIO"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://silo.pgsty.com/images/contributors/kuldeep-link11.svg" width="48" height="48" alt="@kuldeep-link11" title="@kuldeep-link11 — Reported NATS JWT credentials and target reload issues"></a>
<a href="https://github.com/sargarass"><img src="https://silo.pgsty.com/images/contributors/sargarass.svg" width="48" height="48" alt="@sargarass" title="@sargarass — Reported ListMultipartUploads prefix and pagination semantics"></a>
<a href="https://github.com/liuhaodongliu990-cmyk"><img src="https://silo.pgsty.com/images/contributors/liuhaodongliu990-cmyk.svg" width="48" height="48" alt="@liuhaodongliu990-cmyk" title="@liuhaodongliu990-cmyk — Reported indeterminate progress for prefix downloads"></a>
<a href="https://github.com/Xavier-777"><img src="https://silo.pgsty.com/images/contributors/Xavier-777.svg" width="48" height="48" alt="@Xavier-777" title="@Xavier-777 — Reported Console lifecycle management and file preview gaps"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://silo.pgsty.com/images/contributors/spaceg00se-r.svg" width="48" height="48" alt="@spaceg00se-r" title="@spaceg00se-r — Requested cpuv1 support and reported a workflow token failure"></a>
<a href="https://github.com/kh0mka"><img src="https://silo.pgsty.com/images/contributors/kh0mka.svg" width="48" height="48" alt="@kh0mka" title="@kh0mka — Reported inter-node I/O timeouts in ReadFileStreamHandler"></a>
<a href="https://github.com/bagutzu"><img src="https://silo.pgsty.com/images/contributors/bagutzu.svg" width="48" height="48" alt="@bagutzu" title="@bagutzu — Requested KES-compatible external KMS and OpenBao support"></a>
<a href="https://github.com/DestroyLee"><img src="https://silo.pgsty.com/images/contributors/DestroyLee.svg" width="48" height="48" alt="@DestroyLee" title="@DestroyLee — Reported the missing documentation navigation"></a>
<a href="https://github.com/mosesdd"><img src="https://silo.pgsty.com/images/contributors/mosesdd.svg" width="48" height="48" alt="@mosesdd" title="@mosesdd — Requested a maintained Helm chart"></a>
<a href="https://github.com/zylpsrs"><img src="https://silo.pgsty.com/images/contributors/zylpsrs.svg" width="48" height="48" alt="@zylpsrs" title="@zylpsrs — Reported missing Console tiering and site replication"></a>
<a href="https://github.com/heroes1412"><img src="https://silo.pgsty.com/images/contributors/heroes1412.svg" width="48" height="48" alt="@heroes1412" title="@heroes1412 — Reported the unusable profiling option"></a>
<a href="https://github.com/redfoxfox"><img src="https://silo.pgsty.com/images/contributors/redfoxfox.svg" width="48" height="48" alt="@redfoxfox" title="@redfoxfox — Reported Chinese documentation availability"></a>
<a href="https://github.com/jiadzh"><img src="https://silo.pgsty.com/images/contributors/jiadzh.svg" width="48" height="48" alt="@jiadzh" title="@jiadzh — Requested Windows build guidance"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://silo.pgsty.com/images/contributors/AntonOfTheWoods.svg" width="48" height="48" alt="@AntonOfTheWoods" title="@AntonOfTheWoods — Asked for clarity on Helm chart and operator options"></a>
<a href="https://github.com/chalukyaj"><img src="https://silo.pgsty.com/images/contributors/chalukyaj.svg" width="48" height="48" alt="@chalukyaj" title="@chalukyaj — Proposed making the SILO Operator easier to discover"></a>
<a href="https://github.com/nsanitate"><img src="https://silo.pgsty.com/images/contributors/nsanitate.svg" width="48" height="48" alt="@nsanitate" title="@nsanitate — Proposed CNCF Sandbox governance"></a>
<a href="https://github.com/Kesavaambati"><img src="https://silo.pgsty.com/images/contributors/Kesavaambati.svg" width="48" height="48" alt="@Kesavaambati" title="@Kesavaambati — Asked about community support and image maintenance"></a>
</p>
GitHub does not generate a contributor graph for forks, so [`CONTRIBUTORS.md`](CONTRIBUTORS.md) — not the Insights page — is this project's attribution record. It names everyone alongside the change or report they contributed.
[View the full contribution record](CONTRIBUTORS.md) for each person's proposals, fixes, and reports.
## Background
+45 -45
View File
@@ -95,53 +95,53 @@ S3 API、`MINIO_*` 环境变量、`minio_*` 指标、`x-minio-*` 头、`/minio/*
## 贡献者
<table>
<tr>
<td align="center" width="150">
<a href="https://github.com/ZouhairCharef"><img src="https://github.com/ZouhairCharef.png?size=100" width="72" alt="ZouhairCharef"><br><sub><b>@ZouhairCharef</b></sub></a><br><sub>CVE-2026-34986</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/mfredenhagen"><img src="https://github.com/mfredenhagen.png?size=100" width="72" alt="mfredenhagen"><br><sub><b>@mfredenhagen</b></sub></a><br><sub>CVE-2026-39883</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/pinginfo"><img src="https://github.com/pinginfo.png?size=100" width="72" alt="pinginfo"><br><sub><b>@pinginfo</b></sub></a><br><sub>桶通知流式输出</sub>
</td>
<td align="center" width="150">
<a href="https://github.com/waterkip"><img src="https://github.com/waterkip.png?size=100" width="72" alt="waterkip"><br><sub><b>@waterkip</b></sub></a><br><sub>文档链接修正</sub>
</td>
</tr>
</table>
<p>
<a href="https://github.com/magicxor"><img src="https://github.com/magicxor.png?size=64" width="44" alt="magicxor" title="@magicxor"></a>
<a href="https://github.com/ycjlin"><img src="https://github.com/ycjlin.png?size=64" width="44" alt="ycjlin" title="@ycjlin"></a>
<a href="https://github.com/h5vx"><img src="https://github.com/h5vx.png?size=64" width="44" alt="h5vx" title="@h5vx"></a>
<a href="https://github.com/Dansyuqri"><img src="https://github.com/Dansyuqri.png?size=64" width="44" alt="Dansyuqri" title="@Dansyuqri"></a>
<a href="https://github.com/davinkevin"><img src="https://github.com/davinkevin.png?size=64" width="44" alt="davinkevin" title="@davinkevin"></a>
<a href="https://github.com/lem21h"><img src="https://github.com/lem21h.png?size=64" width="44" alt="lem21h" title="@lem21h"></a>
<a href="https://github.com/sulin37392"><img src="https://github.com/sulin37392.png?size=64" width="44" alt="sulin37392" title="@sulin37392"></a>
<a href="https://github.com/mosesdd"><img src="https://github.com/mosesdd.png?size=64" width="44" alt="mosesdd" title="@mosesdd"></a>
<a href="https://github.com/Xavier-777"><img src="https://github.com/Xavier-777.png?size=64" width="44" alt="Xavier-777" title="@Xavier-777"></a>
<a href="https://github.com/jiadzh"><img src="https://github.com/jiadzh.png?size=64" width="44" alt="jiadzh" title="@jiadzh"></a>
<a href="https://github.com/TLINDEN"><img src="https://github.com/TLINDEN.png?size=64" width="44" alt="TLINDEN" title="@TLINDEN"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://github.com/AntonOfTheWoods.png?size=64" width="44" alt="AntonOfTheWoods" title="@AntonOfTheWoods"></a>
<a href="https://github.com/zylpsrs"><img src="https://github.com/zylpsrs.png?size=64" width="44" alt="zylpsrs" title="@zylpsrs"></a>
<a href="https://github.com/nsanitate"><img src="https://github.com/nsanitate.png?size=64" width="44" alt="nsanitate" title="@nsanitate"></a>
<a href="https://github.com/makinikm"><img src="https://github.com/makinikm.png?size=64" width="44" alt="makinikm" title="@makinikm"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://github.com/spaceg00se-r.png?size=64" width="44" alt="spaceg00se-r" title="@spaceg00se-r"></a>
<a href="https://github.com/heroes1412"><img src="https://github.com/heroes1412.png?size=64" width="44" alt="heroes1412" title="@heroes1412"></a>
<a href="https://github.com/vampywiz17"><img src="https://github.com/vampywiz17.png?size=64" width="44" alt="vampywiz17" title="@vampywiz17"></a>
<a href="https://github.com/chalukyaj"><img src="https://github.com/chalukyaj.png?size=64" width="44" alt="chalukyaj" title="@chalukyaj"></a>
<a href="https://github.com/cbornet"><img src="https://github.com/cbornet.png?size=64" width="44" alt="cbornet" title="@cbornet"></a>
<a href="https://github.com/jvasile"><img src="https://github.com/jvasile.png?size=64" width="44" alt="jvasile" title="@jvasile"></a>
<a href="https://github.com/Kesavaambati"><img src="https://github.com/Kesavaambati.png?size=64" width="44" alt="Kesavaambati" title="@Kesavaambati"></a>
<a href="https://github.com/redfoxfox"><img src="https://github.com/redfoxfox.png?size=64" width="44" alt="redfoxfox" title="@redfoxfox"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://github.com/kuldeep-link11.png?size=64" width="44" alt="kuldeep-link11" title="@kuldeep-link11"></a>
<a href="https://github.com/meesudzu"><img src="https://github.com/meesudzu.png?size=64" width="44" alt="meesudzu" title="@meesudzu"></a>
<a href="https://github.com/pmezhuev"><img src="https://github.com/pmezhuev.png?size=64" width="44" alt="pmezhuev" title="@pmezhuev"></a>
<a href="https://github.com/kh0mka"><img src="https://github.com/kh0mka.png?size=64" width="44" alt="kh0mka" title="@kh0mka"></a>
**41 位社区贡献者**共同建设 SILO、Console、mcli、公共包与相关项目。名单包含维护者,以及所有提出 Issue 或 PR 的真人作者;按已合并 PR、其他 PR、Issue 报告排序,黄圈标记显著贡献。
<p align="center">
<a href="https://github.com/Vonng"><img src="https://silo.pgsty.com/images/contributors/Vonng.svg" width="60" height="60" alt="@Vonng" title="@Vonng — 维护 SILO、Console、mcli、公共包、发行与文档"></a>
<a href="https://github.com/h5vx"><img src="https://silo.pgsty.com/images/contributors/h5vx.svg" width="60" height="60" alt="@h5vx" title="@h5vx — 实现单桶 CORS 配置与请求执行"></a>
<a href="https://github.com/mrjavadseydi"><img src="https://silo.pgsty.com/images/contributors/mrjavadseydi.svg" width="60" height="60" alt="@mrjavadseydi" title="@mrjavadseydi — 修复有效桶配额指标,并提交按访问频率分层的 ILM 方案"></a>
<a href="https://github.com/Dansyuqri"><img src="https://silo.pgsty.com/images/contributors/Dansyuqri.svg" width="60" height="60" alt="@Dansyuqri" title="@Dansyuqri — 为分片上传完成响应补充 ChecksumType"></a>
<a href="https://github.com/ycjlin"><img src="https://silo.pgsty.com/images/contributors/ycjlin.svg" width="60" height="60" alt="@ycjlin" title="@ycjlin — 修复缺失桶的 ListObjects 语义"></a>
<a href="https://github.com/pinginfo"><img src="https://silo.pgsty.com/images/contributors/pinginfo.svg" width="60" height="60" alt="@pinginfo" title="@pinginfo — 修复桶通知的流式输出"></a>
<a href="https://github.com/ZouhairCharef"><img src="https://silo.pgsty.com/images/contributors/ZouhairCharef.svg" width="60" height="60" alt="@ZouhairCharef" title="@ZouhairCharef — 修复 go-jose 中的 CVE-2026-34986"></a>
<a href="https://github.com/mfredenhagen"><img src="https://silo.pgsty.com/images/contributors/mfredenhagen.svg" width="60" height="60" alt="@mfredenhagen" title="@mfredenhagen — 修复 OpenTelemetry 中的 CVE-2026-39883"></a>
<a href="https://github.com/waterkip"><img src="https://silo.pgsty.com/images/contributors/waterkip.svg" width="60" height="60" alt="@waterkip" title="@waterkip — 将文档链接指向 SILO 门户"></a>
<a href="https://github.com/mikemikimike"><img src="https://silo.pgsty.com/images/contributors/mikemikimike.svg" width="60" height="60" alt="@mikemikimike" title="@mikemikimike — 提交 SSE-C 复制分片明文尺寸修复"></a>
<a href="https://github.com/metaneutrons"><img src="https://silo.pgsty.com/images/contributors/metaneutrons.svg" width="60" height="60" alt="@metaneutrons" title="@metaneutrons — 报告并提交显式版本删除鉴权方案"></a>
<a href="https://github.com/magicxor"><img src="https://silo.pgsty.com/images/contributors/magicxor.svg" width="60" height="60" alt="@magicxor" title="@magicxor — 报告并提交 DELETE If-Match 条件请求支持方案"></a>
<a href="https://github.com/davinkevin"><img src="https://silo.pgsty.com/images/contributors/davinkevin.svg" width="60" height="60" alt="@davinkevin" title="@davinkevin — 提交 distroless 容器镜像与依赖自动更新方案"></a>
<a href="https://github.com/lem21h"><img src="https://silo.pgsty.com/images/contributors/lem21h.svg" width="48" height="48" alt="@lem21h" title="@lem21h — 提交健壮性与 goroutine 改进"></a>
<a href="https://github.com/sulin37392"><img src="https://silo.pgsty.com/images/contributors/sulin37392.svg" width="48" height="48" alt="@sulin37392" title="@sulin37392 — 提交依赖更新"></a>
<a href="https://github.com/cbornet"><img src="https://silo.pgsty.com/images/contributors/cbornet.svg" width="60" height="60" alt="@cbornet" title="@cbornet — 报告分片与流式校验和缺陷及缺失桶语义问题"></a>
<a href="https://github.com/vampywiz17"><img src="https://silo.pgsty.com/images/contributors/vampywiz17.svg" width="60" height="60" alt="@vampywiz17" title="@vampywiz17 — 报告 LDAP TLS 与 Console 登录回归"></a>
<a href="https://github.com/orenyomtov"><img src="https://silo.pgsty.com/images/contributors/orenyomtov.svg" width="60" height="60" alt="@orenyomtov" title="@orenyomtov — 报告未签名头导致的 CopyObject 跨对象读取(SN-2026-011)"></a>
<a href="https://github.com/mumu-lab"><img src="https://silo.pgsty.com/images/contributors/mumu-lab.svg" width="48" height="48" alt="@mumu-lab" title="@mumu-lab — 报告桶配额指标读取已弃用字段的问题"></a>
<a href="https://github.com/jvasile"><img src="https://silo.pgsty.com/images/contributors/jvasile.svg" width="48" height="48" alt="@jvasile" title="@jvasile — 报告 Debian 包缺少用户、用户组与默认配置"></a>
<a href="https://github.com/pmezhuev"><img src="https://silo.pgsty.com/images/contributors/pmezhuev.svg" width="48" height="48" alt="@pmezhuev" title="@pmezhuev — 报告 RPM 包缺少 GPG 签名"></a>
<a href="https://github.com/TLINDEN"><img src="https://silo.pgsty.com/images/contributors/TLINDEN.svg" width="48" height="48" alt="@TLINDEN" title="@TLINDEN — 报告发布压缩包缺少客户端"></a>
<a href="https://github.com/makinikm"><img src="https://silo.pgsty.com/images/contributors/makinikm.svg" width="48" height="48" alt="@makinikm" title="@makinikm — 报告容器镜像缺少客户端"></a>
<a href="https://github.com/meesudzu"><img src="https://silo.pgsty.com/images/contributors/meesudzu.svg" width="48" height="48" alt="@meesudzu" title="@meesudzu — 提出从上游 MinIO 迁移的指南需求"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://silo.pgsty.com/images/contributors/kuldeep-link11.svg" width="48" height="48" alt="@kuldeep-link11" title="@kuldeep-link11 — 报告 NATS JWT 凭据与通知目标重载问题"></a>
<a href="https://github.com/sargarass"><img src="https://silo.pgsty.com/images/contributors/sargarass.svg" width="48" height="48" alt="@sargarass" title="@sargarass — 报告 ListMultipartUploads 前缀与分页语义问题"></a>
<a href="https://github.com/liuhaodongliu990-cmyk"><img src="https://silo.pgsty.com/images/contributors/liuhaodongliu990-cmyk.svg" width="48" height="48" alt="@liuhaodongliu990-cmyk" title="@liuhaodongliu990-cmyk — 报告前缀下载进度显示异常"></a>
<a href="https://github.com/Xavier-777"><img src="https://silo.pgsty.com/images/contributors/Xavier-777.svg" width="48" height="48" alt="@Xavier-777" title="@Xavier-777 — 报告 Console 生命周期管理与文件预览缺失"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://silo.pgsty.com/images/contributors/spaceg00se-r.svg" width="48" height="48" alt="@spaceg00se-r" title="@spaceg00se-r — 提出 cpuv1 支持需求并报告工作流令牌错误"></a>
<a href="https://github.com/kh0mka"><img src="https://silo.pgsty.com/images/contributors/kh0mka.svg" width="48" height="48" alt="@kh0mka" title="@kh0mka — 报告 ReadFileStreamHandler 节点间 I/O 超时"></a>
<a href="https://github.com/bagutzu"><img src="https://silo.pgsty.com/images/contributors/bagutzu.svg" width="48" height="48" alt="@bagutzu" title="@bagutzu — 提出兼容 KES 的外部 KMS 与 OpenBao 支持需求"></a>
<a href="https://github.com/DestroyLee"><img src="https://silo.pgsty.com/images/contributors/DestroyLee.svg" width="48" height="48" alt="@DestroyLee" title="@DestroyLee — 报告文档目录导航缺失"></a>
<a href="https://github.com/mosesdd"><img src="https://silo.pgsty.com/images/contributors/mosesdd.svg" width="48" height="48" alt="@mosesdd" title="@mosesdd — 提出维护 Helm Chart 的需求"></a>
<a href="https://github.com/zylpsrs"><img src="https://silo.pgsty.com/images/contributors/zylpsrs.svg" width="48" height="48" alt="@zylpsrs" title="@zylpsrs — 报告 Console 缺少分层与站点复制"></a>
<a href="https://github.com/heroes1412"><img src="https://silo.pgsty.com/images/contributors/heroes1412.svg" width="48" height="48" alt="@heroes1412" title="@heroes1412 — 报告性能分析选项不可用"></a>
<a href="https://github.com/redfoxfox"><img src="https://silo.pgsty.com/images/contributors/redfoxfox.svg" width="48" height="48" alt="@redfoxfox" title="@redfoxfox — 报告中文文档站点不可用"></a>
<a href="https://github.com/jiadzh"><img src="https://silo.pgsty.com/images/contributors/jiadzh.svg" width="48" height="48" alt="@jiadzh" title="@jiadzh — 提出 Windows 构建指导需求"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://silo.pgsty.com/images/contributors/AntonOfTheWoods.svg" width="48" height="48" alt="@AntonOfTheWoods" title="@AntonOfTheWoods — 提出明确 Helm Chart 与 Operator 选项的需求"></a>
<a href="https://github.com/chalukyaj"><img src="https://silo.pgsty.com/images/contributors/chalukyaj.svg" width="48" height="48" alt="@chalukyaj" title="@chalukyaj — 提出改善 SILO Operator 可发现性的建议"></a>
<a href="https://github.com/nsanitate"><img src="https://silo.pgsty.com/images/contributors/nsanitate.svg" width="48" height="48" alt="@nsanitate" title="@nsanitate — 提出加入 CNCF Sandbox 的治理建议"></a>
<a href="https://github.com/Kesavaambati"><img src="https://silo.pgsty.com/images/contributors/Kesavaambati.svg" width="48" height="48" alt="@Kesavaambati" title="@Kesavaambati — 提出社区支持与容器镜像维护问题"></a>
</p>
GitHub 不为 fork 仓库生成贡献者图表,因此 [`CONTRIBUTORS.md`](CONTRIBUTORS.md)(而非 Insights 页面)才是本项目的署名记录,其中逐一记录了每个人对应的改动或报告。
[查看完整贡献记录](CONTRIBUTORS.md),了解每位贡献者的提案、修复与问题报告。
## 背景
+1 -1
View File
@@ -40,7 +40,7 @@ if [ -n "${MCLI_BIN:-}" ]; then
exit 0
fi
release=${MCLI_RELEASE:-RELEASE.2026-09-03T07-13-05Z}
release=${MCLI_RELEASE:-RELEASE.2026-09-13T00-00-00Z}
version_hyphen=${release#RELEASE.}
package_version=$(printf '%s\n' "${version_hyphen}" | sed -E 's/^([0-9]{4})-([0-9]{2})-([0-9]{2})T([0-9]{2})-([0-9]{2})-([0-9]{2})Z$/\1\2\3\4\5\6.0.0/')
if [ "${package_version}" = "${version_hyphen}" ]; then
@@ -289,6 +289,16 @@
"MINIO_IDENTITY_TLS_SKIP_VERIFY",
"MINIO_IDENTITY_TLS_STS_EXPIRY",
"MINIO_IDLE_TIMEOUT",
"MINIO_ILM_ACCESS_BINS",
"MINIO_ILM_ACCESS_BIN_WIDTH",
"MINIO_ILM_ACCESS_FLUSH",
"MINIO_ILM_ACCESS_MAX_SIZE",
"MINIO_ILM_ACCESS_MAX_TRACKED",
"MINIO_ILM_ACCESS_MIN_RESIDENCY",
"MINIO_ILM_ACCESS_POOLS",
"MINIO_ILM_ACCESS_PROMOTE_WATERMARK",
"MINIO_ILM_ACCESS_TIERING",
"MINIO_ILM_ACCESS_WORKERS",
"MINIO_ILM_EXPIRATION_WORKERS",
"MINIO_ILM_TRANSITION_WORKERS",
"MINIO_INTERFACE",
@@ -502,6 +512,7 @@
"MINIO_SITE_COMMENT",
"MINIO_SITE_NAME",
"MINIO_SITE_REGION",
"MINIO_SITE_REPLICATION_METADATA_TOMBSTONES",
"MINIO_STORAGE_CLASS_COMMENT",
"MINIO_STORAGE_CLASS_INLINE_BLOCK",
"MINIO_STORAGE_CLASS_OPTIMIZE",
@@ -594,6 +605,7 @@
"x-minio-internal-encrypted-multipart",
"x-minio-internal-encryptedmultipart",
"x-minio-internal-erasure-upgraded",
"x-minio-internal-ilm-atier",
"x-minio-internal-inline-data",
"x-minio-internal-objectlock-legalhold-timestamp",
"x-minio-internal-replica-status",
@@ -616,6 +628,7 @@
"x-minio-internal-transitioned-object",
"x-minio-internal-xyz",
"x-minio-key",
"x-minio-last-modified",
"x-minio-lifecycleconfig-updatedat",
"x-minio-meta",
"x-minio-meta-appid",
@@ -631,6 +644,7 @@
"x-minio-replication-encrypted-multipart",
"x-minio-replication-ready",
"x-minio-replication-reset-status",
"x-minio-replication-server-side-encryption",
"x-minio-replication-server-side-encryption-iv",
"x-minio-replication-server-side-encryption-seal-algorithm",
"x-minio-replication-server-side-encryption-sealed-key",
@@ -731,6 +745,7 @@
"/idp/ldap/policy/{operation}",
"/idp/openid/list-access-keys-bulk",
"/ilm",
"/ilm/access",
"/import-bucket-metadata",
"/import-iam",
"/import-iam-v2",
@@ -1058,12 +1073,10 @@
"cmd/metrics.go=\"Version of current MinIO server instance\"",
"cmd/metrics.go=\"minio\"",
"cmd/object-api-utils.go=\".minio.sys\"",
"cmd/object-handlers-common.go=\"X-Minio-Last-Modified\"",
"cmd/object-handlers.go=\"minio-federated\"",
"cmd/object-handlers.go=\"minio.metadata.\"",
"cmd/object-handlers.go=\"minio.versionId\"",
"cmd/object-multipart-handlers.go=\"X-Minio-Replication-Server-Side-Encryption-Iv\"",
"cmd/object-multipart-handlers.go=\"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm\"",
"cmd/object-multipart-handlers.go=\"X-Minio-Replication-Server-Side-Encryption-Sealed-Key\"",
"cmd/replication-trust.go=\"X-Minio-Replication-Encrypted-Multipart\"",
"cmd/replication-trust.go=\"X-Minio-Replication-Server-Side-Encryption-Iv\"",
"cmd/replication-trust.go=\"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm\"",
+2
View File
@@ -115,7 +115,9 @@ func collect(repo string) (manifest, error) {
fset := token.NewFileSet()
for _, rel := range files {
// Investigation artifacts contain synthetic routes and archived configurations.
if rel == "SILO_REBRANDING_MIGRATION.md" ||
strings.HasPrefix(rel, "docs/investigations/") ||
strings.HasPrefix(rel, "buildscripts/rebrand-guard/") ||
strings.HasPrefix(rel, "buildscripts/helm-migration-guard/") {
continue
+11 -9
View File
@@ -36,7 +36,7 @@ for file in \
buildscripts/verify-helm-migration.sh \
Dockerfile.goreleaser \
Dockerfile.distroless \
dockerscripts/download-static-curl.sh \
dockerscripts/build-static-curl.sh \
dockerscripts/docker-entrypoint.sh \
helm/silo/Chart.yaml \
helm/silo/values.yaml \
@@ -101,7 +101,7 @@ require_text Dockerfile.goreleaser "Published checksum drift"
require_text Dockerfile.distroless 'COPY --chmod=0755 silo /usr/bin/silo'
require_text Dockerfile.distroless 'ENTRYPOINT ["/usr/bin/silo"]'
require_text Dockerfile.distroless '"/usr/bin/silo", "healthcheck", "ready"'
require_text dockerscripts/download-static-curl.sh "sha256sum -c"
require_text dockerscripts/build-static-curl.sh "sha256sum -c"
require_text helm/silo/Chart.yaml "name: silo"
require_text helm/silo/values.yaml "repository: pgsty/silo"
require_text helm/silo/templates/deployment.yaml "/usr/bin/docker-entrypoint.sh silo server"
@@ -159,11 +159,13 @@ fi
# The repository and its default branch are pgsty/silo and main. The invariant
# is that the old name is never a live target, not that it is never spoken: the
# READMEs have to name it to explain the rename and to point at the archived
# artifacts, which is the opposite of stranding a reader on it.
# artifacts, which is the opposite of stranding a reader on it. CONTRIBUTORS.md
# also quotes historical issue titles.
#
# So two rules. First, no live URL may resolve to the old repository anywhere,
# READMEs included.
stale_repo_url="$(rg -n -e 'github\.com/pgsty/minio' -e 'hub\.docker\.com/r/pgsty/minio' \
# READMEs and CONTRIBUTORS.md included.
old_repo_pattern='pgsty/minio(\.git)?([^[:alnum:]_.-]|$)'
stale_repo_url="$(rg -n -e "github\.com/${old_repo_pattern}" -e "hub\.docker\.com/r/${old_repo_pattern}" \
--glob '!.git/**' --glob '!dist/**' \
--glob '!SILO_REBRANDING_MIGRATION.md' \
--glob '!buildscripts/rebrand-guard/compat-baseline.json' . |
@@ -175,10 +177,10 @@ fi
# Second, the bare name may only appear where it is deliberate: the pinned
# pre-rebrand image digest in the upgrade test, the two guards that refuse a
# legacy image, and the two READMEs that document the rename and the archived
# minio branch.
repo_guard_allowlist='^(buildscripts/minio-upgrade\.sh|buildscripts/verify-rebrand\.sh|buildscripts/helm-migration-guard/main\.go|README\.md|README_ZH\.md):'
stale_repo="$(rg -n 'pgsty/minio' --glob '!.git/**' --glob '!dist/**' \
# legacy image, the two READMEs that document the rename and the archived
# minio branch, and historical issue titles in CONTRIBUTORS.md.
repo_guard_allowlist='^(buildscripts/minio-upgrade\.sh|buildscripts/verify-rebrand\.sh|buildscripts/helm-migration-guard/main\.go|README\.md|README_ZH\.md|CONTRIBUTORS\.md):'
stale_repo="$(rg -n "${old_repo_pattern}" --glob '!.git/**' --glob '!dist/**' \
--glob '!SILO_REBRANDING_MIGRATION.md' \
--glob '!buildscripts/rebrand-guard/compat-baseline.json' . |
sed 's#^\./##' | grep -Ev "${repo_guard_allowlist}" || true)"
+348
View File
@@ -0,0 +1,348 @@
package cmd
import (
"archive/zip"
"bytes"
"encoding/base64"
"encoding/json"
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
"github.com/minio/mux"
)
func corsAdminRequest(t *testing.T, cred auth.Credentials, method, path string, body []byte) *httptest.ResponseRecorder {
t.Helper()
router := mux.NewRouter()
registerAdminRouter(router, true)
req, err := newTestSignedRequestV4(method, adminPathPrefix+adminAPIVersionPrefix+path,
int64(len(body)), bytes.NewReader(body), cred.AccessKey, cred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("admin %s: %d: %s", path, rec.Code, rec.Body.String())
}
return rec
}
func corsImportReport(t *testing.T, rec *httptest.ResponseRecorder) madmin.BucketMetaImportErrs {
t.Helper()
var rpt madmin.BucketMetaImportErrs
if err := json.Unmarshal(rec.Body.Bytes(), &rpt); err != nil {
t.Fatalf("import report %q: %v", rec.Body.String(), err)
}
return rpt
}
func corsZip(t *testing.T, entries map[string][]byte) []byte {
t.Helper()
var buf bytes.Buffer
zw := zip.NewWriter(&buf)
for name, data := range entries {
w, err := zw.Create(name)
if err != nil {
t.Fatal(err)
}
if _, err = w.Write(data); err != nil {
t.Fatal(err)
}
}
if err := zw.Close(); err != nil {
t.Fatal(err)
}
return buf.Bytes()
}
// corsCorruptedZip builds an archive holding a stored (uncompressed) cors.xml
// whose payload is altered after the checksum is computed, plus the given
// companion entries. The altered document stays well formed, so only the zip
// checksum tells the two apart.
func corsCorruptedZip(t *testing.T, name string, doc []byte, others map[string][]byte) []byte {
t.Helper()
var buf bytes.Buffer
zw := zip.NewWriter(&buf)
w, err := zw.CreateHeader(&zip.FileHeader{Name: name, Method: zip.Store})
if err != nil {
t.Fatal(err)
}
if _, err = w.Write(doc); err != nil {
t.Fatal(err)
}
for other, data := range others {
ow, err := zw.Create(other)
if err != nil {
t.Fatal(err)
}
if _, err = ow.Write(data); err != nil {
t.Fatal(err)
}
}
if err = zw.Close(); err != nil {
t.Fatal(err)
}
raw := buf.Bytes()
at := bytes.Index(raw, []byte("app.example.com"))
if at < 0 {
t.Fatalf("stored CORS payload not found in archive")
}
raw[at] = 'A'
return raw
}
// TestAdminBucketMetadataCORSRoundTrip covers the export/import round trip for
// per-bucket CORS, per-file error reporting for an invalid document, and that
// an archive without cors.xml leaves an existing configuration alone.
func TestAdminBucketMetadataCORSRoundTrip(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, cred auth.Credentials, t *testing.T) {
corsXML := []byte(testSiteReplicationCORSDoc)
if _, err := updateLocalBucketCORSMetadata(t.Context(), obj, bucket, corsXML); err != nil {
t.Fatal(err)
}
// Export must carry the stored document verbatim.
rec := corsAdminRequest(t, cred, http.MethodGet, "/export-bucket-metadata?bucket="+bucket, nil)
archive := rec.Body.Bytes()
zr, err := zip.NewReader(bytes.NewReader(archive), int64(len(archive)))
if err != nil {
t.Fatal(err)
}
var exported []byte
for _, f := range zr.File {
if f.Name != bucket+"/"+bucketCorsConfig {
continue
}
r, err := f.Open()
if err != nil {
t.Fatal(err)
}
exported, err = io.ReadAll(r)
r.Close()
if err != nil {
t.Fatal(err)
}
}
if !bytes.Equal(exported, corsXML) {
t.Fatalf("%s: exported CORS = %q, want %q", instanceType, exported, corsXML)
}
// Drop the configuration: the archive must then omit the entry.
if _, err = updateLocalBucketCORSMetadata(t.Context(), obj, bucket, nil); err != nil {
t.Fatal(err)
}
if _, _, err = globalBucketMetadataSys.GetCorsConfigXML(bucket); err == nil {
t.Fatalf("%s: CORS still present before restore", instanceType)
}
rec = corsAdminRequest(t, cred, http.MethodGet, "/export-bucket-metadata?bucket="+bucket, nil)
empty := rec.Body.Bytes()
zr, err = zip.NewReader(bytes.NewReader(empty), int64(len(empty)))
if err != nil {
t.Fatal(err)
}
for _, f := range zr.File {
if f.Name == bucket+"/"+bucketCorsConfig {
t.Fatalf("%s: export emitted %s for a bucket without CORS", instanceType, f.Name)
}
}
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata", archive)
if st := corsImportReport(t, rec).Buckets[bucket]; !st.Cors.IsSet || st.Cors.Err != "" {
t.Fatalf("%s: import report cors = %+v", instanceType, st.Cors)
}
stored, storedAt, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: restored CORS = %q, err = %v", instanceType, stored, err)
}
created, err := globalBucketMetadataSys.CreatedAt(bucket)
if err != nil {
t.Fatal(err)
}
if !storedAt.After(created) {
t.Fatalf("%s: restored CORS timestamp %v is not after bucket creation %v", instanceType, storedAt, created)
}
// An archive without cors.xml must not remove the configuration.
corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{bucket + "/quota.json": []byte(`{"quota":0}`)}))
if stored, _, err = globalBucketMetadataSys.GetCorsConfigXML(bucket); err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: import without cors.xml changed CORS: %q, err = %v", instanceType, stored, err)
}
// A bucket the import itself creates must still land above its own
// creation time, otherwise CORS replication would drop the restore.
fresh := "cors-import-created-bucket"
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{fresh + "/" + bucketCorsConfig: corsXML}))
if st := corsImportReport(t, rec).Buckets[fresh]; !st.Cors.IsSet || st.Cors.Err != "" {
t.Fatalf("%s: fresh bucket import report cors = %+v", instanceType, st.Cors)
}
freshStored, freshAt, err := globalBucketMetadataSys.GetCorsConfigXML(fresh)
if err != nil || !bytes.Equal(freshStored, corsXML) {
t.Fatalf("%s: fresh bucket CORS = %q, err = %v", instanceType, freshStored, err)
}
freshCreated, err := globalBucketMetadataSys.CreatedAt(fresh)
if err != nil {
t.Fatal(err)
}
if !freshAt.After(freshCreated) {
t.Fatalf("%s: fresh bucket CORS timestamp %v is not after creation %v", instanceType, freshAt, freshCreated)
}
// An invalid document must fail loudly for that bucket and change nothing.
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{bucket + "/" + bucketCorsConfig: []byte("<CORSConfiguration><CORSRule>")}))
if st := corsImportReport(t, rec).Buckets[bucket]; st.Cors.Err == "" {
t.Fatalf("%s: invalid CORS import reported no error: %+v", instanceType, st)
}
if stored, _, err = globalBucketMetadataSys.GetCorsConfigXML(bucket); err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: invalid CORS import changed stored config: %q, err = %v", instanceType, stored, err)
}
// A well formed document carried by a corrupt zip entry must be
// rejected too, leaving the stored document and its timestamp alone
// while the other configs in the same archive still apply.
_, corsAt, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil {
t.Fatal(err)
}
rec = corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsCorruptedZip(t, bucket+"/"+bucketCorsConfig, corsXML,
map[string][]byte{bucket + "/quota.json": []byte(`{"quota":4096,"quotatype":"hard"}`)}))
st := corsImportReport(t, rec).Buckets[bucket]
if st.Cors.Err == "" {
t.Fatalf("%s: corrupt CORS entry reported no error: %+v", instanceType, st)
}
if !st.Quota.IsSet || st.Quota.Err != "" {
t.Fatalf("%s: corrupt CORS entry blocked the neighboring quota: %+v", instanceType, st.Quota)
}
stored, storedAt, err = globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil || !bytes.Equal(stored, corsXML) || !storedAt.Equal(corsAt) {
t.Fatalf("%s: corrupt CORS entry changed stored config: %q at %v (was %v), err = %v", instanceType, stored, storedAt, corsAt, err)
}
quota, _, err := globalBucketMetadataSys.GetQuotaConfig(t.Context(), bucket)
if err != nil || quota == nil || quota.Quota != 4096 {
t.Fatalf("%s: neighboring quota not applied: %+v, err = %v", instanceType, quota, err)
}
}})
}
// corsPeerStub is a stand-in site-replication peer. It records every
// SRBucketMeta it is asked to apply and answers with status.
func corsPeerStub(t *testing.T, applied chan<- madmin.SRBucketMeta, status int) *httptest.Server {
t.Helper()
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Method == http.MethodPut && applied != nil {
var item madmin.SRBucketMeta
if err := json.NewDecoder(r.Body).Decode(&item); err != nil {
t.Errorf("decode peer apply: %v", err)
w.WriteHeader(http.StatusBadRequest)
return
}
applied <- item
}
w.WriteHeader(status)
}))
}
// TestAdminBucketMetadataCORSImportReplicatesPastPeerFailure pins that an
// imported CORS document reaches the reachable peers even when the shared
// bucket metadata hook failed against an unreachable one, and that both
// failures are still reported for the bucket.
func TestAdminBucketMetadataCORSImportReplicatesPastPeerFailure(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, cred auth.Credentials, t *testing.T) {
ctx := t.Context()
corsXML := []byte(testSiteReplicationCORSDoc)
healthyApplies := make(chan madmin.SRBucketMeta, 4)
healthy := corsPeerStub(t, healthyApplies, http.StatusOK)
defer healthy.Close()
broken := corsPeerStub(t, nil, http.StatusBadRequest)
defer broken.Close()
// With site replication on, admin requests resolve their token signing
// key through the site replicator account, so it has to exist.
serviceCred, err := auth.CreateCredentials(siteReplicatorSvcAcc, "cors-import-service-secret")
if err != nil {
t.Fatal(err)
}
serviceCred.ParentUser = cred.AccessKey
if _, err = globalIAMSys.store.AddServiceAccount(ctx, serviceCred); err != nil {
t.Fatal(err)
}
defer globalIAMSys.DeleteServiceAccount(ctx, serviceCred.AccessKey, false)
globalSiteReplicatorCred.Set(serviceCred.SecretKey)
defer globalSiteReplicatorCred.Set("")
globalSiteReplicationSys.Lock()
oldEnabled, oldState := globalSiteReplicationSys.enabled, globalSiteReplicationSys.state
globalSiteReplicationSys.enabled = true
globalSiteReplicationSys.state = srState{
Name: "cors-import-test",
ServiceAccountAccessKey: serviceCred.AccessKey,
Peers: map[string]madmin.PeerInfo{
globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()},
"peer-healthy": {Name: "healthy", DeploymentID: "peer-healthy", Endpoint: healthy.URL},
"peer-broken": {Name: "broken", DeploymentID: "peer-broken", Endpoint: broken.URL},
},
}
globalSiteReplicationSys.Unlock()
defer func() {
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled, globalSiteReplicationSys.state = oldEnabled, oldState
globalSiteReplicationSys.Unlock()
}()
rec := corsAdminRequest(t, cred, http.MethodPut, "/import-bucket-metadata",
corsZip(t, map[string][]byte{
bucket + "/" + bucketCorsConfig: corsXML,
bucket + "/quota.json": []byte(`{"quota":8192,"quotatype":"hard"}`),
}))
st := corsImportReport(t, rec).Buckets[bucket]
if !st.Cors.IsSet || st.Cors.Err != "" {
t.Fatalf("%s: import report cors = %+v", instanceType, st.Cors)
}
stored, storedAt, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil || !bytes.Equal(stored, corsXML) {
t.Fatalf("%s: stored CORS = %q, err = %v", instanceType, stored, err)
}
// The reachable peer must have been told about the CORS document,
// carrying exactly the timestamp that was saved locally.
var corsSeen, sharedSeen bool
for range 2 {
select {
case item := <-healthyApplies:
if item.Type != madmin.SRBucketMetaTypeCorsConfig {
sharedSeen = item.Bucket == bucket && item.Quota != nil
continue
}
if item.Bucket != bucket || item.Cors == nil || !item.UpdatedAt.Equal(storedAt) {
t.Fatalf("%s: peer CORS event = %#v, want %s at %v", instanceType, item, bucket, storedAt)
}
payload, decErr := base64.StdEncoding.Strict().DecodeString(*item.Cors)
if decErr != nil || !bytes.Equal(payload, corsXML) {
t.Fatalf("%s: peer CORS payload = %q, err = %v", instanceType, payload, decErr)
}
corsSeen = true
case <-time.After(10 * time.Second):
t.Fatalf("%s: healthy peer received no further events (shared=%v cors=%v)", instanceType, sharedSeen, corsSeen)
}
}
if !sharedSeen || !corsSeen {
t.Fatalf("%s: healthy peer events shared=%v cors=%v, want both", instanceType, sharedSeen, corsSeen)
}
// Both hook failures against the unreachable peer stay reported.
if got := strings.Count(st.Err, "->broken:"); got != 2 {
t.Fatalf("%s: bucket error mentions the broken peer %d times, want 2: %q", instanceType, got, st.Err)
}
}})
}
+93 -10
View File
@@ -76,7 +76,7 @@ func (a adminAPIHandlers) PutBucketQuotaConfigHandler(w http.ResponseWriter, r *
return
}
quotaConfig, err := parseBucketQuota(bucket, data)
_, err = parseBucketQuota(bucket, data)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -94,9 +94,6 @@ func (a adminAPIHandlers) PutBucketQuotaConfigHandler(w http.ResponseWriter, r *
Quota: data,
UpdatedAt: updatedAt,
}
if quotaConfig.Size == 0 && quotaConfig.Quota == 0 {
bucketMeta.Quota = nil
}
// Call site replication hook.
replLogIf(ctx, globalSiteReplicationSys.BucketMetaHook(ctx, bucketMeta))
@@ -417,6 +414,7 @@ func (a adminAPIHandlers) ExportBucketMetadataHandler(w http.ResponseWriter, r *
bucketLifecycleConfig,
bucketSSEConfig,
bucketTaggingConfig,
bucketCorsConfig,
bucketQuotaConfigFile,
objectLockConfig,
bucketVersioningConfig,
@@ -437,7 +435,7 @@ func (a adminAPIHandlers) ExportBucketMetadataHandler(w http.ResponseWriter, r *
writeErrorResponse(ctx, w, exportError(ctx, err, cfgFile, bucket), r.URL)
return
}
configData, err := json.Marshal(config)
configData, err := canonicalBucketPolicy(config)
if err != nil {
writeErrorResponse(ctx, w, exportError(ctx, err, cfgFile, bucket), r.URL)
return
@@ -517,6 +515,19 @@ func (a adminAPIHandlers) ExportBucketMetadataHandler(w http.ResponseWriter, r *
return
}
rawDataFn(bytes.NewReader(configData), cfgPath, len(configData))
case bucketCorsConfig:
// Export the stored document verbatim: GetBucketCors returns
// the bytes exactly as they were PUT, so the archive must
// round-trip them unchanged.
configData, _, err := globalBucketMetadataSys.GetCorsConfigXML(bucket)
if err != nil {
if errors.Is(err, errConfigNotFound) {
continue
}
writeErrorResponse(ctx, w, exportError(ctx, err, cfgFile, bucket), r.URL)
return
}
rawDataFn(bytes.NewReader(configData), cfgPath, len(configData))
case objectLockConfig:
config, _, err := globalBucketMetadataSys.GetObjectLockConfig(bucket)
if err != nil {
@@ -616,6 +627,13 @@ func applyImportedBucketMetadata(dst *BucketMetadata, src BucketMetadata, fields
case bucketQuotaConfigFile:
dst.QuotaConfigJSON = bytes.Clone(src.QuotaConfigJSON)
dst.QuotaConfigUpdatedAt = src.QuotaConfigUpdatedAt
case bucketCorsConfig:
// The import stamps its fields before creating any missing bucket,
// and a CORS event stamped before bucket creation is discarded as
// belonging to an older incarnation, so the imported document takes
// the same monotonic timestamp a local PutBucketCors would assign.
dst.CorsConfigUpdatedAt = localCORSUpdatedAt(*dst, src.CorsConfigUpdatedAt)
dst.CorsConfigXML = bytes.Clone(src.CorsConfigXML)
case objectLockConfig:
dst.ObjectLockConfigXML = bytes.Clone(src.ObjectLockConfigXML)
dst.ObjectLockConfigUpdatedAt = src.ObjectLockConfigUpdatedAt
@@ -645,6 +663,8 @@ func (i *importMetaReport) SetStatus(bucket, fname string, err error) {
st.Tagging = madmin.MetaStatus{IsSet: true, Err: errMsg}
case bucketQuotaConfigFile:
st.Quota = madmin.MetaStatus{IsSet: true, Err: errMsg}
case bucketCorsConfig:
st.Cors = madmin.MetaStatus{IsSet: true, Err: errMsg}
case objectLockConfig:
st.ObjectLock = madmin.MetaStatus{IsSet: true, Err: errMsg}
case bucketVersioningConfig:
@@ -899,7 +919,7 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
continue
}
configData, err := json.Marshal(bucketPolicy)
configData, err := canonicalBucketPolicy(bucketPolicy)
if err != nil {
rpt.SetStatus(bucket, fileName, err)
continue
@@ -1015,6 +1035,32 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
bucketMap[bucket].QuotaConfigUpdatedAt = updatedAt
markImported(bucket, fileName)
rpt.SetStatus(bucket, fileName, nil)
case bucketCorsConfig:
if sz > maxBucketCorsSize {
rpt.SetStatus(bucket, fileName, errors.New(ErrEntityTooLarge.String()))
continue
}
// Read one byte past the declared size: stopping exactly at sz
// leaves archive/zip short of EOF, so it never verifies the entry
// checksum and a corrupt entry carrying well formed XML would be
// stored as a valid document. The extra byte also lets the reader
// reject an entry longer than it declares.
corsData, err := io.ReadAll(io.LimitReader(reader, sz+1))
if err != nil {
rpt.SetStatus(bucket, fileName, err)
continue
}
if err = validateCORSReplicationPayload(corsData); err != nil {
rpt.SetStatus(bucket, fileName, fmt.Errorf("%s (%s)", errorCodes[ErrMalformedXML].Description, err))
continue
}
bucketMap[bucket].CorsConfigXML = corsData
bucketMap[bucket].CorsConfigUpdatedAt = updatedAt
markImported(bucket, fileName)
rpt.SetStatus(bucket, fileName, nil)
}
}
@@ -1032,18 +1078,39 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
continue
}
var merged BucketMetadata
var commitAt time.Time
err := func() error {
lockCtx, unlock, err := lockBucketMetadata(ctx, objectAPI, bucket)
if err != nil {
return err
}
defer unlock()
merged, err = loadBucketMetadataParse(lockCtx, objectAPI, bucket, true)
merged, err = loadBucketMetadataParse(lockCtx, objectAPI, bucket, false)
if err != nil {
return err
}
if err := ensureBucketMetadataCreated(lockCtx, objectAPI, &merged); err != nil {
return err
}
commitAt = UTCNow()
for _, file := range replicatedBucketConfigs {
if _, ok := fields[file]; ok {
commitAt = localBucketConfigUpdatedAt(merged, file, commitAt)
}
}
applyImportedBucketMetadata(&merged, *meta, fields)
return globalBucketMetadataSys.saveMetadata(lockCtx, objectAPI, merged)
for _, file := range replicatedBucketConfigs {
if _, ok := fields[file]; !ok {
continue
}
data, at := replicatedBucketConfig(&merged, file)
payload, _, err := bucketConfigPayload(bucket, file, *data, len(merged.ObjectLockConfigXML) != 0)
if err != nil {
return err
}
*data, *at = payload, commitAt
}
return globalBucketMetadataSys.saveMetadata(lockCtx, objectAPI, &merged)
}()
if err != nil {
rpt.SetStatus(bucket, "", err)
@@ -1051,7 +1118,7 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
}
*meta = merged
globalNotificationSys.LoadBucketMetadata(bgContext(ctx), bucket)
hook := madmin.SRBucketMeta{Bucket: bucket, UpdatedAt: updatedAt}
hook := madmin.SRBucketMeta{Bucket: bucket, UpdatedAt: commitAt}
var hookNeeded bool
if _, ok := fields[bucketQuotaConfigFile]; ok {
hook.Quota = meta.QuotaConfigJSON
@@ -1059,7 +1126,7 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
}
if _, ok := fields[bucketPolicyConfig]; ok {
hook.Policy = meta.PolicyConfigJSON
hookNeeded = true
hookNeeded = hookNeeded || len(hook.Policy) != 0
}
if _, ok := fields[bucketVersioningConfig]; ok {
hook.Versioning = enc(meta.VersioningConfigXML)
@@ -1080,6 +1147,22 @@ func (a adminAPIHandlers) ImportBucketMetadataHandler(w http.ResponseWriter, r *
if hookNeeded {
err = globalSiteReplicationSys.BucketMetaHook(ctx, hook)
}
if _, ok := fields[bucketPolicyConfig]; ok && len(meta.PolicyConfigJSON) == 0 {
// An omitted bulk Policy cannot express deletion.
err = errors.Join(err, globalSiteReplicationSys.BucketMetaHook(ctx, madmin.SRBucketMeta{
Type: madmin.SRBucketMetaTypePolicy, Bucket: bucket, UpdatedAt: commitAt,
}))
}
if _, ok := fields[bucketCorsConfig]; ok {
// CORS carries its own timestamp, so it replicates through the
// dedicated event rather than the shared bucket metadata hook. It
// is announced even when the shared hook failed: the document is
// already committed locally, and a peer that is unreachable for
// one config must not withhold CORS from the reachable ones.
if corsEvent, live := newBucketCORSReplicationEvent(bucket, *meta); live {
err = errors.Join(err, globalSiteReplicationSys.BucketMetaHook(ctx, corsEvent))
}
}
if err != nil {
rpt.SetStatus(bucket, "", err)
continue
+5 -1
View File
@@ -502,11 +502,15 @@ func (a adminAPIHandlers) AddUser(w http.ResponseWriter, r *http.Request) {
}
checkDenyOnly := accessKey == cred.AccessKey
action := policy.Action(policy.CreateUserAdminAction)
if checkDenyOnly {
action = policy.ChangeMyPasswordAdminAction
}
if !globalIAMSys.IsAllowed(policy.Args{
AccountName: cred.AccessKey,
Groups: cred.Groups,
Action: policy.CreateUserAdminAction,
Action: action,
ConditionValues: getConditionValues(r, "", cred),
IsOwner: owner,
Claims: cred.Claims,
+102
View File
@@ -243,6 +243,7 @@ func TestIAMInternalIDPServerSuite(t *testing.T) {
suite.SetUpSuite(c)
suite.TestUserCreate(c)
suite.TestUserPasswordActionAuthorization(c)
suite.TestUserStatusActionAuthorization(c)
suite.TestGroupStatusActionAuthorization(c)
suite.TestUserPolicyEscalationBug(c)
@@ -356,6 +357,106 @@ func (s *TestSuiteIAM) TestUserCreate(c *check) {
}
}
func (s *TestSuiteIAM) TestUserPasswordActionAuthorization(c *check) {
for _, tt := range []struct {
name string
statements string
self bool
other bool
}{
{"readonly", "", true, false},
{"consolereadonly", "", true, false},
{"password grant", `{"Effect":"Allow","Action":"admin:ChangeMyPassword"}`, true, false},
{"legacy CreateUser deny", `{"Effect":"Deny","Action":"admin:CreateUser","Resource":"arn:aws:s3:::*"}`, true, false},
{"password deny", `{"Effect":"Deny","Action":"admin:ChangeMyPassword"}`, false, false},
{"user admin", `{"Effect":"Allow","Action":"admin:CreateUser"}`, true, true},
{"user admin with password deny", `{"Effect":"Allow","Action":"admin:CreateUser"},{"Effect":"Deny","Action":"admin:ChangeMyPassword"}`, false, true},
{"password deny overrides grant", `{"Effect":"Allow","Action":"admin:ChangeMyPassword"},{"Effect":"Deny","Action":"admin:ChangeMyPassword"}`, false, false},
{"wildcard deny", `{"Effect":"Deny","Action":"admin:*"}`, false, false},
} {
c.Run(tt.name, func(t *testing.T) {
c := &check{t, s.serverType}
ctx, cancel := context.WithTimeout(context.Background(), testDefaultTimeout)
defer cancel()
var users []string
policyName := tt.name
defer func() {
for _, user := range users {
if err := s.adm.RemoveUser(ctx, user); err != nil {
c.Errorf("remove test user: %v", err)
}
}
if tt.statements != "" {
if err := s.adm.RemoveCannedPolicy(ctx, policyName); err != nil {
c.Errorf("remove test policy: %v", err)
}
}
}()
createUser := func() (string, string) {
accessKey, secretKey := mustGenerateCredentials(c)
if err := s.adm.SetUser(ctx, accessKey, secretKey, madmin.AccountEnabled); err != nil {
c.Fatalf("create test user: %v", err)
}
users = append(users, accessKey)
return accessKey, secretKey
}
client := func(accessKey, secretKey string) *madmin.AdminClient {
adm, err := madmin.New(s.endpoint, accessKey, secretKey, s.secure)
if err != nil {
c.Fatal(err)
}
adm.SetCustomTransport(s.TestSuiteCommon.client.Transport)
return adm
}
if tt.statements != "" {
policyName = getRandomBucketName()
doc := []byte(`{"Version":"2012-10-17","Statement":[` + tt.statements + `]}`)
if err := s.adm.AddCannedPolicy(ctx, policyName, doc); err != nil {
c.Fatalf("save test policy: %v", err)
}
}
accessKey, secretKey := createUser()
if _, err := s.adm.AttachPolicy(ctx, madmin.PolicyAssociationReq{
User: accessKey, Policies: []string{policyName},
}); err != nil {
c.Fatalf("attach test policy: %v", err)
}
adm := client(accessKey, secretKey)
_, newSecretKey := mustGenerateCredentials(c)
err := adm.SetUser(ctx, accessKey, newSecretKey, madmin.AccountEnabled)
if tt.self {
if err != nil {
c.Fatalf("change own password: %v", err)
}
if _, err = adm.AccountInfo(ctx, madmin.AccountOpts{}); err == nil {
c.Fatal("old password still authenticates")
}
adm = client(accessKey, newSecretKey)
} else if err == nil || madmin.ToErrorResponse(err).Code != "AccessDenied" {
c.Fatalf("self password change: expected AccessDenied, got %v", err)
}
if _, err := adm.AccountInfo(ctx, madmin.AccountOpts{}); err != nil {
c.Fatalf("current password no longer authenticates: %v", err)
}
target, _ := createUser()
newUser, newUserSecret := mustGenerateCredentials(c)
for _, key := range []string{target, newUser} {
err := adm.SetUser(ctx, key, newUserSecret, madmin.AccountEnabled)
if tt.other {
if err != nil {
c.Fatalf("create or update another user: %v", err)
}
if key == newUser {
users = append(users, newUser)
}
} else if err == nil || madmin.ToErrorResponse(err).Code != "AccessDenied" {
c.Fatalf("create or update another user: expected AccessDenied, got %v", err)
}
}
})
}
}
func (s *TestSuiteIAM) TestUserStatusActionAuthorization(c *check) {
ctx, cancel := context.WithTimeout(context.Background(), testDefaultTimeout)
defer cancel()
@@ -946,6 +1047,7 @@ func (s *TestSuiteIAM) TestCannedPolicies(c *check) {
defaultPolicies := []string{
"readwrite",
"readonly",
"consolereadonly",
"writeonly",
"diagnostics",
"consoleAdmin",
+5 -1
View File
@@ -1523,10 +1523,14 @@ var errorCodes = errorCodeMap{
Description: "Your Host header is malformed.",
HTTPStatusCode: http.StatusBadRequest,
},
// The stored object cannot be served: a server-side data condition, not a
// successful partial read. Upstream maps it to http.StatusPartialContent
// (since ca6b4773e, 2017), which lets SDKs accept the XML error document
// as object content; SILO deliberately diverges and returns 500.
ErrObjectTampered: {
Code: "XMinioObjectTampered",
Description: errObjectTampered.Error(),
HTTPStatusCode: http.StatusPartialContent,
HTTPStatusCode: http.StatusInternalServerError,
},
ErrSiteReplicationInvalidRequest: {
+4 -1
View File
@@ -212,7 +212,10 @@ func setObjectHeaders(ctx context.Context, w http.ResponseWriter, objInfo Object
}
if rs == nil && opts.PartNumber > 0 {
rs = partNumberToRangeSpec(objInfo, opts.PartNumber)
rs, err = partNumberToRangeSpec(objInfo, opts.PartNumber)
if err != nil {
return err
}
}
// For providing ranged content
+4 -11
View File
@@ -632,18 +632,11 @@ func isReqAuthenticated(ctx context.Context, r *http.Request, region string, sty
return ErrInvalidDigest
}
// Extract either 'X-Amz-Content-Sha256' header or 'X-Amz-Content-Sha256' query parameter (if V4 presigned)
// Do not verify 'X-Amz-Content-Sha256' if skipSHA256.
// Honor the selected header/query checksum, including the header fallback
// for a presigned request. STS separately hashes its body for its signature.
var contentSHA256 []byte
if skipSHA256 := skipContentSha256Cksum(r); !skipSHA256 && isRequestPresignedSignatureV4(r) {
if sha256Sum, ok := r.Form[xhttp.AmzContentSha256]; ok && len(sha256Sum) > 0 {
contentSHA256, err = hex.DecodeString(sha256Sum[0])
if err != nil {
return ErrContentSHA256Mismatch
}
}
} else if _, ok := r.Header[xhttp.AmzContentSha256]; !skipSHA256 && ok {
contentSHA256, err = hex.DecodeString(r.Header.Get(xhttp.AmzContentSha256))
if !skipContentSha256Cksum(r) {
contentSHA256, err = hex.DecodeString(getContentSha256Cksum(r, serviceS3))
if err != nil || len(contentSHA256) == 0 {
return ErrContentSHA256Mismatch
}
-19
View File
@@ -467,25 +467,6 @@ func testAdversarialHealCorsPropagatesNewerEqualValueTimestamp(obj ObjectLayer,
}
}
func TestAdversarialBucketMetadataComparisonIsBase64CaseSensitive(t *testing.T) {
upper := "QQ=="
lower := "qQ=="
upperBytes, err := base64.StdEncoding.Strict().DecodeString(upper)
if err != nil {
t.Fatal(err)
}
lowerBytes, err := base64.StdEncoding.Strict().DecodeString(lower)
if err != nil {
t.Fatal(err)
}
if string(upperBytes) == string(lowerBytes) {
t.Fatal("test inputs unexpectedly decode to the same bytes")
}
if isBucketMetadataEqual(&upper, &lower) {
t.Fatal("different decoded payloads were treated as equal")
}
}
func TestSiteReplicationStatusDetectsCorsTimestampMismatch(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
+1
View File
@@ -39,6 +39,7 @@ const (
lcEventSrc_s3PutObject
lcEventSrc_s3CopyObject
lcEventSrc_s3CompleteMultipartUpload
lcEventSrc_AccessTier
)
//revive:enable:var-naming
+111
View File
@@ -0,0 +1,111 @@
// Copyright 2026 PGSTY contributors.
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"errors"
"fmt"
"net/http"
"sync"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
func TestDeleteBucketMetadataLockCancellation(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testDeleteBucketMetadataLockCancellation})
}
func testDeleteBucketMetadataLockCancellation(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
_, unlock, err := lockBucketMetadata(t.Context(), obj, bucket)
if err != nil {
t.Fatal(err)
}
release := sync.OnceFunc(unlock)
defer release()
// Observe DeleteBucket's ACTUAL metadata.lock attempt. Set the hook after
// our own acquisition above so it only trips on the delete.
delAtLock := make(chan struct{})
var once sync.Once
hook := func(b string) {
if b == bucket {
once.Do(func() { close(delAtLock) })
}
}
lockBucketMetadataAcquireHook.Store(&hook)
defer lockBucketMetadataAcquireHook.Store(nil)
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
done := make(chan error, 1)
go func() { done <- obj.DeleteBucket(ctx, bucket, DeleteBucketOptions{Force: true, NoLock: true}) }()
select {
case <-delAtLock:
// Fixed tree: the delete reached metadata.lock and is blocking on the
// lock we hold. Cancel it and confirm it fails WITHOUT deleting, while we
// still hold the lock (release stays deferred until after the checks).
cancel()
if err := <-done; err == nil {
t.Errorf("%s: canceled deletion succeeded", instanceType)
}
if _, err := obj.GetBucketInfo(t.Context(), bucket, BucketOptions{}); err != nil {
t.Errorf("%s: bucket disappeared while metadata.lock was held: %v", instanceType, err)
}
if _, err := readBucketMetadata(t.Context(), obj, bucket); err != nil {
t.Errorf("%s: canceled deletion removed metadata: %v", instanceType, err)
}
case err := <-done:
// Broken tree: the delete finished without ever taking metadata.lock,
// i.e. it did not serialize the destructive operation behind the lock.
t.Errorf("%s: delete bypassed metadata.lock (err=%v)", instanceType, err)
}
}
func TestQueuedMetadataUpdateAfterDelete(t *testing.T) {
defer DetectTestLeak(t)()
for _, expiry := range []bool{false, true} {
t.Run(fmt.Sprintf("expiry=%v", expiry), func(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
previous := newObjectLayerFn()
barrier := &lcMergeBarrier{ObjectLayer: obj, bucket: bucket, mAtLock: make(chan struct{}), mProceed: make(chan struct{})}
setObjectLayer(barrier)
defer setObjectLayer(previous)
release := sync.OnceFunc(func() { close(barrier.mProceed) })
defer release()
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
ctx = context.WithValue(ctx, lcMergeWriterKey{}, "M")
done := make(chan error, 1)
go func() {
if expiry {
done <- globalBucketMetadataSys.UpdateExpiryLCConfig(ctx, bucket, nil, UTCNow())
return
}
_, err := globalBucketMetadataSys.Update(ctx, bucket, bucketTaggingConfig, []byte(`<Tagging><TagSet/></Tagging>`))
done <- err
}()
select {
case <-barrier.mAtLock:
case <-ctx.Done():
t.Fatal("writer did not reach metadata.lock")
}
if err := obj.DeleteBucket(t.Context(), bucket, DeleteBucketOptions{Force: true}); err != nil {
t.Fatal(err)
}
release()
if err := <-done; !isErrBucketNotFound(err) {
t.Errorf("%s: queued update should reject a deleted bucket, got %v", instanceType, err)
}
if _, err := readBucketMetadata(t.Context(), obj, bucket); !errors.Is(err, errConfigNotFound) && !isErrBucketNotFound(err) {
t.Errorf("%s: queued update recreated metadata: %v", instanceType, err)
}
}})
})
}
}
+278
View File
@@ -0,0 +1,278 @@
// Copyright (c) 2015-2026 MinIO, Inc.
// Copyright (c) 2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
package cmd
import (
"context"
"encoding/base64"
"errors"
"net/http"
"sync"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
// ---------------------------------------------------------------------------
// Target 1 (issue #105): a higher-level lifecycle XML merge racing another
// bucket-metadata transition outside BucketMetadataSys.Delete.
//
// Before the fix, PeerBucketLCConfigHandler / healBucketILMExpiry read the
// current lifecycle document without metadata.lock, merged the replicated expiry
// rules with the local transition rules, and then persisted the merged blob with
// BucketMetadataSys.Update. Update re-read the record under metadata.lock but
// overwrote LifecycleConfigXML wholesale with the pre-computed blob, so any
// lifecycle transition change committed between the merge read and the merge
// write was silently lost on disk. BucketMetadataSys.UpdateExpiryLCConfig now
// performs the read, merge, and save under a single metadata.lock. This test
// drives PeerBucketLCConfigHandler and asserts the concurrent change survives.
// ---------------------------------------------------------------------------
type lcMergeWriterKey struct{}
// lcMergeBarrier pauses the merge writer (context value "M") exactly when it
// tries to take metadata.lock for its persisting Update. By that point the
// merge has already read the stale lifecycle document, so the test can commit a
// concurrent transition change before releasing the merge write.
type lcMergeBarrier struct {
ObjectLayer
bucket string
mAtLock chan struct{}
mProceed chan struct{}
mOnce sync.Once
}
func (o *lcMergeBarrier) metadataLock() string {
return pathJoin(bucketMetaPrefix, o.bucket, "metadata.lock")
}
func (o *lcMergeBarrier) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket != minioMetaBucket || len(objects) != 1 || objects[0] != o.metadataLock() {
return lock
}
return metadataObservedRWLocker{RWLocker: lock, onLock: func(ctx context.Context) {
if ctx.Value(lcMergeWriterKey{}) != "M" {
return
}
o.mOnce.Do(func() { close(o.mAtLock) })
select {
case <-o.mProceed:
case <-ctx.Done():
}
}}
}
func TestLifecycleExpiryMergeRaceLosesConcurrentTransition(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testLifecycleExpiryMergeRaceLosesConcurrentTransition,
})
}
func testLifecycleExpiryMergeRaceLosesConcurrentTransition(obj ObjectLayer, instanceType, bucket string,
_ http.Handler, _ auth.Credentials, t *testing.T,
) {
// Revision N: a single transition-only rule.
baseXML := []byte(`<LifecycleConfiguration><Rule><ID>keep</ID><Filter><Prefix>data/</Prefix></Filter><Status>Enabled</Status><Transition><Days>30</Days><StorageClass>WARM</StorageClass></Transition></Rule></LifecycleConfiguration>`)
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketLifecycleConfig, baseXML); err != nil {
t.Fatalf("%s: seed lifecycle: %v", instanceType, err)
}
// A replicated expiry-only rule arriving from a peer site.
expXML := `<LifecycleConfiguration><Rule><ID>expire</ID><Filter><Prefix>tmp/</Prefix></Filter><Status>Enabled</Status><Expiration><Days>7</Days></Expiration></Rule></LifecycleConfiguration>`
expLCConfig := base64.StdEncoding.EncodeToString([]byte(expXML))
previousObjectAPI := newObjectLayerFn()
barrier := &lcMergeBarrier{
ObjectLayer: obj,
bucket: bucket,
mAtLock: make(chan struct{}),
mProceed: make(chan struct{}),
}
setObjectLayer(barrier)
defer setObjectLayer(previousObjectAPI)
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
mCtx := context.WithValue(ctx, lcMergeWriterKey{}, "M")
mDone := make(chan error, 1)
go func() {
mDone <- globalSiteReplicationSys.PeerBucketLCConfigHandler(mCtx, bucket, &expLCConfig, UTCNow())
}()
// Wait until the merge writer has read revision N and is about to persist.
select {
case <-barrier.mAtLock:
case err := <-mDone:
t.Fatalf("%s: merge writer finished before persisting: %v", instanceType, err)
case <-ctx.Done():
t.Fatalf("%s: merge writer never reached metadata.lock: %v", instanceType, ctx.Err())
}
// Concurrent local lifecycle transition change commits revision N+1.
concurrentXML := []byte(`<LifecycleConfiguration><Rule><ID>keep</ID><Filter><Prefix>data/</Prefix></Filter><Status>Enabled</Status><Transition><Days>10</Days><StorageClass>COLD</StorageClass></Transition></Rule></LifecycleConfiguration>`)
if _, err := globalBucketMetadataSys.Update(ctx, bucket, bucketLifecycleConfig, concurrentXML); err != nil {
t.Fatalf("%s: concurrent transition update: %v", instanceType, err)
}
// Release the merge write so it lands after the concurrent commit.
close(barrier.mProceed)
if err := <-mDone; err != nil {
t.Fatalf("%s: merge writer failed: %v", instanceType, err)
}
// The persisted lifecycle must contain both the replicated expiry rule and
// the concurrent transition change.
cfg, _, err := globalBucketMetadataSys.GetLifecycleConfig(bucket)
if err != nil {
t.Fatalf("%s: read merged lifecycle: %v", instanceType, err)
}
var keep, expire bool
for i := range cfg.Rules {
switch cfg.Rules[i].ID {
case "keep":
keep = true
if cfg.Rules[i].Transition.Days != 10 || cfg.Rules[i].Transition.StorageClass != "COLD" {
t.Fatalf("%s: lifecycle merge overwrote the concurrent transition change: got Days=%d StorageClass=%q, want Days=10 StorageClass=COLD",
instanceType, cfg.Rules[i].Transition.Days, cfg.Rules[i].Transition.StorageClass)
}
case "expire":
expire = true
}
}
if !expire {
t.Fatalf("%s: merged lifecycle dropped the replicated expiry rule: %+v", instanceType, cfg.Rules)
}
if !keep {
t.Fatalf("%s: merged lifecycle dropped the transition rule entirely: %+v", instanceType, cfg.Rules)
}
}
// ---------------------------------------------------------------------------
// Target 2 (issue #105): DeleteBucket racing an in-flight metadata writer must
// not resurrect a ghost .metadata.bin record.
//
// Before the fix, erasureServerPools.DeleteBucket took only <bucket>.lck and
// purged the metadata prefix while config writers (updateAndParse) took only
// metadata.lock, so a writer already MID-SAVE (holding metadata.lock, past
// saveMetadata's existence recheck) could persist .metadata.bin after the purge.
// DeleteBucket now takes metadata.lock before deleting, so it waits for that
// writer and then purges whatever the writer wrote.
//
// This test isolates the DeleteBucket-lock fix specifically: the writer holds
// metadata.lock and is paused at the .metadata.bin PutObject, so the earlier
// saveMetadata existence recheck cannot save it — only serializing the delete
// behind the writer can. Removing just DeleteBucket's metadata.lock (keeping the
// recheck) therefore makes this test fail. The delete's ACTUAL metadata.lock
// attempt is observed with lockBucketMetadataAcquireHook (its lock is taken
// through the erasureServerPools receiver, invisible to the object-layer
// barrier), so the handshake is deterministic with no timing assumption.
// ---------------------------------------------------------------------------
func TestDeleteBucketResurrectsGhostMetadata(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testDeleteBucketResurrectsGhostMetadata,
})
}
func testDeleteBucketResurrectsGhostMetadata(obj ObjectLayer, instanceType, bucket string,
_ http.Handler, _ auth.Credentials, t *testing.T,
) {
previousObjectAPI := newObjectLayerFn()
// Writer A holds metadata.lock and pauses at the .metadata.bin PutObject,
// i.e. already past saveMetadata's existence recheck and mid-save.
barrier := &metadataRMWBarrierObjectLayer{
ObjectLayer: obj,
bucket: bucket,
aReady: make(chan struct{}),
aRelease: make(chan struct{}),
bLockAttempt: make(chan struct{}),
}
setObjectLayer(barrier)
defer setObjectLayer(previousObjectAPI)
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
aCtx := context.WithValue(ctx, metadataRMWWriterKey{}, "A")
policyJSON := []byte(`{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":"*","Action":"s3:GetObject","Resource":"arn:aws:s3:::` + bucket + `/*"}]}`)
aReleased := sync.OnceFunc(func() { close(barrier.aRelease) })
defer aReleased()
aDone := make(chan error, 1)
go func() {
_, err := globalBucketMetadataSys.Update(aCtx, bucket, bucketPolicyConfig, policyJSON)
aDone <- err
}()
select {
case <-barrier.aReady:
case err := <-aDone:
t.Fatalf("%s: writer A finished before persisting: %v", instanceType, err)
case <-ctx.Done():
t.Fatalf("%s: writer A never reached metadata save: %v", instanceType, ctx.Err())
}
// A now holds metadata.lock mid-save. Observe DeleteBucket's ACTUAL
// metadata.lock attempt via the acquire hook, set only now so A's earlier
// acquisition does not trip it.
delAtLock := make(chan struct{})
var once sync.Once
hook := func(b string) {
if b == bucket {
once.Do(func() { close(delAtLock) })
}
}
lockBucketMetadataAcquireHook.Store(&hook)
defer lockBucketMetadataAcquireHook.Store(nil)
delDone := make(chan error, 1)
go func() {
delDone <- obj.DeleteBucket(ctx, bucket, DeleteBucketOptions{Force: true})
}()
select {
case <-delAtLock:
// Fixed tree: DeleteBucket reached metadata.lock and blocks on A. Release
// A so it finishes its save and unlocks; the delete then acquires the
// lock and purges the record A wrote.
aReleased()
if err := <-aDone; err != nil {
t.Fatalf("%s: writer A save failed while holding metadata.lock: %v", instanceType, err)
}
if err := <-delDone; err != nil {
t.Fatalf("%s: delete bucket: %v", instanceType, err)
}
case err := <-delDone:
// Broken tree: DeleteBucket purged without taking metadata.lock. Release
// A so its mid-save PutObject recreates .metadata.bin (the ghost).
if err != nil {
t.Fatalf("%s: delete bucket: %v", instanceType, err)
}
aReleased()
if err := <-aDone; err != nil {
t.Fatalf("%s: writer A save failed: %v", instanceType, err)
}
}
// After DeleteBucket, no .metadata.bin record may remain on disk.
if meta, err := readBucketMetadata(ctx, obj, bucket); err == nil {
t.Fatalf("%s: ghost .metadata.bin resurrected after DeleteBucket: name=%q created=%s policyLen=%d",
instanceType, meta.Name, meta.Created, len(meta.PolicyConfigJSON))
} else if !errors.Is(err, errConfigNotFound) && !isErrBucketNotFound(err) && !errors.Is(err, errVolumeNotFound) {
t.Fatalf("%s: unexpected error reading deleted bucket metadata: %v", instanceType, err)
}
}
+54
View File
@@ -0,0 +1,54 @@
// Copyright 2026 PGSTY contributors.
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"net/http"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/grid"
)
func TestPeerMetadataReloadWithEqualMaximumTimestamp(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testPeerMetadataReloadWithEqualMaximumTimestamp})
}
func testPeerMetadataReloadWithEqualMaximumTimestamp(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
disk, err := loadBucketMetadata(t.Context(), obj, bucket)
if err != nil {
t.Fatal(err)
}
disk.TaggingConfigXML = []byte(`<Tagging><TagSet><Tag><Key>revision</Key><Value>new</Value></Tag></TagSet></Tagging>`)
disk.TaggingConfigUpdatedAt = UTCNow()
// An unrelated configuration has the greatest timestamp in both records.
disk.PolicyConfigUpdatedAt = disk.TaggingConfigUpdatedAt.Add(time.Hour)
if err := disk.Save(t.Context(), obj); err != nil {
t.Fatal(err)
}
resident := disk
resident.TaggingConfigXML = []byte(`<Tagging><TagSet><Tag><Key>revision</Key><Value>old</Value></Tag></TagSet></Tagging>`)
resident.TaggingConfigUpdatedAt = disk.TaggingConfigUpdatedAt.Add(-time.Minute)
if err := resident.parseAllConfigs(t.Context(), obj); err != nil {
t.Fatal(err)
}
if !resident.lastUpdate().Equal(disk.lastUpdate()) {
t.Fatal("fixture must have equal maximum timestamps")
}
globalBucketMetadataSys.Set(bucket, resident)
args := grid.MSS{peerRESTBucket: bucket}
if _, err := (&peerRESTServer{}).LoadBucketMetadataHandler(&args); err != nil {
t.Fatal(err)
}
tagging, _, err := globalBucketMetadataSys.GetTaggingConfig(bucket)
if err != nil {
t.Fatal(err)
}
if got := tagging.ToMap()["revision"]; got != "new" {
t.Errorf("%s: peer reload retained tag %q despite a newer tagging configuration", instanceType, got)
}
}
+330
View File
@@ -0,0 +1,330 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful,
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"encoding/json"
"errors"
"sort"
"time"
"github.com/minio/minio-go/v7/pkg/tags"
bucketsse "github.com/minio/minio/internal/bucket/encryption"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
"github.com/minio/minio/internal/bucket/versioning"
"github.com/pgsty/silo-pkg/v3/policy"
)
// Only these fields share the site-replication source-time ordering contract.
// Bulk apply/import process Object Lock before Versioning, whose effective
// document depends on it. Periodic heal retains its existing type order.
var replicatedBucketConfigs = [...]string{
objectLockConfig, bucketVersioningConfig, bucketPolicyConfig,
bucketTaggingConfig, bucketSSEConfig, bucketQuotaConfigFile,
}
func replicatedBucketConfig(meta *BucketMetadata, file string) (*[]byte, *time.Time) {
switch file {
case bucketPolicyConfig:
return &meta.PolicyConfigJSON, &meta.PolicyConfigUpdatedAt
case bucketTaggingConfig:
return &meta.TaggingConfigXML, &meta.TaggingConfigUpdatedAt
case bucketSSEConfig:
return &meta.EncryptionConfigXML, &meta.EncryptionConfigUpdatedAt
case bucketQuotaConfigFile:
return &meta.QuotaConfigJSON, &meta.QuotaConfigUpdatedAt
case bucketVersioningConfig:
return &meta.VersioningConfigXML, &meta.VersioningConfigUpdatedAt
case objectLockConfig:
return &meta.ObjectLockConfigXML, &meta.ObjectLockConfigUpdatedAt
}
return nil, nil
}
// Callers that only need to know whether a file is under the contract must not
// probe replicatedBucketConfig with a throwaway BucketMetadata.
func isReplicatedBucketConfig(file string) bool {
for _, replicated := range replicatedBucketConfigs {
if replicated == file {
return true
}
}
return false
}
func bucketConfigUpdateOnly(file string) bool {
return file == bucketVersioningConfig || file == objectLockConfig
}
// Reuse the persistence rule before comparison, so an accepted Versioning
// event and the document Save actually writes have the same comparison key.
func effectiveBucketVersioning(data []byte, lockEnabled bool) []byte {
if lockEnabled {
config, err := versioning.ParseConfig(bytes.NewReader(data))
if err != nil || !config.Enabled() || config.PrefixesExcluded() {
return enabledBucketVersioningConfig
}
}
return data
}
// A parsed policy still contains map-backed sets with nondeterministic Marshal
// order. Sort every set array recursively, including statements and conditions.
// RawMessage keeps integer values intact; decoding through float64 would not.
func canonicalBucketPolicyJSON(data json.RawMessage) (json.RawMessage, error) {
data = bytes.TrimSpace(data)
if len(data) == 0 {
return nil, nil
}
switch data[0] {
case '{':
var obj map[string]json.RawMessage
if err := json.Unmarshal(data, &obj); err != nil {
return nil, err
}
for key, value := range obj {
var err error
obj[key], err = canonicalBucketPolicyJSON(value)
if err != nil {
return nil, err
}
}
return json.Marshal(obj)
case '[':
var arr []json.RawMessage
if err := json.Unmarshal(data, &arr); err != nil {
return nil, err
}
for i := range arr {
var err error
arr[i], err = canonicalBucketPolicyJSON(arr[i])
if err != nil {
return nil, err
}
}
sort.Slice(arr, func(i, j int) bool { return bytes.Compare(arr[i], arr[j]) < 0 })
return json.Marshal(arr)
default:
var compact bytes.Buffer
if err := json.Compact(&compact, data); err != nil {
return nil, err
}
return compact.Bytes(), nil
}
}
// Encode the validated policy fields explicitly: BPStatement's required
// Action/Resource tags otherwise try to marshal empty sets for the supported
// NotAction/NotResource alternatives. This stays within the existing schema.
func canonicalBucketPolicy(cfg *policy.BucketPolicy) ([]byte, error) {
if cfg.IsEmpty() {
return nil, nil
}
doc := map[string]any{"Version": cfg.Version}
if cfg.ID != "" {
doc["ID"] = cfg.ID
}
statements := make([]map[string]any, 0, len(cfg.Statements))
for _, st := range cfg.Statements {
statement := map[string]any{"Effect": st.Effect, "Principal": st.Principal}
if st.SID != "" {
statement["Sid"] = st.SID
}
if len(st.Actions) != 0 {
statement["Action"] = st.Actions
}
if len(st.NotActions) != 0 {
statement["NotAction"] = st.NotActions
}
if len(st.Resources) != 0 {
statement["Resource"] = st.Resources
}
if len(st.NotResources) != 0 {
statement["NotResource"] = st.NotResources
}
if len(st.Conditions) != 0 {
statement["Condition"] = st.Conditions
}
statements = append(statements, statement)
}
doc["Statement"] = statements
data, err := json.Marshal(doc)
if err != nil {
return nil, err
}
return canonicalBucketPolicyJSON(data)
}
// Validate with the same parsers as Save. Policy's established empty-policy
// semantics are deletion; a parsed zero quota is still a live document.
func bucketConfigPayload(bucket, file string, data []byte, lockEnabled bool) ([]byte, []byte, error) {
if len(data) == 0 {
return nil, nil, nil
}
var err error
key := data
switch file {
case bucketPolicyConfig:
var cfg *policy.BucketPolicy
cfg, err = policy.ParseBucketPolicyConfig(bytes.NewReader(data), bucket)
if err == nil {
if cfg.IsEmpty() {
return nil, nil, nil
}
key, err = canonicalBucketPolicy(cfg)
}
case bucketQuotaConfigFile:
cfg, parseErr := parseBucketQuota(bucket, data)
err = parseErr
if err == nil {
key, err = json.Marshal(cfg)
}
case bucketTaggingConfig:
_, err = tags.ParseBucketXML(bytes.NewReader(data))
case bucketSSEConfig:
_, err = bucketsse.ParseBucketSSEConfig(bytes.NewReader(data))
case objectLockConfig:
_, err = objectlock.ParseObjectLockConfig(bytes.NewReader(data))
case bucketVersioningConfig:
data = effectiveBucketVersioning(data, lockEnabled)
key = data
_, err = versioning.ParseConfig(bytes.NewReader(data))
}
return data, key, err
}
type bucketConfigState struct {
data, key []byte
at time.Time
real, valid bool
}
func newBucketConfigState(bucket, file string, data []byte, at, created time.Time, lockEnabled bool) (bucketConfigState, error) {
data, key, err := bucketConfigPayload(bucket, file, data, lockEnabled)
if err != nil {
return bucketConfigState{}, err
}
if at.IsZero() {
at = created
}
valid := !created.IsZero() && !at.Before(created)
modified := valid && at.After(created)
if bucketConfigUpdateOnly(file) && len(data) == 0 {
modified = false
}
return bucketConfigState{data: data, key: key, at: at, real: modified, valid: valid}, nil
}
func (s bucketConfigState) candidate() bool {
return s.valid && (s.real || len(s.data) != 0)
}
func compareBucketConfigStates(a, b bucketConfigState) int {
if a.valid != b.valid {
if a.valid {
return 1
}
return -1
}
if a.real != b.real {
if a.real {
return 1
}
return -1
}
if a.real {
if n := a.at.Compare(b.at); n != 0 {
return n
}
// At equal source time a real deletion wins, preventing resurrection.
if (len(a.data) == 0) != (len(b.data) == 0) {
if len(a.data) == 0 {
return 1
}
return -1
}
}
return bytes.Compare(a.key, b.key)
}
func localBucketConfigUpdatedAt(meta BucketMetadata, file string, now time.Time) time.Time {
_, at := replicatedBucketConfig(&meta, file)
for _, lower := range []time.Time{meta.Created, *at} {
if !now.After(lower) {
now = lower.Add(time.Nanosecond)
}
}
return now.UTC()
}
func ensureBucketMetadataCreated(ctx context.Context, obj ObjectLayer, meta *BucketMetadata) error {
if !meta.Created.IsZero() {
return nil
}
info, err := obj.GetBucketInfo(ctx, meta.Name, BucketOptions{NoMetadata: true})
if err != nil {
return err
}
if info.Created.IsZero() {
return errors.New("bucket metadata creation time is unknown")
}
meta.Created = info.Created.UTC()
return nil
}
// applyBucketConfig runs under metadata.lock, on freshly loaded metadata. It
// changes only the selected field; the caller persists once after all checks.
func applyBucketConfig(meta *BucketMetadata, file string, data []byte, at time.Time) (bool, error) {
if bucketConfigUpdateOnly(file) && len(data) == 0 {
return false, nil
}
current, currentAt := replicatedBucketConfig(meta, file)
lockEnabled := len(meta.ObjectLockConfigXML) != 0
incoming, err := newBucketConfigState(meta.Name, file, data, at, meta.Created, lockEnabled)
if err != nil {
return false, err
}
if !incoming.candidate() {
return false, nil
}
local, err := newBucketConfigState(meta.Name, file, *current, *currentAt, meta.Created, lockEnabled)
if err != nil {
return false, err
}
if compareBucketConfigStates(incoming, local) <= 0 {
return false, nil
}
*current, *currentAt = bytes.Clone(incoming.data), incoming.at.UTC()
return true, nil
}
func rebaseBucketConfigDefaults(meta *BucketMetadata, oldCreated time.Time) {
if meta.Created.Equal(oldCreated) {
return
}
// These six fields alone use Created to distinguish a baseline from a
// tombstone. Preserve actual source times when adopting an existing bucket.
for _, file := range replicatedBucketConfigs {
_, at := replicatedBucketConfig(meta, file)
if at.IsZero() || at.Equal(oldCreated) {
*at = meta.Created
}
}
}
+160 -51
View File
@@ -24,6 +24,7 @@ import (
"fmt"
"math/rand"
"sync"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
@@ -125,19 +126,42 @@ func (sys *BucketMetadataSys) Set(bucket string, meta BucketMetadata) {
}
}
func (sys *BucketMetadataSys) updateAndParse(ctx context.Context, bucket string, configFile string, configData []byte, parse, lifecycleDelete bool) (updatedAt time.Time, err error) {
// bucketMetadataUpdate returns the committed snapshot to the caller. meta and
// updatedAt hold the saved state only when changed is true. Local writes always
// change state, because localBucketConfigUpdatedAt is strictly greater than the
// current field time, so their handlers can broadcast meta without rechecking.
type bucketMetadataUpdate struct {
meta BucketMetadata
updatedAt time.Time
changed bool
}
func (sys *BucketMetadataSys) updateAndParse(ctx context.Context, bucket, configFile string, configData []byte, parse, lifecycleDelete bool) (time.Time, error) {
result, err := sys.updateAndParseMetadata(ctx, bucket, configFile, configData, parse, lifecycleDelete, nil)
return result.updatedAt, err
}
func (sys *BucketMetadataSys) updateAndParseMetadata(ctx context.Context, bucket string, configFile string, configData []byte, parse, lifecycleDelete bool, sourceTime *time.Time) (result bucketMetadataUpdate, err error) {
objAPI := newObjectLayerFn()
if objAPI == nil {
return updatedAt, errServerNotInitialized
return result, errServerNotInitialized
}
if isMinioMetaBucketName(bucket) {
return updatedAt, errInvalidArgument
return result, errInvalidArgument
}
// Load deletions without parsed caches (notably quota), and compare the
// six replicated fields against the raw document under the same lock.
if isReplicatedBucketConfig(configFile) {
parse = false
if bucketConfigUpdateOnly(configFile) && len(configData) == 0 {
return result, nil
}
}
notifyCtx := ctx
ctx, unlock, err := lockBucketMetadata(ctx, objAPI, bucket)
if err != nil {
return updatedAt, err
return result, err
}
err = func() error {
@@ -157,55 +181,73 @@ func (sys *BucketMetadataSys) updateAndParse(ctx context.Context, bucket string,
return err
}
}
updatedAt = UTCNow()
switch configFile {
case bucketPolicyConfig:
meta.PolicyConfigJSON = configData
meta.PolicyConfigUpdatedAt = updatedAt
case bucketNotificationConfig:
meta.NotificationConfigXML = configData
meta.NotificationConfigUpdatedAt = updatedAt
case bucketLifecycleConfig:
meta.LifecycleConfigXML = configData
meta.LifecycleConfigUpdatedAt = updatedAt
case bucketSSEConfig:
meta.EncryptionConfigXML = configData
meta.EncryptionConfigUpdatedAt = updatedAt
case bucketTaggingConfig:
meta.TaggingConfigXML = configData
meta.TaggingConfigUpdatedAt = updatedAt
case bucketQuotaConfigFile:
meta.QuotaConfigJSON = configData
meta.QuotaConfigUpdatedAt = updatedAt
case objectLockConfig:
meta.ObjectLockConfigXML = configData
meta.ObjectLockConfigUpdatedAt = updatedAt
case bucketVersioningConfig:
meta.VersioningConfigXML = configData
meta.VersioningConfigUpdatedAt = updatedAt
case bucketReplicationConfig:
meta.ReplicationConfigXML = configData
meta.ReplicationConfigUpdatedAt = updatedAt
case bucketTargetsFile:
meta.BucketTargetsConfigJSON, meta.BucketTargetsConfigMetaJSON, err = encryptBucketMetadata(ctx, meta.Name, configData, kms.Context{
bucket: meta.Name,
bucketTargetsFile: bucketTargetsFile,
})
if err != nil {
return fmt.Errorf("Error encrypting bucket target metadata %w", err)
updatedAt := UTCNow()
if isReplicatedBucketConfig(configFile) {
if err := ensureBucketMetadataCreated(ctx, objAPI, &meta); err != nil {
var at time.Time
if sourceTime != nil {
at = *sourceTime
}
logBucketConfigReplication(ctx, bucket, configFile, "indeterminate", at, meta.Created, err.Error())
return err
}
if sourceTime == nil || sourceTime.IsZero() {
updatedAt = localBucketConfigUpdatedAt(meta, configFile, updatedAt)
if sourceTime != nil {
logBucketConfigReplication(ctx, bucket, configFile, "legacy-zero", *sourceTime, meta.Created, "assigned local source time")
}
} else {
updatedAt = sourceTime.UTC()
}
if updatedAt.Before(meta.Created) {
logBucketConfigReplication(ctx, bucket, configFile, "before-created", updatedAt, meta.Created, "peer event")
return nil
}
changed, err := applyBucketConfig(&meta, configFile, configData, updatedAt)
if err != nil {
return err
}
if !changed {
return nil
}
} else {
switch configFile {
case bucketNotificationConfig:
meta.NotificationConfigXML = configData
meta.NotificationConfigUpdatedAt = updatedAt
case bucketLifecycleConfig:
meta.LifecycleConfigXML = configData
meta.LifecycleConfigUpdatedAt = updatedAt
case bucketReplicationConfig:
meta.ReplicationConfigXML = configData
meta.ReplicationConfigUpdatedAt = updatedAt
case bucketTargetsFile:
meta.BucketTargetsConfigJSON, meta.BucketTargetsConfigMetaJSON, err = encryptBucketMetadata(ctx, meta.Name, configData, kms.Context{
bucket: meta.Name,
bucketTargetsFile: bucketTargetsFile,
})
if err != nil {
return fmt.Errorf("Error encrypting bucket target metadata %w", err)
}
meta.BucketTargetsConfigUpdatedAt = updatedAt
meta.BucketTargetsConfigMetaUpdatedAt = updatedAt
default:
return fmt.Errorf("Unknown bucket %s metadata update requested %s", bucket, configFile)
}
meta.BucketTargetsConfigUpdatedAt = updatedAt
meta.BucketTargetsConfigMetaUpdatedAt = updatedAt
default:
return fmt.Errorf("Unknown bucket %s metadata update requested %s", bucket, configFile)
}
return sys.saveMetadata(ctx, objAPI, meta)
if err := sys.saveMetadata(ctx, objAPI, &meta); err != nil {
return err
}
result = bucketMetadataUpdate{meta: meta, updatedAt: updatedAt, changed: true}
return nil
}()
if err != nil {
return updatedAt, err
return result, err
}
globalNotificationSys.LoadBucketMetadata(bgContext(notifyCtx), bucket) // Do not use caller context here
return updatedAt, nil
if result.changed {
globalNotificationSys.LoadBucketMetadata(bgContext(notifyCtx), bucket)
}
return result, nil
}
func (sys *BucketMetadataSys) save(ctx context.Context, meta BucketMetadata) error {
@@ -218,7 +260,7 @@ func (sys *BucketMetadataSys) save(ctx context.Context, meta BucketMetadata) err
return errInvalidArgument
}
if err := sys.saveMetadata(ctx, objAPI, meta); err != nil {
if err := sys.saveMetadata(ctx, objAPI, &meta); err != nil {
return err
}
@@ -228,11 +270,16 @@ func (sys *BucketMetadataSys) save(ctx context.Context, meta BucketMetadata) err
// saveMetadata persists and publishes metadata locally. Callers performing a
// read-modify-write must hold metadata.lock and release it before peer fan-out.
func (sys *BucketMetadataSys) saveMetadata(ctx context.Context, objAPI ObjectLayer, meta BucketMetadata) error {
func (sys *BucketMetadataSys) saveMetadata(ctx context.Context, objAPI ObjectLayer, meta *BucketMetadata) error {
// A writer may have queued for metadata.lock before DeleteBucket completed.
// Recheck the physical bucket under that lock, before recreating metadata.
if _, err := objAPI.GetBucketInfo(ctx, meta.Name, BucketOptions{NoMetadata: true}); err != nil {
return err
}
if err := meta.Save(ctx, objAPI); err != nil {
return err
}
sys.Set(meta.Name, meta)
sys.Set(meta.Name, *meta)
return nil
}
@@ -240,8 +287,19 @@ func lockBucketMetadata(ctx context.Context, objectAPI ObjectLayer, bucket strin
return lockBucketMetadataWithTimeout(ctx, objectAPI, bucket, globalOperationTimeout)
}
// lockBucketMetadataAcquireHook, when set, is invoked at the start of every
// metadata.lock acquisition, immediately before the blocking Lock() call. It is
// nil in production (a single atomic load, no behavior change) and exists only
// so tests can deterministically observe a caller reaching the metadata lock —
// notably DeleteBucket, whose lock is taken through its erasureServerPools
// receiver and is therefore invisible to an injected object layer.
var lockBucketMetadataAcquireHook atomic.Pointer[func(bucket string)]
func lockBucketMetadataWithTimeout(ctx context.Context, objectAPI ObjectLayer, bucket string, timeout *dynamicTimeout) (context.Context, func(), error) {
lock := objectAPI.NewNSLock(minioMetaBucket, pathJoin(bucketMetaPrefix, bucket, "metadata.lock"))
if hook := lockBucketMetadataAcquireHook.Load(); hook != nil {
(*hook)(bucket)
}
lkctx, err := lock.GetLock(ctx, timeout)
if err != nil {
return nil, nil, err
@@ -292,6 +350,57 @@ func (sys *BucketMetadataSys) Update(ctx context.Context, bucket string, configF
return sys.updateAndParse(ctx, bucket, configFile, configData, true, false)
}
// UpdateExpiryLCConfig merges a replicated ILM expiry configuration with the
// bucket's current lifecycle document and persists the merged result while
// holding metadata.lock across the read, merge, and save. The site-replication
// expiry heal and peer-apply paths must use this instead of computing the merge
// from an unlocked GetConfigFromDisk read and then writing it with Update: that
// two-step sequence drops any lifecycle transition change committed in between
// (issue #105). Lock order stays <bucket>.lck -> metadata.lock -> .metadata.bin;
// the merge and save run under metadata.lock and the peer fan-out runs after it
// is released.
func (sys *BucketMetadataSys) UpdateExpiryLCConfig(ctx context.Context, bucket string, expLCConfig *string, updatedAt time.Time) error {
objAPI := newObjectLayerFn()
if objAPI == nil {
return errServerNotInitialized
}
if isMinioMetaBucketName(bucket) {
return errInvalidArgument
}
notifyCtx := ctx
ctx, unlock, err := lockBucketMetadata(ctx, objAPI, bucket)
if err != nil {
return err
}
err = func() error {
defer unlock()
meta, err := loadBucketMetadataParse(ctx, objAPI, bucket, true)
if err != nil {
if !globalIsErasure && !globalIsDistErasure && errors.Is(err, errVolumeNotFound) {
// Only single drive mode needs this fallback.
meta = newBucketMetadata(bucket)
} else {
return err
}
}
configData, err := mergeExpiryWithLCConfig(bucket, meta, expLCConfig, updatedAt)
if err != nil {
return err
}
meta.LifecycleConfigXML = configData
meta.LifecycleConfigUpdatedAt = UTCNow()
return sys.saveMetadata(ctx, objAPI, &meta)
}()
if err != nil {
return err
}
globalNotificationSys.LoadBucketMetadata(bgContext(notifyCtx), bucket) // Do not use caller context here
return nil
}
// Get metadata for a bucket.
// If no metadata exists errConfigNotFound is returned and a new metadata is returned.
// Only a shallow copy is returned, so referenced data should not be modified,
+4 -9
View File
@@ -303,6 +303,9 @@ func loadBucketMetadataParseUnderLock(ctx context.Context, objectAPI ObjectLayer
return newBucketMetadata(bucket), fmt.Errorf("%w: %v", errBucketMetadataMigrationLockUnavailable, err)
}
defer unlock()
if _, err := objectAPI.GetBucketInfo(ctx, bucket, BucketOptions{NoMetadata: true}); err != nil {
return newBucketMetadata(bucket), err
}
return loadBucketMetadataParse(ctx, objectAPI, bucket, parse)
}
@@ -385,15 +388,7 @@ func (b *BucketMetadata) parseAllConfigs(ctx context.Context, objectAPI ObjectLa
} else {
b.objectLockConfig = nil
}
if b.objectLockConfig != nil {
// Object Lock requires every object to be versioned. Whatever the lock
// document contains, a suspended or prefix-excluded versioning document
// is replaced by plain Enabled versioning; Save persists the result.
config, versioningErr := versioning.ParseConfig(bytes.NewReader(b.VersioningConfigXML))
if versioningErr != nil || !config.Enabled() || config.PrefixesExcluded() {
b.VersioningConfigXML = enabledBucketVersioningConfig
}
}
b.VersioningConfigXML = effectiveBucketVersioning(b.VersioningConfigXML, b.objectLockConfig != nil)
if len(b.VersioningConfigXML) != 0 {
b.versioningConfig, err = versioning.ParseConfig(bytes.NewReader(b.VersioningConfigXML))
+125 -10
View File
@@ -25,6 +25,7 @@ import (
"strings"
"time"
"github.com/minio/minio/internal/amztime"
"github.com/minio/minio/internal/auth"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
xhttp "github.com/minio/minio/internal/http"
@@ -333,11 +334,10 @@ func checkPutObjectLockAllowed(ctx context.Context, rq *http.Request, bucket, ob
return mode, retainDate, legalHold, ErrObjectLocked
}
if !legalHoldRequested && retentionCfg.LockEnabled {
// inherit retention from bucket configuration
return retentionCfg.Mode, objectlock.RetentionDate{Time: t.Add(retentionCfg.Validity)}, legalHold, ErrNone
}
return "", objectlock.RetentionDate{}, legalHold, ErrNone
// Inherit retention from the bucket configuration. A legal-hold header
// on the same request, ON or OFF, is independent of retention and must
// not suppress the default (#165).
return retentionCfg.Mode, objectlock.RetentionDate{Time: t.Add(retentionCfg.Validity)}, legalHold, ErrNone
}
return mode, retainDate, legalHold, ErrNone
}
@@ -386,22 +386,137 @@ func (s objectLockState) legalHoldIsOlderThan(src time.Time) bool {
// restoreRetention and restoreLegalHold put the stored state back into
// metadata that was rebuilt from a request whose update was not applied.
func (s objectLockState) restoreRetention(metadata map[string]string) {
// The stored timestamp orders the next update and must survive even when
// the stored value is empty, which is how a removal is recorded.
if s.retentionTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = s.retentionTimestamp
}
if s.mode == "" {
return
}
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = s.mode
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = s.retainUntil
if s.retentionTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = s.retentionTimestamp
}
}
func (s objectLockState) restoreLegalHold(metadata map[string]string) {
if s.legalHoldTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = s.legalHoldTimestamp
}
if s.legalHold == "" {
return
}
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = s.legalHold
if s.legalHoldTimestamp != "" {
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = s.legalHoldTimestamp
}
// replicaStoredLock reads the Object Lock state stored on the addressed version
// so a trusted replica write can order its update against it. A missing object
// or version yields an empty state, which is correct for the first write of a
// version; any other read error is returned so the caller fails the write rather
// than ordering an incoming update against lock state it merely failed to read
// (an older incoming value must not win over a newer stored one just because the
// read timed out).
func replicaStoredLock(ctx context.Context, getObjectInfo GetObjectInfoFn, bucket, object, versionID string) (objectLockState, error) {
oi, err := getObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: versionID})
switch {
case err == nil:
return storedObjectLockState(oi.UserDefined), nil
case isErrObjectNotFound(err) || isErrVersionNotFound(err):
return objectLockState{}, nil
default:
return objectLockState{}, err
}
}
// applyReplicatedObjectLock writes the retention and legal-hold decision into
// metadata for a PUT, CopyObject, or multipart-initiation request. A request
// that is not an actual trusted replica -- a normal user write, or a trusted
// peer that carried the replication marker without REPLICA status -- takes
// ordinary write semantics: a validated value is applied and stamped now, and a
// missing value is left as is. Only an actual replica update is ordered against
// the state already stored on the addressed version, so a stale value cannot
// overwrite a newer one and a full retransmit cannot roll a destination back.
// The stored argument is meaningful only for a replica; callers pass an empty
// state otherwise. Only the two Object Lock keys and their reserved ordering
// timestamps are touched; any encryption-metadata reconciliation stays with the
// caller.
func applyReplicatedObjectLock(metadata map[string]string, stored objectLockState,
replicaTrusted bool,
retentionMode objectlock.RetMode, retentionDate objectlock.RetentionDate,
legalHold objectlock.ObjectLegalHold, srcRetentionTimestamp, srcLegalholdTimestamp time.Time,
) {
switch {
case !replicaTrusted:
// Ordinary write semantics: apply a validated retention and stamp it now;
// a missing value carries no instruction, so leave the metadata as it is.
if retentionMode.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
case !stored.retentionIsOlderThan(srcRetentionTimestamp):
// The stored update is at least as new as this replica's, or the replica
// carries no ordering timestamp: keep what is stored. This is also how a
// stale retransmit is rejected.
stored.restoreRetention(metadata)
default:
// The replica update wins. A removal carries no value but still records
// the source timestamp that orders it.
if retentionMode.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
}
metadata[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = srcRetentionTimestamp.UTC().Format(time.RFC3339Nano)
}
// Legal hold has no removal in S3: an explicitly empty header is already
// rejected as an invalid status, so the only value-less shape that gets here
// is an absent one, which conveys no legal-hold change. Only a valid status
// can win.
switch {
case !replicaTrusted:
if legalHold.Status.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
case legalHold.Status.Valid() && stored.legalHoldIsOlderThan(srcLegalholdTimestamp):
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
metadata[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = srcLegalholdTimestamp.UTC().Format(time.RFC3339Nano)
default:
stored.restoreLegalHold(metadata)
}
}
// reconcileStoredObjectLock re-orders the Object Lock already written into
// metadata against the state currently stored on the destination version, both
// compared by their reserved ordering timestamps. It runs inside the object
// layer under the namespace write lock that guards the version replacement,
// after the destination version is read and before the new one is committed, so
// a replica update whose ordering was decided at handler time (or, for multipart,
// at initiation) cannot overwrite a newer lock update that reached the version in
// between. metadata already carries the incoming update with its source
// timestamps; a stored value that is not older than the incoming one is put back,
// which for a stored removal means clearing the incoming value and keeping only
// the removal's timestamp. Only the two lock keys and their reserved timestamps
// move; a non-replica write never sets the flag that invokes this.
func reconcileStoredObjectLock(metadata map[string]string, stored objectLockState) {
incoming := storedObjectLockState(metadata)
incomingRetentionTS, _ := time.Parse(time.RFC3339Nano, incoming.retentionTimestamp)
if !stored.retentionIsOlderThan(incomingRetentionTS) {
// The stored retention is at least as new as the incoming one (or the
// incoming update is unordered): drop the incoming value and put the stored
// state back, which may itself be a removal (value keys absent, timestamp
// present).
delete(metadata, strings.ToLower(xhttp.AmzObjectLockMode))
delete(metadata, strings.ToLower(xhttp.AmzObjectLockRetainUntilDate))
delete(metadata, ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp)
stored.restoreRetention(metadata)
}
incomingLegalHoldTS, _ := time.Parse(time.RFC3339Nano, incoming.legalHoldTimestamp)
if incoming.legalHold == "" || !stored.legalHoldIsOlderThan(incomingLegalHoldTS) {
delete(metadata, strings.ToLower(xhttp.AmzObjectLockLegalHold))
delete(metadata, ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp)
stored.restoreLegalHold(metadata)
}
}
+5 -6
View File
@@ -19,7 +19,6 @@ package cmd
import (
"bytes"
"encoding/json"
"io"
"net/http"
@@ -100,13 +99,13 @@ func (api objectAPIHandlers) PutBucketPolicyHandler(w http.ResponseWriter, r *ht
return
}
configData, err := json.Marshal(bucketPolicy)
configData, err := canonicalBucketPolicy(bucketPolicy)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
updatedAt, err := globalBucketMetadataSys.Update(ctx, bucket, bucketPolicyConfig, configData)
result, err := globalBucketMetadataSys.updateAndParseMetadata(ctx, bucket, bucketPolicyConfig, configData, false, false, nil)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -116,8 +115,8 @@ func (api objectAPIHandlers) PutBucketPolicyHandler(w http.ResponseWriter, r *ht
replLogIf(ctx, globalSiteReplicationSys.BucketMetaHook(ctx, madmin.SRBucketMeta{
Type: madmin.SRBucketMetaTypePolicy,
Bucket: bucket,
Policy: bucketPolicyBytes,
UpdatedAt: updatedAt,
Policy: result.meta.PolicyConfigJSON,
UpdatedAt: result.updatedAt,
}))
// Success.
@@ -200,7 +199,7 @@ func (api objectAPIHandlers) GetBucketPolicyHandler(w http.ResponseWriter, r *ht
return
}
configData, err := json.Marshal(config)
configData, err := canonicalBucketPolicy(config)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
+26 -7
View File
@@ -255,13 +255,32 @@ func getConditionValuesWithTags(r *http.Request, lc string, cred auth.Credential
}
cloneHeader := r.Header.Clone()
signatureAge := cloneHeader.Get("x-amz-signature-age")
cloneHeader.Del("x-amz-signature-age")
// The presigned V4 verifier overwrites this internal scratch header after
// validating the signature. Ignore a value supplied on every other request
// type, where it would otherwise synthesize s3:signatureAge.
if authType == authTypePresigned && signatureAge != "" {
args["signatureAge"] = []string{signatureAge}
// s3:signatureAge is derived from the presigned X-Amz-Date rather than from
// anything the verifier writes back: PutObject and UploadPart authorize
// before they verify the signature, so a post-verification value is not yet
// available on the first evaluation. The date is bound by the signature
// (doesPresignedSignatureMatch rebuilds and compares it), so a forged date
// only changes the authorization outcome of a request that then fails
// verification. A date that does not parse leaves the key absent; the
// verifier rejects the request as ErrMalformedPresignedDate.
if authType == authTypePresigned {
if signedDate, err := time.Parse(iso8601Format, r.Form.Get(xhttp.AmzDate)); err == nil {
args["signatureAge"] = []string{strconv.FormatInt(currTime.Sub(signedDate).Milliseconds(), 10)}
}
}
// s3:x-amz-content-sha256 must name the payload hash the request is actually
// verified and enforced against, and only one such value. Presence of the
// header controls whether the key exists at all (AWS documents that the
// query-string form does not populate it), but the value comes from the same
// selection getContentSha256Cksum makes for verification: the presigned query
// value takes precedence over the header, and a repeated header contributes
// only its first value. Exposing every raw header value instead let a
// request satisfy a policy with a value the verifier never checked.
if _, ok := cloneHeader[xhttp.AmzContentSha256]; ok {
args[xhttp.AmzContentSha256] = []string{getContentSha256Cksum(r, serviceS3)}
cloneHeader.Del(xhttp.AmzContentSha256)
}
userTags := cloneHeader.Get(xhttp.AmzObjectTagging)
+34 -6
View File
@@ -23,8 +23,10 @@ import (
"net/url"
"os"
"slices"
"strconv"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/handlers"
@@ -579,8 +581,20 @@ func TestGetConditionValuesRejectsAbsentInternalKeys(t *testing.T) {
}
}
// s3:signatureAge is derived from the presigned X-Amz-Date, which the signature
// binds. A client header under the former scratch name must never supply it on
// any auth type, and a presign whose date is missing or malformed leaves the
// key absent (the verifier then rejects the request).
func TestGetConditionValuesOnlyAcceptsPresignedSignatureAge(t *testing.T) {
const signatureAgeHeader = "x-amz-signature-age"
signedDate := UTCNow().Add(-90 * time.Second)
presignQuery := func(date string) string {
q := url.Values{xhttp.AmzCredential: {"access/20260803/us-east-1/s3/aws4_request"}}
if date != "" {
q.Set(xhttp.AmzDate, date)
}
return "http://minio.local/bkt/obj?" + q.Encode()
}
for _, tc := range []struct {
name string
@@ -602,19 +616,33 @@ func TestGetConditionValuesOnlyAcceptsPresignedSignatureAge(t *testing.T) {
},
},
{
name: "presigned verifier value",
target: "http://minio.local/bkt/obj?" + url.Values{
xhttp.AmzCredential: {"access/20260803/us-east-1/s3/aws4_request"},
}.Encode(),
name: "presigned client header without date",
target: presignQuery(""),
headers: map[string]string{signatureAgeHeader: "250"},
},
{
name: "presigned malformed date",
target: presignQuery("yesterday"),
},
{
name: "presigned signed date",
target: presignQuery(signedDate.Format(iso8601Format)),
headers: map[string]string{signatureAgeHeader: "250"},
want: true,
},
} {
t.Run(tc.name, func(t *testing.T) {
got := condValuesForRequest(t, tc.target, tc.headers)
_, ok := got["signatureAge"]
v, ok := got["signatureAge"]
if ok != tc.want {
t.Fatalf("signatureAge presence: expected %v, got %v", tc.want, got["signatureAge"])
t.Fatalf("signatureAge presence: expected %v, got %v", tc.want, v)
}
if !tc.want {
return
}
age, err := strconv.ParseInt(strings.Join(v, ""), 10, 64)
if err != nil || age < (90*time.Second).Milliseconds() || age > (2*time.Minute).Milliseconds() {
t.Fatalf("signatureAge = %v, want about 90s derived from X-Amz-Date rather than the client header", v)
}
})
}
+12 -8
View File
@@ -43,6 +43,17 @@ func NewBucketQuotaSys() *BucketQuotaSys {
return &BucketQuotaSys{}
}
// getBucketQuotaSize returns the effective enforced hard-quota size.
func getBucketQuotaSize(quota *madmin.BucketQuota) uint64 {
if quota == nil || quota.Type != madmin.HardQuota {
return 0
}
if quota.Size > 0 {
return quota.Size
}
return quota.Quota
}
var bucketStorageCache = cachevalue.New[DataUsageInfo]()
// Init initialize bucket quota.
@@ -110,14 +121,7 @@ func (sys *BucketQuotaSys) enforceQuotaHard(ctx context.Context, bucket string,
return err
}
var quotaSize uint64
if q != nil && q.Type == madmin.HardQuota {
if q.Size > 0 {
quotaSize = q.Size
} else if q.Quota > 0 {
quotaSize = q.Quota
}
}
quotaSize := getBucketQuotaSize(q)
if quotaSize > 0 {
if uint64(size) >= quotaSize { // check if file size already exceeds the quota
return BucketQuotaExceeded{Bucket: bucket}
+77
View File
@@ -0,0 +1,77 @@
// Copyright (c) 2015-2025 MinIO, Inc.
// Copyright (c) 2025-2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful,
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"testing"
"github.com/minio/madmin-go/v3"
)
func TestGetBucketQuotaSize(t *testing.T) {
tests := []struct {
name string
quota *madmin.BucketQuota
want uint64
}{
{name: "nil"},
{name: "empty", quota: &madmin.BucketQuota{}},
{name: "current size", quota: &madmin.BucketQuota{Type: madmin.HardQuota, Size: 1024}, want: 1024},
{name: "legacy quota", quota: &madmin.BucketQuota{Type: madmin.HardQuota, Quota: 2048}, want: 2048},
{name: "size takes precedence", quota: &madmin.BucketQuota{Type: madmin.HardQuota, Size: 1024, Quota: 2048}, want: 1024},
{name: "missing type", quota: &madmin.BucketQuota{Size: 1024}},
{name: "unsupported type", quota: &madmin.BucketQuota{Type: "fifo", Size: 1024}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := getBucketQuotaSize(tt.quota); got != tt.want {
t.Fatalf("getBucketQuotaSize() = %d, want %d", got, tt.want)
}
})
}
}
func TestIsBktQuotaCfgReplicated(t *testing.T) {
hardQuota := func(size, legacy uint64) *madmin.BucketQuota {
return &madmin.BucketQuota{Type: madmin.HardQuota, Size: size, Quota: legacy}
}
tests := []struct {
name string
quotas []*madmin.BucketQuota
want bool
}{
{name: "none configured", quotas: []*madmin.BucketQuota{nil, nil}, want: true},
{name: "missing from one site", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), nil}},
{name: "matching size", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), hardQuota(1024, 0)}, want: true},
{name: "different size", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), hardQuota(2048, 0)}},
{name: "equivalent representations", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), hardQuota(0, 1024)}, want: true},
{name: "different typeless size", quotas: []*madmin.BucketQuota{{Size: 1024}, {Size: 2048}}},
{name: "different type", quotas: []*madmin.BucketQuota{hardQuota(1024, 0), {Type: "fifo", Size: 1024}}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := isBktQuotaCfgReplicated(len(tt.quotas), tt.quotas); got != tt.want {
t.Fatalf("isBktQuotaCfgReplicated() = %v, want %v", got, tt.want)
}
})
}
}
+4 -4
View File
@@ -640,10 +640,10 @@ type VersionPurgeStatusType = replication.VersionPurgeStatusType
type replicationResyncer struct {
// map of bucket to their resync status
statusMap map[string]BucketReplicationResyncStatus
workerSize int
resyncCancelCh chan struct{}
workerCh chan struct{}
statusMap map[string]BucketReplicationResyncStatus
workerSize int
cancelResyncs map[resyncOpts]context.CancelCauseFunc
workerCh chan struct{}
sync.RWMutex
}
+454 -113
View File
@@ -418,7 +418,12 @@ func checkReplicateDelete(ctx context.Context, bucket string, dobj ObjectToDelet
// target cluster, the object version is marked deleted on the source and hidden from listing. It is permanently
// deleted from the source when the VersionPurgeStatus changes to "Complete", i.e after replication succeeds
// on target.
func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, objectAPI ObjectLayer) {
// replicateDelete replicates a delete (delete marker or version purge) to all
// applicable targets and returns the per-target replication outcome. Callers
// that only trigger replication may ignore the return value; the resync path
// uses it to classify success/failure per target rather than inferring it from
// the mere presence or absence of the target version.
func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, objectAPI ObjectLayer) replicatedInfos {
var replicationStatus replication.StatusType
bucket := dobj.Bucket
versionID := dobj.DeleteMarkerVersionID
@@ -453,7 +458,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
Host: globalLocalNodeName,
EventName: event.ObjectReplicationNotTracked,
})
return
return replicatedInfos{}
}
dsc, err := parseReplicateDecision(ctx, bucket, dobj.ReplicationState.ReplicateDecisionStr)
if err != nil {
@@ -471,7 +476,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
Host: globalLocalNodeName,
EventName: event.ObjectReplicationNotTracked,
})
return
return replicatedInfos{}
}
// Lock the object name before starting replication operation.
@@ -492,7 +497,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
Host: globalLocalNodeName,
EventName: event.ObjectReplicationNotTracked,
})
return
return replicatedInfos{}
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
@@ -597,6 +602,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
EventName: eventName,
})
}
return rinfos
}
func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationInfo, tgt *TargetClient) (rinfo replicatedTargetInfo) {
@@ -779,6 +785,15 @@ func putReplicationOpts(ctx context.Context, sc string, objInfo ObjectInfo) (put
meta := make(map[string]string)
isSSEC := crypto.SSEC.IsEncrypted(objInfo.UserDefined)
// An SSE-C object is replicated as raw ciphertext, and the replication
// headers carry no compression state. Sending a compressed SSE-C object
// would land a replica that decrypts to an S2 stream instead of the
// object, so fail loudly instead of writing a wrong replica.
if isSSEC && objInfo.IsCompressed() {
return putOpts, false, fmt.Errorf("replication of a compressed SSE-C object is not supported: %s/%s(%s)",
objInfo.Bucket, objInfo.Name, objInfo.VersionID)
}
for k, v := range objInfo.UserDefined {
_, isValidSSEHeader := validSSEReplicationHeaders[k]
// In case of SSE-C objects copy the allowed internal headers as well
@@ -857,19 +872,29 @@ func putReplicationOpts(ctx context.Context, sc string, objInfo ObjectInfo) (put
if cc, ok := lkMap.Lookup(xhttp.CacheControl); ok {
putOpts.CacheControl = cc
}
if mode, ok := lkMap.Lookup(xhttp.AmzObjectLockMode); ok {
rmode := minio.RetentionMode(mode)
putOpts.Mode = rmode
mode, hasMode := lkMap.Lookup(xhttp.AmzObjectLockMode)
retainDateStr, hasRetainDate := lkMap.Lookup(xhttp.AmzObjectLockRetainUntilDate)
if hasMode {
putOpts.Mode = minio.RetentionMode(mode)
}
if retainDateStr, ok := lkMap.Lookup(xhttp.AmzObjectLockRetainUntilDate); ok {
// A removed retention is stored as an empty or absent mode and date; it is
// sent as a value-less update that still carries its ordering timestamp.
if hasRetainDate && retainDateStr != "" {
rdate, err := amztime.ISO8601Parse(retainDateStr)
if err != nil {
return putOpts, false, err
}
putOpts.RetainUntilDate = rdate
// set retention timestamp in opts
}
retainTmstampStr, hasRetainTmstamp := objInfo.UserDefined[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp]
if hasMode || hasRetainDate || hasRetainTmstamp {
// Send the ordering timestamp whenever the version carries one, even for a
// removal whose value keys are absent (the shape a retransmit PUT leaves),
// so the next hop can order the removal instead of keeping obsolete
// retention.
retTimestamp := objInfo.ModTime
if retainTmstampStr, ok := objInfo.UserDefined[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp]; ok {
if hasRetainTmstamp {
var err error
retTimestamp, err = time.Parse(time.RFC3339Nano, retainTmstampStr)
if err != nil {
return putOpts, false, err
@@ -931,11 +956,17 @@ func equals(k1 string, keys ...string) bool {
return false
}
// nullVersionExcludedFromResync reports the exclusion at the head of getReplicationAction, kept
// verbatim from upstream: an existing object resync leaves a null version alone when the source
// modification time is later than the one the target reports, without comparing anything else.
func nullVersionExcludedFromResync(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replication.Type) bool {
return opType == replication.ExistingObjectReplicationType &&
oi1.ModTime.Unix() > oi2.LastModified.Unix() && oi1.VersionID == nullVersionID
}
// returns replicationAction by comparing metadata between source and target
func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replication.Type) replicationAction {
// Avoid resyncing null versions created prior to enabling replication if target has a newer copy
if opType == replication.ExistingObjectReplicationType &&
oi1.ModTime.Unix() > oi2.LastModified.Unix() && oi1.VersionID == nullVersionID {
if nullVersionExcludedFromResync(oi1, oi2, opType) {
return replicateNone
}
sz, _ := oi1.GetActualSize()
@@ -986,9 +1017,21 @@ func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replicati
"X-Amz-Meta-",
}
// An empty object lock mode or retain-until-date records a removed retention, but
// it is omitted from GET/HEAD response headers: setObjectHeaders() skips both keys
// when the value is empty, and FilterObjectLockMetadata() drops them when the mode
// is not valid. The target can therefore never report them, so treat empty and
// absent as equal rather than as a permanent difference.
emptyLockValue := func(k, v string) bool {
return v == "" && equals(k, xhttp.AmzObjectLockMode, xhttp.AmzObjectLockRetainUntilDate)
}
// compare metadata on both maps to see if meta is identical
compareMeta1 := make(map[string]string)
for k, v := range oi1.UserDefined {
if emptyLockValue(k, v) {
continue
}
var found bool
for _, prefix := range compareKeys {
if !stringsHasPrefixFold(k, prefix) {
@@ -1004,6 +1047,10 @@ func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replicati
compareMeta2 := make(map[string]string)
for k, v := range oi2.Metadata {
val := strings.Join(v, ",")
if emptyLockValue(k, val) {
continue
}
var found bool
for _, prefix := range compareKeys {
if !stringsHasPrefixFold(k, prefix) {
@@ -1013,7 +1060,7 @@ func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replicati
break
}
if found {
compareMeta2[strings.ToLower(k)] = strings.Join(v, ",")
compareMeta2[strings.ToLower(k)] = val
}
}
@@ -1024,9 +1071,87 @@ func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replicati
return replicateNone
}
// objectRetentionGetter is the part of the replication target client used to confirm whether a
// destination version still holds Object Lock retention.
type objectRetentionGetter interface {
GetObjectRetention(ctx context.Context, bucketName, objectName, versionID string) (*minio.RetentionMode, *time.Time, error)
}
// retentionRemovedAtSource reports whether oi carries the shape a removed retention leaves behind.
// Two representations persist. A retention removed directly on this cluster keeps the object lock
// keys present with empty values (PutObjectRetentionHandler, cmd/object-handlers.go:3309-3316). A
// removal that arrived by replication keeps only the retention ordering timestamp, with the mode
// and retain-until-date keys absent, because restoreRetention and the replica update path write
// the timestamp alone when the mode is empty (cmd/bucket-object-lock.go:388-399,
// cmd/object-handlers.go:1782-1797). A present ordering timestamp paired with a non-empty mode is
// a retention that was set, not removed, and must not be mistaken for one.
func retentionRemovedAtSource(oi ObjectInfo) bool {
lkMap := caseInsensitiveMap(oi.UserDefined)
// Representation (1): an object lock key is present with an empty value.
for _, k := range []string{xhttp.AmzObjectLockMode, xhttp.AmzObjectLockRetainUntilDate} {
if v, ok := lkMap.Lookup(k); ok && v == "" {
return true
}
}
// Representation (2): a recorded retention ordering timestamp with the mode value absent or
// empty is a removal restoreRetention persisted without the empty public keys.
if _, ok := oi.UserDefined[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp]; ok {
if v, ok := lkMap.Lookup(xhttp.AmzObjectLockMode); !ok || v == "" {
return true
}
}
return false
}
// targetRetentionConfirmedAbsent reports whether the destination version is known to hold no
// retention. A HEAD response omits retention both when the version has none and when the
// replication credential lacks s3:GetObjectRetention (cmd/object-handlers.go:942-946), so the
// comparison in getReplicationAction on its own cannot tell a removal that is already in sync from
// one the destination still holds. Only NoSuchObjectLockConfiguration, the answer for a version
// that carries no retention, and a response naming no retention mode count as absent. Everything
// else is uncertainty and is treated as still present, so that the removal is resent exactly as it
// is today: a denied or unreachable destination, a mode the SDK returned without recognizing since
// it does not validate it, and InvalidRequest, which names a bucket with no Object Lock
// configuration but is also what a destination answers when its own read of that configuration
// fails (cmd/bucket-object-lock.go:39-50 returns an error with a zero Retention, discarded at
// cmd/object-handlers.go:3275).
func targetRetentionConfirmedAbsent(ctx context.Context, tgt objectRetentionGetter, bucket, object, versionID string) bool {
mode, _, err := tgt.GetObjectRetention(ctx, bucket, object, versionID)
if err != nil {
return minio.ToErrorResponse(err).Code == "NoSuchObjectLockConfiguration"
}
// An absent or empty mode is no retention. A non-empty mode is retention, whether or not this
// SDK recognizes it.
return mode == nil || *mode == ""
}
// replicationActionForTarget returns the action for a source version against a destination that
// answered HEAD. It is getReplicationAction plus the confirmation that a removed retention which
// compares as in sync really is: see targetRetentionConfirmedAbsent.
func replicationActionForTarget(ctx context.Context, oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replication.Type, tgt objectRetentionGetter, bucket, object string) replicationAction {
rAction := getReplicationAction(oi1, oi2, opType)
if rAction != replicateNone || !retentionRemovedAtSource(oi1) {
return rAction
}
// A null version the resync deliberately leaves alone is not a comparison result, so it is
// not the confirmation's to reopen.
if nullVersionExcludedFromResync(oi1, oi2, opType) {
return rAction
}
if targetRetentionConfirmedAbsent(ctx, tgt, bucket, object, oi1.VersionID) {
return rAction
}
return replicateMetadata
}
// replicateObject replicates the specified version of the object to destination bucket
// The source object is then updated to reflect the replication status.
func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI ObjectLayer) {
// replicateObject replicates a single object version to all applicable targets
// and returns the per-target replication outcome. Callers that only trigger
// replication may ignore the return value; the resync path uses it to classify
// success/failure per target rather than inferring it from the mere existence
// of the target version.
func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI ObjectLayer) replicatedInfos {
var replicationStatus replication.StatusType
defer func() {
if replicationStatus.Empty() {
@@ -1059,7 +1184,7 @@ func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI Obje
UserAgent: "Internal: [Replication]",
Host: globalLocalNodeName,
})
return
return replicatedInfos{}
}
tgtArns := cfg.FilterTargetArns(replication.ObjectOpts{
Name: object,
@@ -1079,7 +1204,7 @@ func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI Obje
Host: globalLocalNodeName,
})
globalReplicationPool.Get().queueMRFSave(ri.ToMRFEntry())
return
return replicatedInfos{}
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
@@ -1185,6 +1310,7 @@ func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI Obje
ri.RetryCount++
globalReplicationPool.Get().queueMRFSave(ri.ToMRFEntry())
}
return rinfos
}
// replicateObject replicates object data for specified version of the object to destination bucket
@@ -1299,6 +1425,8 @@ func (ri ReplicateObjectInfo) replicateObject(ctx context.Context, objectAPI Obj
putOpts, isMP, err := putReplicationOpts(ctx, tgt.StorageClass, objInfo)
if err != nil {
rinfo.Err = err
rinfo.ReplicationStatus = replication.Failed
replLogIf(ctx, fmt.Errorf("failure setting options for replication bucket:%s err:%w", bucket, err))
sendEvent(eventArgs{
EventName: event.ObjectReplicationNotTracked,
@@ -1467,7 +1595,7 @@ func (ri ReplicateObjectInfo) replicateAll(ctx context.Context, objectAPI Object
sOpts.Set(xhttp.AmzTagDirective, "ACCESS")
oi, cerr := tgt.StatObject(ctx, tgt.Bucket, object, sOpts)
if cerr == nil {
rAction = getReplicationAction(objInfo, oi, ri.OpType)
rAction = replicationActionForTarget(ctx, objInfo, oi, ri.OpType, tgt, tgt.Bucket, object)
rinfo.ReplicationStatus = replication.Completed
if rAction == replicateNone {
if ri.OpType == replication.ExistingObjectReplicationType &&
@@ -1497,11 +1625,14 @@ func (ri ReplicateObjectInfo) replicateAll(ctx context.Context, objectAPI Object
return rinfo
}
} else {
// SSEC objects will refuse HeadObject without the decryption key.
// Ignore the error, since we know the object exists and versioning prevents overwriting existing versions.
// The sender holds no customer key, so the target refuses HeadObject on
// an SSE-C object and the replica cannot be compared. The metadata-only
// CopyObject that a replicateMetadata action would run then fails on any
// non-empty object, because the undecryptable source checksum makes the
// target recompute one and rewrite the data. A full retransmit is the
// only action that completes.
if isSSEC && strings.Contains(cerr.Error(), errorCodes[ErrSSEEncryptedObject].Description) {
rinfo.ReplicationStatus = replication.Completed
rinfo.ReplicationAction = replicateNone
rAction = replicateAll
goto applyAction
}
// if target returns error other than NoSuchKey, defer replication attempt
@@ -1586,6 +1717,11 @@ applyAction:
} else {
putOpts, isMP, err := putReplicationOpts(ctx, tgt.StorageClass, objInfo)
if err != nil {
// rinfo was primed Completed above; a failure to build the write
// options means nothing reached the target, so mark it Failed and
// carry the error instead of reporting a phantom success.
rinfo.ReplicationStatus = replication.Failed
rinfo.Err = err
replLogIf(ctx, fmt.Errorf("failed to set replicate options for object %s/%s(%s) (target %s) err:%w", bucket, objInfo.Name, objInfo.VersionID, tgt.EndpointURL(), err))
sendEvent(eventArgs{
EventName: event.ObjectReplicationNotTracked,
@@ -2218,9 +2354,9 @@ func (p *ReplicationPool) queueReplicaTask(ri ReplicateObjectInfo) {
switch ri.OpType {
case replication.HealReplicationType, replication.ExistingObjectReplicationType:
ch = p.mrfReplicaCh
healCh = p.getWorkerCh(ri.Name, ri.Bucket, ri.Size)
healCh = p.getWorkerCh(ri.Bucket, ri.Name, ri.Size)
default:
ch = p.getWorkerCh(ri.Name, ri.Bucket, ri.Size)
ch = p.getWorkerCh(ri.Bucket, ri.Name, ri.Size)
}
if ch == nil && healCh == nil {
return
@@ -2828,10 +2964,10 @@ const (
func newresyncer() *replicationResyncer {
rs := replicationResyncer{
statusMap: make(map[string]BucketReplicationResyncStatus),
workerSize: resyncWorkerCnt,
resyncCancelCh: make(chan struct{}, resyncWorkerCnt),
workerCh: make(chan struct{}, resyncWorkerCnt),
statusMap: make(map[string]BucketReplicationResyncStatus),
workerSize: resyncWorkerCnt,
cancelResyncs: make(map[resyncOpts]context.CancelCauseFunc),
workerCh: make(chan struct{}, resyncWorkerCnt),
}
for i := 0; i < rs.workerSize; i++ {
rs.workerCh <- struct{}{}
@@ -2839,13 +2975,67 @@ func newresyncer() *replicationResyncer {
return &rs
}
var errResyncCanceled = errors.New("replication resync canceled")
// Registration and cancellation share the status lock. A queued run therefore
// cannot miss cancellation between publishing its status and taking a slot.
func (s *replicationResyncer) registerResync(parent context.Context, opts resyncOpts) (context.Context, context.CancelCauseFunc, bool) {
s.Lock()
defer s.Unlock()
st, ok := s.statusMap[opts.bucket].TargetsMap[opts.arn]
if !ok || st.ResyncID != opts.resyncID {
return nil, nil, false
}
if _, running := s.cancelResyncs[opts]; running {
return nil, nil, false
}
ctx, cancel := context.WithCancelCause(parent)
if s.cancelResyncs == nil {
s.cancelResyncs = make(map[resyncOpts]context.CancelCauseFunc)
}
s.cancelResyncs[opts] = cancel
if st.ResyncStatus == ResyncCanceled {
cancel(errResyncCanceled)
}
return ctx, cancel, true
}
func (s *replicationResyncer) cancelResyncID(resyncID string) {
s.Lock()
defer s.Unlock()
for bucket, m := range s.statusMap {
for arn, st := range m.TargetsMap {
if st.ResyncID == resyncID && (st.ResyncStatus == ResyncPending || st.ResyncStatus == ResyncStarted) {
st.ResyncStatus = ResyncCanceled
st.LastUpdate = UTCNow()
m.TargetsMap[arn] = st
m.LastUpdate = st.LastUpdate
}
}
s.statusMap[bucket] = m
}
for opts, cancel := range s.cancelResyncs {
if opts.resyncID == resyncID {
cancel(errResyncCanceled)
}
}
}
// mark status of replication resync on remote target for the bucket
func (s *replicationResyncer) markStatus(status ResyncStatusType, opts resyncOpts, objAPI ObjectLayer) {
func (s *replicationResyncer) markStatus(status ResyncStatusType, opts resyncOpts, objAPI ObjectLayer) ResyncStatusType {
s.Lock()
defer s.Unlock()
m := s.statusMap[opts.bucket]
st := m.TargetsMap[opts.arn]
st, ok := m.TargetsMap[opts.arn]
if !ok || st.ResyncID != opts.resyncID {
return NoResync
}
// A cancel may win the lock after the finalizer checked its context.
// Persist that cancellation, never a stale Started/Completed result.
if st.ResyncStatus == ResyncCanceled {
status = ResyncCanceled
}
st.LastUpdate = UTCNow()
st.ResyncStatus = status
m.TargetsMap[opts.arn] = st
@@ -2855,14 +3045,18 @@ func (s *replicationResyncer) markStatus(status ResyncStatusType, opts resyncOpt
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
saveResyncStatus(ctx, opts.bucket, m, objAPI)
return status
}
// update replication resync stats for bucket's remote target
func (s *replicationResyncer) incStats(ts TargetReplicationResyncStatus, opts resyncOpts) {
func (s *replicationResyncer) incStats(ts TargetReplicationResyncStatus, opts resyncOpts) bool {
s.Lock()
defer s.Unlock()
m := s.statusMap[opts.bucket]
st := m.TargetsMap[opts.arn]
st, ok := m.TargetsMap[opts.arn]
if !ok || st.ResyncID != opts.resyncID || st.ResyncStatus == ResyncCanceled {
return false
}
st.Object = ts.Object
st.ReplicatedCount += ts.ReplicatedCount
st.FailedCount += ts.FailedCount
@@ -2871,20 +3065,175 @@ func (s *replicationResyncer) incStats(ts TargetReplicationResyncStatus, opts re
m.TargetsMap[opts.arn] = st
m.LastUpdate = UTCNow()
s.statusMap[opts.bucket] = m
return true
}
// resyncResults consumes the per-object outcomes produced by the resync worker
// pool and applies each to the in-memory resync status via apply. It centralizes
// the finalization ordering so a status persisted after finish() returns always
// reflects every result.
type resyncResults struct {
ch chan TargetReplicationResyncStatus
apply func(TargetReplicationResyncStatus)
wg sync.WaitGroup
}
// newResyncResults starts the result-consuming goroutine that folds each worker
// result into the bucket's resync status.
func (s *replicationResyncer) newResyncResults(opts resyncOpts) *resyncResults {
return startResyncResults(func(r TargetReplicationResyncStatus) {
if s.incStats(r, opts) {
globalSiteResyncMetrics.updateMetric(r, opts.resyncID)
}
})
}
// startResyncResults starts a goroutine that applies every received result with
// apply. Injecting the apply action keeps the shutdown ordering in finish()
// testable.
func startResyncResults(apply func(TargetReplicationResyncStatus)) *resyncResults {
rr := &resyncResults{
ch: make(chan TargetReplicationResyncStatus, 1),
apply: apply,
}
rr.wg.Add(1)
go func() {
defer rr.wg.Done()
for r := range rr.ch {
rr.apply(r)
}
}()
return rr
}
// finish shuts the resync pipeline down in an order that guarantees a status
// persisted afterwards reflects every result. It first closes the worker input
// channels and waits for the producer workers to exit, so none can send on a
// closed result channel (a hazard on early-return paths) and every submitted
// result is delivered (a result a worker discards on cancellation is
// intentionally not); only then does it close the result channel and wait for
// the consumer to apply the last buffered result.
func (rr *resyncResults) finish(workers []chan ReplicateObjectInfo, workerWg *sync.WaitGroup) {
for i := range workers {
xioutil.SafeClose(workers[i])
}
workerWg.Wait()
xioutil.SafeClose(rr.ch)
rr.wg.Wait()
}
// sendResyncResult delivers a worker's computed per-object result to ch,
// returning false if the run's context was canceled first.
func (s *replicationResyncer) sendResyncResult(ctx context.Context, ch chan<- TargetReplicationResyncStatus, st TargetReplicationResyncStatus) bool {
select {
case <-ctx.Done():
return false
case ch <- st:
return true
}
}
// User cancellation is terminal; interrupted completion remains retryable.
func finalResyncStatus(status ResyncStatusType, cause error) ResyncStatusType {
if errors.Is(cause, errResyncCanceled) {
return ResyncCanceled
}
if status == ResyncCompleted && cause != nil {
return ResyncFailed
}
return status
}
// resyncTargetSucceeded reports whether this object (or delete) actually
// replicated to the target, from the target's own outcome. A version purge
// reports success through VersionPurgeStatus, not ReplicationStatus. For an
// object or delete marker, success requires a Completed status; a retained
// error is a real failure unless it is the benign duplicate 412 the
// destination returns when it already holds this exact ETag and version, which
// replicateAll deliberately keeps Completed.
func resyncTargetSucceeded(t replicatedTargetInfo, roi ReplicateObjectInfo) bool {
if !roi.VersionPurgeStatus.Empty() {
return t.VersionPurgeStatus == replication.VersionPurgeComplete
}
if t.ReplicationStatus != replication.Completed {
return false
}
return t.Err == nil || minio.ToErrorResponse(t.Err).Code == "PreconditionFailed"
}
// resyncResultFor derives the resync outcome for target arn from the aggregate
// replication result of a single object (or delete). The target counts as a
// success only when its own replication Completed without error - not when the
// target version merely exists. A target that Failed, errored, or was not
// attempted for this object (its arn absent from the result) counts as a
// failure, and the failed byte count is recorded (previously always zero).
func resyncResultFor(rinfos replicatedInfos, arn string, roi ReplicateObjectInfo) TargetReplicationResyncStatus {
st := TargetReplicationResyncStatus{Object: roi.Name, Bucket: roi.Bucket}
for _, t := range rinfos.Targets {
if t.Arn != arn {
continue
}
if resyncTargetSucceeded(t, roi) {
sz := t.Size
if sz == 0 {
sz = roi.Size
}
st.ReplicatedCount++
st.ReplicatedSize += sz
} else {
st.FailedCount++
st.FailedSize += roi.Size
}
return st
}
// arn was not attempted for this object: a resync that cannot confirm the
// object reached the target is not a success.
st.FailedCount++
st.FailedSize += roi.Size
return st
}
// objectNeedsResyncForARN reports whether roi must be resynced for target arn
// specifically. The resync worker pool is scoped to a single target (opts.arn),
// so an object that only qualifies for a different target must be skipped here:
// admitting it would replicate it for arn's peers only, leaving arn absent from
// the per-object result, which resyncResultFor then (correctly, but
// misleadingly) counts as a failure for arn - an object arn was never
// responsible for. Only opts.arn carries this resync's ResetID, and that reset
// is already folded into its per-target decision, so the per-target check both
// scopes dispatch and honors the reset.
func objectNeedsResyncForARN(roi ReplicateObjectInfo, arn string) bool {
return roi.ExistingObjResync.mustResyncTarget(arn)
}
// resyncBucket resyncs all qualifying objects as per replication rules for the target
// ARN
func (s *replicationResyncer) resyncBucket(ctx context.Context, objectAPI ObjectLayer, heal bool, opts resyncOpts) {
ctx, cancel, registered := s.registerResync(ctx, opts)
if !registered {
return
}
defer func() {
cancel(nil)
s.Lock()
delete(s.cancelResyncs, opts)
s.Unlock()
}()
select {
case <-s.workerCh: // block till a worker is available
case <-ctx.Done():
if errors.Is(context.Cause(ctx), errResyncCanceled) {
status := s.markStatus(ResyncCanceled, opts, objectAPI)
globalSiteResyncMetrics.incBucket(opts, status)
}
return
}
resyncStatus := ResyncFailed
defer func() {
s.markStatus(resyncStatus, opts, objectAPI)
// Runs after workers/results drain and before our own deferred cancel.
resyncStatus = finalResyncStatus(resyncStatus, context.Cause(ctx))
resyncStatus = s.markStatus(resyncStatus, opts, objectAPI)
globalSiteResyncMetrics.incBucket(opts, resyncStatus)
s.workerCh <- struct{}{}
}()
@@ -2919,8 +3268,8 @@ func (s *replicationResyncer) resyncBucket(ctx context.Context, objectAPI Object
return
}
// mark resync status as resync started
if !heal {
s.markStatus(ResyncStarted, opts, objectAPI)
if !heal && s.markStatus(ResyncStarted, opts, objectAPI) != ResyncStarted {
return
}
// Walk through all object versions - Walk() is always in ascending order needed to ensure
@@ -2939,16 +3288,21 @@ func (s *replicationResyncer) resyncBucket(ctx context.Context, objectAPI Object
lastCheckpoint = st.Object
}
workers := make([]chan ReplicateObjectInfo, resyncParallelRoutines)
resultCh := make(chan TargetReplicationResyncStatus, 1)
defer xioutil.SafeClose(resultCh)
go func() {
for r := range resultCh {
s.incStats(r, opts)
globalSiteResyncMetrics.updateMetric(r, opts.resyncID)
}
}()
var wg sync.WaitGroup
// results consumes each worker's per-object outcome and folds it into the
// in-memory status. finish() (deferred below) stops the workers and drains
// every result before the deferred markStatus persists, so a Completed status
// cannot race the last incStats. Registered after the markStatus finalizer, so
// LIFO runs finish first.
results := s.newResyncResults(opts)
defer func() {
if resyncStatus != ResyncCompleted {
cancel(nil)
}
// Exactly one finish: success drains with a live context; errors stop
// blocked workers first. Both drain counts before persisting status.
results.finish(workers, &wg)
}()
for i := range resyncParallelRoutines {
wg.Add(1)
workers[i] = make(chan ReplicateObjectInfo, 100)
@@ -2959,10 +3313,10 @@ func (s *replicationResyncer) resyncBucket(ctx context.Context, objectAPI Object
select {
case <-ctx.Done():
return
case <-s.resyncCancelCh:
default:
}
traceFn := s.trace(tgt.ResetID, fmt.Sprintf("%s/%s (%s)", opts.bucket, roi.Name, roi.VersionID))
var rinfos replicatedInfos
if roi.DeleteMarker || !roi.VersionPurgeStatus.Empty() {
versionID := ""
dmVersionID := ""
@@ -2985,83 +3339,69 @@ func (s *replicationResyncer) resyncBucket(ctx context.Context, objectAPI Object
OpType: replication.ExistingObjectReplicationType,
EventType: ReplicateExistingDelete,
}
replicateDelete(ctx, doi, objectAPI)
rinfos = replicateDelete(ctx, doi, objectAPI)
} else {
roi.OpType = replication.ExistingObjectReplicationType
roi.EventType = ReplicateExisting
replicateObject(ctx, roi, objectAPI)
rinfos = replicateObject(ctx, roi, objectAPI)
}
st := TargetReplicationResyncStatus{
Object: roi.Name,
Bucket: roi.Bucket,
}
_, err := tgt.StatObject(ctx, tgt.Bucket, roi.Name, minio.StatObjectOptions{
VersionID: roi.VersionID,
Internal: minio.AdvancedGetOptions{
ReplicationProxyRequest: "false",
},
})
sz := roi.Size
if err != nil {
if roi.DeleteMarker && isErrMethodNotAllowed(ErrorRespToObjectError(err, opts.bucket, roi.Name)) {
st.ReplicatedCount++
} else {
st.FailedCount++
// Classify success/failure from the actual replication outcome
// for this target, not from whether the target version merely
// exists (a rejected update leaves the old version in place).
st := resyncResultFor(rinfos, opts.arn, roi)
var traceSize int64
var traceErr error
for i := range rinfos.Targets {
if rinfos.Targets[i].Arn == opts.arn {
traceSize, traceErr = rinfos.Targets[i].Size, rinfos.Targets[i].Err
break
}
sz = 0
} else {
st.ReplicatedCount++
st.ReplicatedSize += roi.Size
}
traceFn(sz, err)
select {
case <-ctx.Done():
traceFn(traceSize, traceErr)
if !s.sendResyncResult(ctx, results.ch, st) {
return
case <-s.resyncCancelCh:
return
case resultCh <- st:
}
}
}(ctx, i)
}
for res := range objInfoCh {
walkLoop:
for {
var res itemOrErr[ObjectInfo]
select {
case <-ctx.Done():
return
case item, ok := <-objInfoCh:
if !ok {
break walkLoop
}
res = item
}
if res.Err != nil {
resyncStatus = ResyncFailed
replLogIf(ctx, res.Err)
return
}
select {
case <-s.resyncCancelCh:
resyncStatus = ResyncCanceled
return
case <-ctx.Done():
return
default:
}
if heal && lastCheckpoint != "" && lastCheckpoint != res.Item.Name {
continue
}
lastCheckpoint = ""
roi := getHealReplicateObjectInfo(res.Item, rcfg)
if !roi.ExistingObjResync.mustResync() {
// Scope dispatch to this resync's target: the worker pool is for
// opts.arn, so skip objects that only need resync for a different
// target (each target has its own resync). Without this, a cross-target
// object leaves opts.arn absent from its per-object result and is
// miscounted as an opts.arn failure.
if !objectNeedsResyncForARN(roi, opts.arn) {
continue
}
h := xxh3.HashString(roi.Bucket + roi.Name)
select {
case <-s.resyncCancelCh:
return
case <-ctx.Done():
return
default:
h := xxh3.HashString(roi.Bucket + roi.Name)
workers[h%uint64(resyncParallelRoutines)] <- roi
case workers[h%uint64(resyncParallelRoutines)] <- roi:
}
}
for i := range resyncParallelRoutines {
xioutil.SafeClose(workers[i])
}
wg.Wait()
resyncStatus = ResyncCompleted
}
@@ -3087,9 +3427,9 @@ func (s *replicationResyncer) start(ctx context.Context, objAPI ObjectLayer, opt
if len(tgtArns) == 0 {
return fmt.Errorf("arn %s specified for resync not found in replication config", opts.arn)
}
globalReplicationPool.Get().resyncer.RLock()
data, ok := globalReplicationPool.Get().resyncer.statusMap[opts.bucket]
globalReplicationPool.Get().resyncer.RUnlock()
s.Lock()
defer s.Unlock()
data, ok := s.statusMap[opts.bucket]
if !ok {
data, err = loadBucketResyncMetadata(ctx, opts.bucket, objAPI)
if err != nil {
@@ -3110,23 +3450,14 @@ func (s *replicationResyncer) start(ctx context.Context, objAPI ObjectLayer, opt
ResyncStatus: ResyncPending,
Bucket: opts.bucket,
}
data.TargetsMap = data.cloneTgtStats()
data.TargetsMap[opts.arn] = status
if err = saveResyncStatus(ctx, opts.bucket, data, objAPI); err != nil {
return err
}
globalReplicationPool.Get().resyncer.Lock()
defer globalReplicationPool.Get().resyncer.Unlock()
brs, ok := globalReplicationPool.Get().resyncer.statusMap[opts.bucket]
if !ok {
brs = BucketReplicationResyncStatus{
Version: resyncMetaVersion,
TargetsMap: make(map[string]TargetReplicationResyncStatus),
}
}
brs.TargetsMap[opts.arn] = status
globalReplicationPool.Get().resyncer.statusMap[opts.bucket] = brs
go globalReplicationPool.Get().resyncer.resyncBucket(GlobalContext, objAPI, false, opts)
s.statusMap[opts.bucket] = data
go s.resyncBucket(GlobalContext, objAPI, false, opts)
return nil
}
@@ -3200,6 +3531,9 @@ func (p *ReplicationPool) loadResync(ctx context.Context, buckets []string, objA
// Make sure only one node running resync on the cluster.
ctx, cancel := globalLeaderLock.GetLock(ctx)
defer cancel()
var workers sync.WaitGroup
// Keep the merged leader context alive until every resumed run exits.
defer workers.Wait()
for index := range buckets {
bucket := buckets[index]
@@ -3213,19 +3547,24 @@ func (p *ReplicationPool) loadResync(ctx context.Context, buckets []string, objA
}
p.resyncer.Lock()
p.resyncer.statusMap[bucket] = meta
p.resyncer.Unlock()
if current, ok := p.resyncer.statusMap[bucket]; ok {
// A concurrent start/cancel is newer than the disk snapshot.
meta = current
} else {
p.resyncer.statusMap[bucket] = meta
}
tgts := meta.cloneTgtStats()
p.resyncer.Unlock()
for arn, st := range tgts {
switch st.ResyncStatus {
case ResyncFailed, ResyncStarted, ResyncPending:
go p.resyncer.resyncBucket(ctx, objAPI, true, resyncOpts{
opts := resyncOpts{
bucket: bucket,
arn: arn,
resyncID: st.ResyncID,
resyncBefore: st.ResyncBeforeDate,
})
}
workers.Go(func() { p.resyncer.resyncBucket(ctx, objAPI, true, opts) })
}
}
}
@@ -3544,6 +3883,7 @@ func (p *ReplicationPool) queueMRFSave(entry MRFReplicateEntry) {
if entry.RetryCount > mrfRetryLimit { // let scanner catch up if retry count exceeded
atomic.AddUint64(&p.stats.mrfStats.TotalDroppedCount, 1)
atomic.AddUint64(&p.stats.mrfStats.TotalDroppedBytes, uint64(entry.sz))
replLogOnceIf(GlobalContext, errors.New("Replication MRF retry limit reached; further repair is deferred to the scanner"), "replication-mrf-retry-limit", logger.WarningKind)
return
}
@@ -3558,6 +3898,7 @@ func (p *ReplicationPool) queueMRFSave(entry MRFReplicateEntry) {
default:
atomic.AddUint64(&p.stats.mrfStats.TotalDroppedCount, 1)
atomic.AddUint64(&p.stats.mrfStats.TotalDroppedBytes, uint64(entry.sz))
replLogOnceIf(GlobalContext, errors.New("Replication MRF queue is full; dropped entries will need scanner repair"), "replication-mrf-queue-full", logger.WarningKind)
}
}
}
+788
View File
@@ -18,13 +18,21 @@
package cmd
import (
"context"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"path"
"strings"
"sync"
"testing"
"testing/synctest"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio-go/v7"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
)
@@ -307,3 +315,783 @@ func TestReplicationValidationObjectUsesRulePrefix(t *testing.T) {
})
}
}
// The resync-finalization tests below exercise the real result sink, the
// finish() shutdown ordering, and the sendResyncResult / finalResyncStatus
// helpers, plus (for the persistence cases) markStatus with on-disk
// round-tripping. resyncBucket cannot be driven end to end in a unit test
// because its workers call a live remote target (StatObject), so the helpers it
// uses are exercised directly. The blocking-order assertions run under
// testing/synctest so a removed wait fails deterministically, with no timing
// windows.
func newTestResyncer(bucket, arn string) (*replicationResyncer, resyncOpts) {
s := &replicationResyncer{
statusMap: map[string]BucketReplicationResyncStatus{},
}
brs := newBucketResyncStatus(bucket)
brs.TargetsMap[arn] = TargetReplicationResyncStatus{ResyncStatus: ResyncStarted, ResyncID: "reset-" + bucket}
s.statusMap[bucket] = brs
return s, resyncOpts{bucket: bucket, arn: arn, resyncID: "reset-" + bucket}
}
// TestResyncBucketFinalize round-trips the terminal status through a real
// ObjectLayer: a clean run persists Completed with every result, while a run
// whose parent context was canceled during the drain is downgraded to Failed;
// a user-canceled run persists Canceled. Completed never misrepresents an
// incomplete resync.
func TestResyncBucketFinalize(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
objAPI, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatalf("prepare erasure backend: %v", err)
}
defer removeRoots(fsDirs)
// persistTerminal applies resyncBucket's finalizer logic (finalResyncStatus
// then markStatus, which persists) and reads the status back the way the
// resync status API does.
persistTerminal := func(t *testing.T, s *replicationResyncer, opts resyncOpts, status ResyncStatusType, cause error) TargetReplicationResyncStatus {
t.Helper()
s.markStatus(finalResyncStatus(status, cause), opts, objAPI)
brs, err := loadBucketResyncMetadata(ctx, opts.bucket, objAPI)
if err != nil {
t.Fatalf("load persisted resync metadata: %v", err)
}
return brs.TargetsMap[opts.arn]
}
// 1. Clean completion: every result - including the failed object - is folded
// into the persisted status, which stays Completed.
t.Run("persists complete counts", func(t *testing.T) {
s, opts := newTestResyncer("finalize-counts", "arn1")
results := s.newResyncResults(opts)
results.ch <- TargetReplicationResyncStatus{Object: "ok-1", ReplicatedCount: 1, ReplicatedSize: 100}
results.ch <- TargetReplicationResyncStatus{Object: "ok-2", ReplicatedCount: 1, ReplicatedSize: 200}
results.ch <- TargetReplicationResyncStatus{Object: "bad", FailedCount: 1, FailedSize: 300}
var wg sync.WaitGroup // no producer workers for this case
results.finish(nil, &wg)
st := persistTerminal(t, s, opts, ResyncCompleted, nil)
if st.ResyncStatus != ResyncCompleted {
t.Fatalf("persisted status = %s, want Completed", st.ResyncStatus)
}
if st.ReplicatedCount != 2 || st.ReplicatedSize != 300 || st.FailedCount != 1 || st.FailedSize != 300 {
t.Fatalf("persisted counts = {replicated:%d/%d failed:%d/%d}, want {2/300 1/300}",
st.ReplicatedCount, st.ReplicatedSize, st.FailedCount, st.FailedSize)
}
})
// 2. Parent context canceled during the drain -> Completed downgraded to
// Failed (markStatus persists under its own context, so nothing else stops
// a bare Completed from being recorded).
t.Run("parent cancel during drain downgrades to failed", func(t *testing.T) {
s, opts := newTestResyncer("finalize-parent-cancel", "arn1")
results := s.newResyncResults(opts)
results.ch <- TargetReplicationResyncStatus{Object: "ok-1", ReplicatedCount: 1, ReplicatedSize: 100}
var wg sync.WaitGroup
results.finish(nil, &wg)
cctx, ccancel := context.WithCancel(context.Background())
ccancel()
st := persistTerminal(t, s, opts, ResyncCompleted, context.Cause(cctx))
if st.ResyncStatus != ResyncFailed {
t.Fatalf("persisted status = %s, want Failed (parent canceled during drain)", st.ResyncStatus)
}
})
// 3. A user-canceled worker cannot report a completed resync.
t.Run("user cancel persists canceled", func(t *testing.T) {
s, opts := newTestResyncer("finalize-worker-abort", "arn1")
ctx, cancel := context.WithCancelCause(context.Background())
cancel(errResyncCanceled)
ch := make(chan TargetReplicationResyncStatus)
if s.sendResyncResult(ctx, ch, TargetReplicationResyncStatus{Object: "dropped", ReplicatedCount: 1}) {
t.Fatal("sendResyncResult reported success after cancellation")
}
st := persistTerminal(t, s, opts, ResyncCompleted, context.Cause(ctx))
if st.ResyncStatus != ResyncCanceled {
t.Fatalf("persisted status = %s, want Canceled", st.ResyncStatus)
}
})
}
// TestResyncFinishDrainsResults asserts finish() does not return until the
// consumer has applied the final result (the #136 defect). A gated apply holds
// the last result unapplied; under synctest finish() must stay durably blocked
// until it is released - if rr.wg.Wait() is removed, finish() returns early and
// the test fails deterministically.
func TestResyncFinishDrainsResults(t *testing.T) {
synctest.Test(t, func(t *testing.T) {
s, opts := newTestResyncer("drain", "arn1")
reachedFinal := make(chan struct{})
release := make(chan struct{})
results := startResyncResults(func(r TargetReplicationResyncStatus) {
if r.Object == "final" {
close(reachedFinal)
<-release
}
s.incStats(r, opts)
})
results.ch <- TargetReplicationResyncStatus{Object: "ok-1", ReplicatedCount: 1, ReplicatedSize: 100}
results.ch <- TargetReplicationResyncStatus{Object: "final", FailedCount: 1, FailedSize: 200}
<-reachedFinal // consumer received "final" but is gated before incStats(final)
var wg sync.WaitGroup
finishDone := make(chan struct{})
go func() {
results.finish(nil, &wg)
close(finishDone)
}()
synctest.Wait()
select {
case <-finishDone:
close(release)
synctest.Wait()
t.Fatal("finish() returned before the final result was drained (drain wait missing)")
default:
// finish() is durably blocked in rr.wg.Wait() - correct.
}
close(release)
synctest.Wait()
<-finishDone
st := s.statusMap[opts.bucket].TargetsMap[opts.arn]
if st.ReplicatedCount != 1 || st.FailedCount != 1 || st.FailedSize != 200 {
t.Fatalf("status after finish = {replicated:%d failed:%d/%d}, want {1 1/200}",
st.ReplicatedCount, st.FailedCount, st.FailedSize)
}
})
}
// TestResyncFinishWaitsForInflightWorker asserts finish() stops the producer
// workers before it closes the result channel, so an in-flight worker (as on an
// early-return path) never sends on a closed channel and its result is not lost.
// A gated worker stays in flight past the shutdown request; under synctest
// finish() must stay durably blocked until the worker is released - if
// workerWg.Wait() is removed, finish() returns early and the test fails
// deterministically.
func TestResyncFinishWaitsForInflightWorker(t *testing.T) {
synctest.Test(t, func(t *testing.T) {
s, opts := newTestResyncer("workers", "arn1")
results := startResyncResults(func(r TargetReplicationResyncStatus) { s.incStats(r, opts) })
workers := []chan ReplicateObjectInfo{make(chan ReplicateObjectInfo, 1)}
var wg sync.WaitGroup
gotRoi := make(chan struct{})
release := make(chan struct{})
wg.Add(1)
go func() {
defer wg.Done()
for roi := range workers[0] {
close(gotRoi)
<-release
// Mirror the real worker's send; recover so that if finish()
// wrongly closed the result channel first, the test fails via the
// assertion below instead of crashing on send-on-closed.
func() {
defer func() { _ = recover() }()
results.ch <- TargetReplicationResyncStatus{Object: roi.Name, ReplicatedCount: 1, ReplicatedSize: 500}
}()
}
}()
workers[0] <- ReplicateObjectInfo{Name: "inflight"}
<-gotRoi // worker holds a result in flight, not yet delivered
finishDone := make(chan struct{})
go func() {
results.finish(workers, &wg)
close(finishDone)
}()
synctest.Wait()
select {
case <-finishDone:
close(release)
synctest.Wait()
t.Fatal("finish() closed the result channel before the in-flight worker finished (worker wait missing)")
default:
// finish() is durably blocked in workerWg.Wait() - correct.
}
close(release)
synctest.Wait()
<-finishDone
st := s.statusMap[opts.bucket].TargetsMap[opts.arn]
if st.ReplicatedCount != 1 || st.ReplicatedSize != 500 {
t.Fatalf("status after finish = {replicated:%d/%d}, want {1/500}", st.ReplicatedCount, st.ReplicatedSize)
}
})
}
// TestResyncResultFor asserts the resync worker classifies a target from the
// actual replication outcome, not from whether the target version merely exists.
// The key regression is the "failed update over an existing version" case: a
// quota-rejected update leaves the old version in place, and counting existence
// (the previous behavior) would score it a success. It also checks a genuine
// success, an errored-but-Completed result, a delete failure, a delete-marker
// success (zero bytes), and an ARN that was never attempted.
func TestResyncResultFor(t *testing.T) {
const arn = "arn:minio:replication::id:bucket"
obj := ReplicateObjectInfo{Name: "obj", Bucket: "bucket", Size: 196608}
deleteMarker := ReplicateObjectInfo{Name: "dm", Bucket: "bucket", Size: 0, DeleteMarker: true}
tests := []struct {
name string
roi ReplicateObjectInfo
rinfos replicatedInfos
wantRepl, wantReplSize, wantFail, wantFailSize int64
}{
{
name: "completed update",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed, Size: 196608},
}},
wantRepl: 1, wantReplSize: 196608,
},
{
name: "failed update over existing version",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Failed, Err: fmt.Errorf("quota exceeded"), Size: 196608},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "completed but errored is a failure",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed, Err: fmt.Errorf("boom"), Size: 196608},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "delete failed",
roi: deleteMarker,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Failed},
}},
wantFail: 1, wantFailSize: 0,
},
{
name: "delete marker replicated counts zero bytes",
roi: deleteMarker,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed},
}},
wantRepl: 1, wantReplSize: 0,
},
{
name: "arn not attempted is a failure",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: "arn:minio:replication::id2:bucket", ReplicationStatus: replication.Completed, Size: 196608},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "completed with zero size falls back to object size",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, ReplicationStatus: replication.Completed, Size: 0},
}},
wantRepl: 1, wantReplSize: 196608,
},
{
name: "version purge complete is a success",
roi: ReplicateObjectInfo{Name: "purge", Bucket: "bucket", Size: 196608, VersionPurgeStatus: replication.VersionPurgePending},
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
// a successful purge sets only VersionPurgeStatus; ReplicationStatus stays empty.
{Arn: arn, VersionPurgeStatus: replication.VersionPurgeComplete},
}},
wantRepl: 1, wantReplSize: 196608,
},
{
name: "version purge failed is a failure",
roi: ReplicateObjectInfo{Name: "purge", Bucket: "bucket", Size: 196608, VersionPurgeStatus: replication.VersionPurgePending},
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
{Arn: arn, VersionPurgeStatus: replication.VersionPurgeFailed, Err: fmt.Errorf("quota exceeded")},
}},
wantFail: 1, wantFailSize: 196608,
},
{
name: "benign duplicate 412 is a success",
roi: obj,
rinfos: replicatedInfos{Targets: []replicatedTargetInfo{
// the destination answers PreconditionFailed for an exact duplicate;
// replicateAll keeps Completed but retains the error.
{Arn: arn, ReplicationStatus: replication.Completed, Err: minio.ErrorResponse{Code: "PreconditionFailed"}, Size: 196608},
}},
wantRepl: 1, wantReplSize: 196608,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
st := resyncResultFor(tc.rinfos, arn, tc.roi)
if st.Object != tc.roi.Name || st.Bucket != tc.roi.Bucket {
t.Fatalf("object/bucket = %s/%s, want %s/%s", st.Object, st.Bucket, tc.roi.Name, tc.roi.Bucket)
}
if st.ReplicatedCount != tc.wantRepl || st.ReplicatedSize != tc.wantReplSize ||
st.FailedCount != tc.wantFail || st.FailedSize != tc.wantFailSize {
t.Fatalf("resyncResultFor = {replicated:%d/%d failed:%d/%d}, want {%d/%d %d/%d}",
st.ReplicatedCount, st.ReplicatedSize, st.FailedCount, st.FailedSize,
tc.wantRepl, tc.wantReplSize, tc.wantFail, tc.wantFailSize)
}
})
}
}
// TestObjectNeedsResyncForARN asserts the resync dispatch is scoped to the
// target being resynced. The worker pool runs for a single target (opts.arn),
// so an object that only qualifies for a different target must be skipped: with
// A/B rules and a resync of A, an object that needs replication only for B must
// not be admitted to A's worker. Otherwise (after outcome-based classification)
// A would be absent from that object's result and miscounted as an A failure.
func TestObjectNeedsResyncForARN(t *testing.T) {
const (
arnA = "arn:minio:replication::id:bucket"
arnB = "arn:minio:replication::id2:bucket"
)
tests := []struct {
name string
decision ResyncDecision
arn string
want bool
}{
{
name: "target must resync",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{arnA: {Replicate: true}}},
arn: arnA,
want: true,
},
{
name: "object qualifies for B only, resyncing A",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{arnB: {Replicate: true}}},
arn: arnA,
want: false,
},
{
name: "A present but not replicating, B replicating, resyncing A",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{
arnA: {Replicate: false},
arnB: {Replicate: true},
}},
arn: arnA,
want: false,
},
{
name: "object qualifies for both, resyncing A",
decision: ResyncDecision{targets: map[string]ResyncTargetDecision{
arnA: {Replicate: true},
arnB: {Replicate: true},
}},
arn: arnA,
want: true,
},
{
name: "no resync decision",
decision: ResyncDecision{},
arn: arnA,
want: false,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
roi := ReplicateObjectInfo{Name: "obj", Bucket: "bucket", ExistingObjResync: tc.decision}
if got := objectNeedsResyncForARN(roi, tc.arn); got != tc.want {
t.Fatalf("objectNeedsResyncForARN(arn=%s) = %v, want %v", tc.arn, got, tc.want)
}
})
}
}
// newMatchingReplicationPair returns a source/target pair that getReplicationAction must
// classify as replicateNone: same ETag, version id, size, modification time and content
// type. Any action other than replicateNone is therefore attributable to the object lock
// entries a caller adds on top.
func newMatchingReplicationPair() (ObjectInfo, minio.ObjectInfo) {
mtime := time.Date(2026, 9, 5, 10, 0, 0, 0, time.UTC)
size := int64(7)
src := ObjectInfo{
Bucket: "bucket",
Name: "object",
ETag: "d41d8cd98f00b204e9800998ecf8427e",
VersionID: "b0ff1d6e-0000-4000-8000-000000000001",
Size: size,
ActualSize: &size,
ModTime: mtime,
ContentType: "application/octet-stream",
UserDefined: map[string]string{"content-type": "application/octet-stream"},
}
tgt := minio.ObjectInfo{
ETag: src.ETag,
VersionID: src.VersionID,
Size: size,
LastModified: mtime,
ContentType: src.ContentType,
Metadata: http.Header{},
}
return src, tgt
}
// TestGetReplicationActionEmptyObjectLockValues covers the comparison of object lock entries
// whose value is empty. Removing retention from a version stores the mode and retain-until-date
// keys with empty values, while the target's HEAD response omits them entirely, so the two must
// compare equal or the version can never be reported as in sync. Cases 3 and 4 are synthetic
// comparison inputs, since a SILO target cannot return empty lock headers; cases 7 and 8 guard
// against over-normalizing.
func TestGetReplicationActionEmptyObjectLockValues(t *testing.T) {
var (
modeKey = strings.ToLower(xhttp.AmzObjectLockMode)
dateKey = strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
until = "2026-10-05T10:00:00.000Z"
)
emptyRetention := map[string]string{modeKey: "", dateKey: ""}
realRetention := map[string]string{modeKey: "GOVERNANCE", dateKey: until}
tests := []struct {
name string
srcMeta map[string]string
tgtHdr map[string]string
want replicationAction
}{
{"1-both-clean-never-had-retention", nil, nil, replicateNone},
{"2-source-present-empty-target-absent", emptyRetention, nil, replicateNone},
{"3-source-absent-target-present-empty", nil, emptyRetention, replicateNone},
{"4-both-present-empty", emptyRetention, emptyRetention, replicateNone},
{"5-both-governance-equal", realRetention, realRetention, replicateNone},
{"6-source-governance-target-absent", realRetention, nil, replicateMetadata},
{"7-source-empty-target-real-retention", emptyRetention, realRetention, replicateMetadata},
{"8-empty-user-metadata-is-not-normalized", map[string]string{"x-amz-meta-foo": ""}, nil, replicateMetadata},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgt := newMatchingReplicationPair()
for k, v := range test.srcMeta {
src.UserDefined[k] = v
}
for k, v := range test.tgtHdr {
tgt.Metadata.Set(k, v)
}
if got := getReplicationAction(src, tgt, replication.HealReplicationType); got != test.want {
t.Fatalf("getReplicationAction() = %q, want %q (source %v, target %v)", got, test.want, src.UserDefined, tgt.Metadata)
}
})
}
}
// TestEmptyRetentionValuesAreOmittedFromObjectResponseHeaders records why the target half of the
// comparison in getReplicationAction can never report an empty object lock entry:
// FilterObjectLockMetadata drops both keys because an empty mode is not a valid retention mode,
// and setObjectHeaders skips them when writing response headers. Neither filter reaches the
// replication wire: the empty entries are still carried by getCopyObjMetadata and sent by the
// metadata CopyObject, which is why the sender's comparison is what has to tolerate them.
// FilterObjectLockMetadata is also applied by CopyObject (cmd/object-handlers.go:1708), where it
// strips the source's lock metadata before the destination re-derives it from the request.
func TestEmptyRetentionValuesAreOmittedFromObjectResponseHeaders(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
meta := map[string]string{
modeKey: "",
dateKey: "",
"content-type": "application/octet-stream",
}
filtered := objectlock.FilterObjectLockMetadata(meta, false, false)
if _, ok := filtered[modeKey]; ok {
t.Errorf("FilterObjectLockMetadata() kept the empty lock mode key: %v", filtered)
}
if _, ok := filtered[dateKey]; ok {
t.Errorf("FilterObjectLockMetadata() kept the empty retain-until-date key: %v", filtered)
}
rec := httptest.NewRecorder()
if err := setObjectHeaders(t.Context(), rec, ObjectInfo{UserDefined: meta, ModTime: time.Now(), Size: 7}, nil, ObjectOptions{}); err != nil {
t.Fatalf("setObjectHeaders() = %v", err)
}
if v, ok := rec.Header()[http.CanonicalHeaderKey(xhttp.AmzObjectLockMode)]; ok {
t.Errorf("setObjectHeaders() emitted an empty lock mode header: %v", v)
}
if v, ok := rec.Header()[http.CanonicalHeaderKey(xhttp.AmzObjectLockRetainUntilDate)]; ok {
t.Errorf("setObjectHeaders() emitted an empty retain-until-date header: %v", v)
}
}
// fakeRetentionGetter answers GetObjectRetention with a fixed result and counts its calls.
type fakeRetentionGetter struct {
mode *minio.RetentionMode
err error
calls int
}
func (f *fakeRetentionGetter) GetObjectRetention(_ context.Context, _, _, _ string) (*minio.RetentionMode, *time.Time, error) {
f.calls++
return f.mode, nil, f.err
}
func TestRetentionRemovedAtSource(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
tsKey := ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp
stamp := "2026-09-06T01:00:00Z"
tests := []struct {
name string
meta map[string]string
want bool
}{
{"no lock keys", map[string]string{"content-type": "text/plain"}, false},
{"empty pair", map[string]string{modeKey: "", dateKey: ""}, true},
{"empty mode only", map[string]string{modeKey: ""}, true},
{"empty date only", map[string]string{dateKey: ""}, true},
{"real retention", map[string]string{modeKey: "GOVERNANCE", dateKey: "2026-10-05T10:00:00.000Z"}, false},
{"canonical case", map[string]string{xhttp.AmzObjectLockMode: ""}, true},
{"empty user metadata", map[string]string{"x-amz-meta-foo": ""}, false},
// Representation (2): a replicated removal persists the ordering timestamp alone,
// with the mode and retain-until-date keys absent (restoreRetention).
{"timestamp only, mode absent", map[string]string{tsKey: stamp}, true},
{"timestamp with empty mode", map[string]string{tsKey: stamp, modeKey: ""}, true},
{"timestamp with real retention is a set, not a removal", map[string]string{tsKey: stamp, modeKey: "GOVERNANCE", dateKey: "2026-10-05T10:00:00.000Z"}, false},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
if got := retentionRemovedAtSource(ObjectInfo{UserDefined: test.meta}); got != test.want {
t.Fatalf("retentionRemovedAtSource() = %v, want %v", got, test.want)
}
})
}
}
// TestTargetRetentionConfirmedAbsent pins the rule that only an explicit answer from the
// destination clears a removed retention. A denied or unreachable destination must read as still
// holding retention, because HEAD hides a real retention from a credential without
// s3:GetObjectRetention exactly as it hides one that does not exist.
func TestTargetRetentionConfirmedAbsent(t *testing.T) {
governance := minio.Governance
var emptyMode minio.RetentionMode
unknownMode := minio.RetentionMode("ARCHIVE")
tests := []struct {
name string
mode *minio.RetentionMode
err error
want bool
}{
{"version holds governance retention", &governance, nil, false},
{"no retention on the version", nil, minio.ErrorResponse{Code: "NoSuchObjectLockConfiguration"}, true},
{
// The destination also answers this when its own read of the bucket's Object Lock
// configuration fails, so it does not establish that Object Lock is disabled.
"invalid request naming a missing object lock configuration",
nil,
minio.ErrorResponse{Code: "InvalidRequest", Message: "Bucket is missing ObjectLockConfiguration"},
false,
},
{
"unrelated invalid request",
nil,
minio.ErrorResponse{Code: "InvalidRequest", Message: "Object is WORM protected and cannot be overwritten"},
false,
},
{"retention read denied", nil, minio.ErrorResponse{Code: "AccessDenied"}, false},
{"destination unreachable", nil, errors.New("dial tcp: connection refused"), false},
{"empty mode returned", &emptyMode, nil, true},
{"unknown non-empty mode returned", &unknownMode, nil, false},
{"nil mode returned", nil, nil, true},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
if got := targetRetentionConfirmedAbsent(t.Context(), tgt, "bucket", "object", "v1"); got != test.want {
t.Fatalf("targetRetentionConfirmedAbsent() = %v, want %v", got, test.want)
}
if tgt.calls != 1 {
t.Fatalf("GetObjectRetention called %d times, want 1", tgt.calls)
}
})
}
}
// TestReplicationActionForTargetRetentionRemoval covers the decision the replication worker makes
// for a version whose retention was removed. The destination's HEAD never reports the empty keys,
// so the comparison alone reads every one of these as in sync; only the confirmation separates a
// destination that really dropped the retention from one that is hiding it.
func TestReplicationActionForTargetRetentionRemoval(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
governance := minio.Governance
tests := []struct {
name string
srcMeta map[string]string
mode *minio.RetentionMode
err error
want replicationAction
wantCalls int
}{
{
name: "removal confirmed by destination",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
err: minio.ErrorResponse{Code: "NoSuchObjectLockConfiguration"},
want: replicateNone,
wantCalls: 1,
},
{
name: "destination still holds the retention hidden from HEAD",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
mode: &governance,
want: replicateMetadata,
wantCalls: 1,
},
{
name: "retention hidden from HEAD by permissions",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
err: minio.ErrorResponse{Code: "AccessDenied"},
want: replicateMetadata,
wantCalls: 1,
},
{
// A destination that names a missing Object Lock configuration answers the same way
// when its own read of that configuration failed, so it confirms nothing.
name: "destination reports no object lock configuration",
srcMeta: map[string]string{modeKey: "", dateKey: ""},
err: minio.ErrorResponse{Code: "InvalidRequest", Message: "Bucket is missing ObjectLockConfiguration"},
want: replicateMetadata,
wantCalls: 1,
},
{
name: "version never had retention is not confirmed",
srcMeta: nil,
want: replicateNone,
wantCalls: 0,
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgtInfo := newMatchingReplicationPair()
for k, v := range test.srcMeta {
src.UserDefined[k] = v
}
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
got := replicationActionForTarget(t.Context(), src, tgtInfo, replication.HealReplicationType, tgt, "bucket", "object")
if got != test.want {
t.Fatalf("replicationActionForTarget() = %q, want %q", got, test.want)
}
if tgt.calls != test.wantCalls {
t.Fatalf("GetObjectRetention called %d times, want %d", tgt.calls, test.wantCalls)
}
})
}
}
// TestReplicationActionForTargetNullVersionResync pins that the confirmation does not reopen the
// null-version exclusion at the head of getReplicationAction. An existing object resync returns
// replicateNone for a null version whose source modification time is later than the target's,
// before comparing anything, and that must stand even when the source carries a removed retention
// and the destination would report retention or refuse to answer.
func TestReplicationActionForTargetNullVersionResync(t *testing.T) {
modeKey := strings.ToLower(xhttp.AmzObjectLockMode)
dateKey := strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)
governance := minio.Governance
tests := []struct {
name string
mode *minio.RetentionMode
err error
}{
{"destination holds retention", &governance, nil},
{"retention read denied", nil, minio.ErrorResponse{Code: "AccessDenied"}},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgtInfo := newMatchingReplicationPair()
// A null version whose source modification time is later, and whose content differs,
// so only the exclusion can hold the action at replicateNone.
src.VersionID = nullVersionID
src.ModTime = tgtInfo.LastModified.Add(time.Hour)
src.ETag = "5d41402abc4b2a76b9719d911017c592"
src.UserDefined[modeKey] = ""
src.UserDefined[dateKey] = ""
tgtInfo.VersionID = nullVersionID
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
got := replicationActionForTarget(t.Context(), src, tgtInfo, replication.ExistingObjectReplicationType, tgt, "bucket", "object")
if got != replicateNone {
t.Fatalf("replicationActionForTarget() = %q, want %q", got, replicateNone)
}
if tgt.calls != 0 {
t.Fatalf("GetObjectRetention called %d times, want 0", tgt.calls)
}
})
}
}
// TestReplicationActionForTargetTimestampOnlyRemoval covers representation (2) of a removed
// retention. A removal that arrived by replication persists only the retention ordering timestamp,
// with the mode and retain-until-date keys absent, because restoreRetention writes the timestamp
// alone when the mode is empty (cmd/bucket-object-lock.go). The comparison in getReplicationAction
// reads such a source as in sync with a matching destination, so only the GetObjectRetention
// confirmation separates a destination that dropped the retention from one hiding it behind a
// permission-filtered HEAD. The source is built through the real restoreRetention path so the
// fixture is the metadata a replicated removal actually leaves on disk, not a hand-rolled map.
func TestReplicationActionForTargetTimestampOnlyRemoval(t *testing.T) {
governance := minio.Governance
stamp := time.Date(2026, 9, 6, 1, 0, 0, 0, time.UTC).Format(time.RFC3339Nano)
tests := []struct {
name string
mode *minio.RetentionMode
err error
want replicationAction
wantCalls int
}{
{
// The reference case: a destination that denies the retention read is
// indistinguishable from one still holding it, so the removal is resent.
name: "retention hidden from HEAD by permissions",
err: minio.ErrorResponse{Code: "AccessDenied"},
want: replicateMetadata,
wantCalls: 1,
},
{
name: "destination still holds the retention",
mode: &governance,
want: replicateMetadata,
wantCalls: 1,
},
{
name: "removal confirmed by destination",
err: minio.ErrorResponse{Code: "NoSuchObjectLockConfiguration"},
want: replicateNone,
wantCalls: 1,
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
src, tgtInfo := newMatchingReplicationPair()
// Persist the timestamp-only tombstone the same way an applied replica removal does.
objectLockState{retentionTimestamp: stamp}.restoreRetention(src.UserDefined)
if !retentionRemovedAtSource(src) {
t.Fatalf("restoreRetention fixture not recognized as a removal: %v", src.UserDefined)
}
if _, ok := src.UserDefined[strings.ToLower(xhttp.AmzObjectLockMode)]; ok {
t.Fatalf("restoreRetention fixture wrote a mode key, fixture is not timestamp-only: %v", src.UserDefined)
}
tgt := &fakeRetentionGetter{mode: test.mode, err: test.err}
got := replicationActionForTarget(t.Context(), src, tgtInfo, replication.HealReplicationType, tgt, "bucket", "object")
if got != test.want {
t.Fatalf("replicationActionForTarget() = %q, want %q", got, test.want)
}
if tgt.calls != test.wantCalls {
t.Fatalf("GetObjectRetention called %d times, want %d", tgt.calls, test.wantCalls)
}
})
}
}
+6 -2
View File
@@ -319,7 +319,9 @@ func (r *ReplicationStats) getNodeQueueStats(bucket string) (qs ReplQNodeStats)
qs.QStats = r.qCache.getBucketStats(bucket)
qs.TgtXferStats = make(map[string]map[RMetricName]XferStats)
qs.MRFStats = ReplicationMRFStats{
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
TotalDroppedCount: atomic.LoadUint64(&r.mrfStats.TotalDroppedCount),
TotalDroppedBytes: atomic.LoadUint64(&r.mrfStats.TotalDroppedBytes),
}
r.RLock()
@@ -410,7 +412,9 @@ func (r *ReplicationStats) getNodeQueueStatsSummary() (qs ReplQNodeStats) {
qs.XferStats = make(map[RMetricName]XferStats)
qs.QStats = r.qCache.getSiteStats()
qs.MRFStats = ReplicationMRFStats{
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
LastFailedCount: atomic.LoadUint64(&r.mrfStats.LastFailedCount),
TotalDroppedCount: atomic.LoadUint64(&r.mrfStats.TotalDroppedCount),
TotalDroppedBytes: atomic.LoadUint64(&r.mrfStats.TotalDroppedBytes),
}
r.RLock()
defer r.RUnlock()
+3 -3
View File
@@ -97,7 +97,7 @@ func (api objectAPIHandlers) PutBucketVersioningHandler(w http.ResponseWriter, r
return
}
updatedAt, err := globalBucketMetadataSys.Update(ctx, bucket, bucketVersioningConfig, configData)
result, err := globalBucketMetadataSys.updateAndParseMetadata(ctx, bucket, bucketVersioningConfig, configData, false, false, nil)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -107,12 +107,12 @@ func (api objectAPIHandlers) PutBucketVersioningHandler(w http.ResponseWriter, r
//
// We encode the xml bytes as base64 to ensure there are no encoding
// errors.
cfgStr := base64.StdEncoding.EncodeToString(configData)
cfgStr := base64.StdEncoding.EncodeToString(result.meta.VersioningConfigXML)
replLogIf(ctx, globalSiteReplicationSys.BucketMetaHook(ctx, madmin.SRBucketMeta{
Type: madmin.SRBucketMetaTypeVersionConfig,
Bucket: bucket,
Versioning: &cfgStr,
UpdatedAt: updatedAt,
UpdatedAt: result.updatedAt,
}))
writeSuccessResponseHeadersOnly(w)
+53 -7
View File
@@ -120,15 +120,40 @@ func init() {
const consolePrefix = "CONSOLE_"
// consoleMinIOServerEnv derives the CONSOLE_MINIO_SERVER value the embedded
// Console uses to reach the S3/STS API, and whether TLS verification of that
// endpoint must be skipped. With no explicit endpoint configured the Console
// reaches the API over the loopback address, whose TLS certificate is not
// expected to carry a 127.0.0.1 SAN; because Console verifies outbound TLS by
// default, the loopback origin has to be exempted or embedded login (local and
// LDAP alike) fails at the STS handshake. The exemption is endpoint-scoped in
// Console, so every other HTTPS peer stays verified. An explicitly configured
// endpoint is always reached under its own verified name and is never exempted.
func consoleMinIOServerEnv(endpoint string, isTLS bool, port string) (server string, skipVerify bool) {
if endpoint != "" {
return endpoint, false
}
return fmt.Sprintf("%s://127.0.0.1:%s", getURLScheme(isTLS), port), isTLS
}
func minioConfigToConsoleFeatures() {
os.Setenv("CONSOLE_PBKDF_SALT", globalDeploymentID())
os.Setenv("CONSOLE_PBKDF_PASSPHRASE", globalDeploymentID())
if globalMinioEndpoint != "" {
os.Setenv("CONSOLE_MINIO_SERVER", globalMinioEndpoint)
consoleServer, skipVerify := consoleMinIOServerEnv(globalMinioEndpoint, globalIsTLS, globalMinioPort)
os.Setenv("CONSOLE_MINIO_SERVER", consoleServer)
if skipVerify {
// The embedded Console reaches the loopback S3/STS endpoint above, whose
// certificate is not expected to carry a 127.0.0.1 SAN. Console verifies
// outbound TLS by default (silo-console v2.3.x), so opt into the
// endpoint-scoped compatibility switch to preserve the documented loopback
// bypass; every other HTTPS peer (IdP, Prometheus, webhooks, ...) stays
// verified. initConsoleServer unsets CONSOLE_* before calling this, so the
// switch cannot be supplied by the operator on the embedded path.
os.Setenv("CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY", "on")
} else {
// Explicitly set 127.0.0.1 so Console will automatically bypass TLS verification to the local S3 API.
// This will save users from providing a certificate with IP or FQDN SAN that points to the local host.
os.Setenv("CONSOLE_MINIO_SERVER", fmt.Sprintf("%s://127.0.0.1:%s", getURLScheme(globalIsTLS), globalMinioPort))
// An explicitly configured endpoint is reached under its own verified name;
// never let a loopback exemption apply to it.
os.Unsetenv("CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY")
}
if value := env.Get(config.EnvMinIOLogQueryURL, ""); value != "" {
os.Setenv("CONSOLE_LOG_QUERY_URL", value)
@@ -232,11 +257,31 @@ func buildOpenIDConsoleConfig() consoleoauth2.OpenIDPCfg {
return m
}
func initConsoleServer() (*consoleapi.Server, error) {
// unset all console_ environment variables.
// resetConsoleEnvironment preserves the embedded Console's supported resource
// settings verbatim. Server derives all other Console settings itself.
func resetConsoleEnvironment() {
for _, cenv := range env.List(consolePrefix) {
switch cenv {
case consoleapi.ConsoleWSMaxConnections,
consoleapi.ConsoleWSMaxConnectionsPerClient,
consoleapi.ConsoleWSMaxAnonymousConnections,
consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient:
continue
}
os.Unsetenv(cenv)
}
}
func initConsoleServer() (*consoleapi.Server, error) {
resetConsoleEnvironment()
// Validate explicitly: ConfigureAPI logs errors, but embedded Console logs
// are normally silenced. Return configuration failures to Server startup.
if err := consoleapi.ConfigureEmbeddedSourceIPTrust(); err != nil {
return nil, err
}
if err := consoleapi.ConfigureWebSocketLimits(); err != nil {
return nil, err
}
// enable all console environment variables
minioConfigToConsoleFeatures()
@@ -855,6 +900,7 @@ func serverHandleEnvVars() {
}
globalEnableSyncBoot = env.Get("MINIO_SYNC_BOOT", config.EnableOff) == config.EnableOn
globalSiteReplicationMetadataTombstones = env.Get("MINIO_SITE_REPLICATION_METADATA_TOMBSTONES", config.EnableOff) == config.EnableOn
}
func loadRootCredentials() auth.Credentials {
+136
View File
@@ -26,6 +26,7 @@ import (
"strings"
"testing"
consoleapi "github.com/minio/console/api"
"github.com/minio/minio/internal/config"
)
@@ -372,3 +373,138 @@ func TestConfigEnvFileNamedTargetDiscovery(t *testing.T) {
t.Fatalf("named target %q not discovered from %s: %v", "my-hook", key, targets)
}
}
// TestConsoleMinIOServerEnv locks in the loopback TLS exemption that keeps
// embedded Console login working (issue #108) while ensuring an explicitly
// configured endpoint is never silently exempted from TLS verification.
func TestConsoleMinIOServerEnv(t *testing.T) {
tests := []struct {
name string
endpoint string
isTLS bool
port string
wantServer string
wantSkipVerify bool
}{
{
name: "loopback TLS is exempted so embedded login works",
isTLS: true,
port: "9000",
wantServer: "https://127.0.0.1:9000",
wantSkipVerify: true,
},
{
name: "loopback plain HTTP needs no exemption",
isTLS: false,
port: "9000",
wantServer: "http://127.0.0.1:9000",
},
{
name: "explicit https endpoint stays verified",
endpoint: "https://silo.example:9000",
isTLS: true,
port: "9000",
wantServer: "https://silo.example:9000",
},
{
name: "explicit http endpoint stays verified",
endpoint: "http://silo.example:9000",
isTLS: false,
port: "9000",
wantServer: "http://silo.example:9000",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
server, skipVerify := consoleMinIOServerEnv(tt.endpoint, tt.isTLS, tt.port)
if server != tt.wantServer {
t.Fatalf("server = %q, want %q", server, tt.wantServer)
}
if skipVerify != tt.wantSkipVerify {
t.Fatalf("skipVerify = %v, want %v", skipVerify, tt.wantSkipVerify)
}
})
}
}
// The startup path clears process environment, so preserve all existing Console
// variables, including ones unrelated to this test, before exercising it.
func preserveConsoleEnvironment(t *testing.T) {
t.Helper()
for _, entry := range os.Environ() {
if strings.HasPrefix(entry, consolePrefix) {
name, value, _ := strings.Cut(entry, "=")
t.Setenv(name, value)
}
}
}
func TestResetConsoleEnvironment(t *testing.T) {
preserveConsoleEnvironment(t)
settings := map[string]string{
consoleapi.ConsoleWSMaxConnections: "2048",
consoleapi.ConsoleWSMaxConnectionsPerClient: "512",
consoleapi.ConsoleWSMaxAnonymousConnections: "128",
consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient: "16",
}
for key, value := range settings {
t.Setenv(key, value)
}
decoys := []string{"CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY", "CONSOLE_MINIO_SERVER", "CONSOLE_PBKDF_SALT", "CONSOLE_TRUSTED_PROXIES", "CONSOLE_WS_MAX_UNKNOWN"}
for _, key := range decoys {
t.Setenv(key, "operator-value")
}
resetConsoleEnvironment()
for key, want := range settings {
if got, present := os.LookupEnv(key); !present || got != want {
t.Errorf("%s = %q, present = %v; want %q", key, got, present, want)
}
}
for _, key := range decoys {
if _, present := os.LookupEnv(key); present {
t.Errorf("unsupported override %s survived", key)
}
}
for _, raw := range []string{"", " 16 ", "env://missing-limit"} {
t.Setenv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient, raw)
resetConsoleEnvironment()
if got, present := os.LookupEnv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient); !present || got != raw {
t.Fatalf("raw value %q was changed to %q (present = %v)", raw, got, present)
}
}
os.Unsetenv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient)
resetConsoleEnvironment()
if _, present := os.LookupEnv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient); present {
t.Fatal("unset setting became present")
}
}
func TestInitConsoleServerConfigurationErrors(t *testing.T) {
for _, tt := range []struct {
name, proxy, limit, want string
}{
{"proxy error precedes limit error", "proxy.internal", "bad", "MINIO_API_TRUSTED_PROXIES"},
{"blank limit", "", "", "CONSOLE_WS_MAX_ANONYMOUS_CONNECTIONS_PER_CLIENT"},
{"non-integer limit", "", "bad", "CONSOLE_WS_MAX_ANONYMOUS_CONNECTIONS_PER_CLIENT"},
{"out-of-range limit", "", "0", "CONSOLE_WS_MAX_ANONYMOUS_CONNECTIONS_PER_CLIENT"},
{"inconsistent limits", "", "256", "must be less than"},
} {
t.Run(tt.name, func(t *testing.T) {
// Restore the process-wide library configuration after environment cleanup.
t.Cleanup(func() {
_ = consoleapi.ConfigureEmbeddedSourceIPTrust()
_ = consoleapi.ConfigureWebSocketLimits()
})
preserveConsoleEnvironment(t)
t.Setenv(consoleapi.EnvMinIOTrustedProxies, tt.proxy)
t.Setenv(consoleapi.ConsoleWSMaxConnections, "1024")
t.Setenv(consoleapi.ConsoleWSMaxConnectionsPerClient, "256")
t.Setenv(consoleapi.ConsoleWSMaxAnonymousConnections, "64")
t.Setenv(consoleapi.ConsoleWSMaxAnonymousConnectionsPerClient, tt.limit)
server, err := initConsoleServer()
if err == nil || !strings.Contains(err.Error(), tt.want) || server != nil {
t.Fatalf("initConsoleServer() = %v, %v; want nil server and %q error", server, err, tt.want)
}
})
}
}
+738
View File
@@ -0,0 +1,738 @@
// Copyright (c) 2015-2026 MinIO, Inc.
// Copyright (c) 2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"archive/tar"
"bytes"
"crypto/md5"
"encoding/base64"
"encoding/xml"
"io"
"net/http"
"net/http/httptest"
"strconv"
"testing"
"time"
"github.com/klauspost/compress/s2"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// disableCompression turns the global compression config off and returns a
// restore func. It is the replication destination that applies no transform of
// its own; setCopyChecksumCompression covers the enabled cases.
func disableCompression() func() {
globalCompressConfigMu.Lock()
previous := globalCompressConfig
globalCompressConfig.Enabled = false
globalCompressConfigMu.Unlock()
return func() {
globalCompressConfigMu.Lock()
globalCompressConfig = previous
globalCompressConfigMu.Unlock()
}
}
// ssecTestHeaders builds the customer key headers for a key made of the given
// repeated byte.
func ssecTestHeaders(b byte) map[string]string {
key := bytes.Repeat([]byte{b}, 32)
keyMD5 := md5.Sum(key)
return map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
}
// assertStoredSSECUncompressed requires the stored object to be SSE-C sealed
// and to carry no compression marker, so that "not compressed" is never
// reported for an object that is not encrypted either.
func assertStoredSSECUncompressed(t *testing.T, obj ObjectLayer, bucketName, object string) ObjectInfo {
t.Helper()
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if _, sealed := info.UserDefined[crypto.MetaSealedKeySSEC]; !sealed {
t.Fatalf("%s is not SSE-C sealed, the fixture proves nothing (userDefined=%v)", object, info.UserDefined)
}
if marker, compressed := info.UserDefined[ReservedMetadataPrefix+"compression"]; compressed {
t.Errorf("%s was stored as a compressed SSE-C object (compression=%q); such an object cannot be replicated",
object, marker)
}
return info
}
// assertSSECPlaintext GETs an SSE-C object with its customer key and requires
// the body to equal the plaintext.
func assertSSECPlaintext(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
bucketName, object string, sseHeaders map[string]string, want []byte,
) {
t.Helper()
req, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("GET %s status %d: %s", object, rec.Code, rec.Body.String())
}
if bytes.Equal(rec.Body.Bytes(), want) {
return
}
body := rec.Body.Bytes()
head := body
if len(head) > 16 {
head = head[:16]
}
t.Errorf("GET %s returned %d bytes, want the %d byte plaintext; first bytes % x",
object, len(body), len(want), head)
// Name the failure mode: a body that s2-decodes to the plaintext is the raw
// S2 stream of a compressed source shipped without its compression marker.
if decoded, derr := io.ReadAll(s2.NewReader(bytes.NewReader(body))); derr == nil && bytes.Equal(decoded, want) {
t.Errorf("the returned body is the raw S2 stream: s2-decoding it yields the %d byte plaintext", len(want))
}
}
// TestAPISSECCompressionReplicaStaysReadable replicates an SSE-C object written
// with compression enabled and allow_encryption=on, and requires the replica to
// read back as the source plaintext.
//
// The source object is read the way the replication worker reads it
// (ReplicationRequest, hence NoDecryption), its wire headers come from the
// production option builder putReplicationOpts, and the replica is written the
// way a destination that applies no transform of its own stores it: compression
// off and no default encryption.
//
// Before the SSE-C compression exclusion the source was stored as
// encrypt(s2(plaintext)) while putReplicationOpts dropped
// X-Minio-Internal-compression, so the replica decrypted to an S2 stream and a
// correct-key GET returned HTTP 200 with the wrong body.
func TestAPISSECCompressionReplicaStaysReadable(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionReplicaStaysReadable,
})
}
func testAPISSECCompressionReplicaStaysReadable(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
replicator := newObjectAttributesAuthzUser(t, instanceType, bucketName, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
sseHeaders := ssecTestHeaders(0x42)
// Highly compressible and comfortably above minCompressibleSize (4096).
data := bytes.Repeat([]byte("silo compressed ssec replication payload "), 8192)
t.Run(instanceType+"/single-put", func(t *testing.T) {
object := "replication/ssec-single.txt"
// --- Source side: compression ON with allow_encryption ON. ---
restore := setCopyChecksumCompression(true)
srcReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
srcRec := httptest.NewRecorder()
apiRouter.ServeHTTP(srcRec, srcReq)
if srcRec.Code != http.StatusOK {
restore()
t.Fatalf("source PUT status %d: %s", srcRec.Code, srcRec.Body.String())
}
sourceInfo := assertStoredSSECUncompressed(t, obj, bucketName, object)
t.Logf("source: stored size=%d compression=%q plaintext=%d", sourceInfo.Size,
sourceInfo.UserDefined[ReservedMetadataPrefix+"compression"], len(data))
// The replication worker's read: raw stored bytes, no decryption.
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{ReplicationRequest: true})
if err != nil {
restore()
t.Fatal(err)
}
sourceInfo = gr.ObjInfo
raw, err := io.ReadAll(gr)
gr.Close()
if err != nil {
restore()
t.Fatal(err)
}
replicationOpts, isMP, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
restore()
t.Fatalf("putReplicationOpts rejected the source: %v", err)
}
if isMP {
restore()
t.Fatal("single PUT source classified as multipart")
}
headers := map[string]string{}
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
headers[name] = values[0]
}
}
restore()
// --- Destination side: NO compression, NO default encryption. ---
restoreDst := disableCompression()
defer restoreDst()
replReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(raw)), bytes.NewReader(raw), replicator.AccessKey, replicator.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
replRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replRec, replReq)
if replRec.Code != http.StatusOK {
t.Fatalf("replica PUT status %d: %s", replRec.Code, replRec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("replica: stored size=%d compression=%q actual-size=%q", info.Size,
info.UserDefined[ReservedMetadataPrefix+"compression"],
info.UserDefined[ReservedMetadataPrefix+"actual-size"])
assertSSECPlaintext(t, apiRouter, credentials, bucketName, object, sseHeaders, data)
})
t.Run(instanceType+"/multipart", func(t *testing.T) {
object := "replication/ssec-mpu.txt"
restore := setCopyChecksumCompression(true)
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
newRec := httptest.NewRecorder()
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
restore()
t.Fatalf("source NewMultipart status %d: %s", newRec.Code, newRec.Body.String())
}
var sourceInit InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &sourceInit, int64(newRec.Body.Len())); err != nil {
restore()
t.Fatal(err)
}
partReq, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, sourceInit.UploadID, "1"),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
partRec := httptest.NewRecorder()
apiRouter.ServeHTTP(partRec, partReq)
if partRec.Code != http.StatusOK {
restore()
t.Fatalf("source PutPart status %d: %s", partRec.Code, partRec.Body.String())
}
completeBody, err := xml.Marshal(CompleteMultipartUpload{Parts: []CompletePart{
{PartNumber: 1, ETag: canonicalizeETag(partRec.Header()[xhttp.ETag][0])},
}})
if err != nil {
restore()
t.Fatal(err)
}
completeReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, sourceInit.UploadID),
int64(len(completeBody)), bytes.NewReader(completeBody), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
restore()
t.Fatal(err)
}
completeRec := httptest.NewRecorder()
apiRouter.ServeHTTP(completeRec, completeReq)
if completeRec.Code != http.StatusOK {
restore()
t.Fatalf("source Complete status %d: %s", completeRec.Code, completeRec.Body.String())
}
assertStoredSSECUncompressed(t, obj, bucketName, object)
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{ReplicationRequest: true})
if err != nil {
restore()
t.Fatal(err)
}
sourceInfo := gr.ObjInfo
rawPart, err := io.ReadAll(gr)
gr.Close()
if err != nil {
restore()
t.Fatal(err)
}
actualSize, err := sourceInfo.GetActualSize()
if err != nil {
restore()
t.Fatal(err)
}
t.Logf("source mpu: stored size=%d actual-size=%d rawRead=%d plaintext=%d compression=%q",
sourceInfo.Size, actualSize, len(rawPart), len(data),
sourceInfo.UserDefined[ReservedMetadataPrefix+"compression"])
replicationOpts, isMP, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
restore()
t.Fatalf("putReplicationOpts rejected the source: %v", err)
}
if !isMP {
restore()
t.Fatal("SSE-C multipart source not recognized as multipart")
}
replicationOpts.Internal.SourceMTime = time.Time{}
headers := map[string]string{}
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
headers[name] = values[0]
}
}
restore()
// --- Destination: no compression, no default encryption. ---
restoreDst := disableCompression()
defer restoreDst()
replNewReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, replicator.AccessKey, replicator.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
replNewRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replNewRec, replNewReq)
if replNewRec.Code != http.StatusOK {
t.Fatalf("replica NewMultipart status %d: %s", replNewRec.Code, replNewRec.Body.String())
}
var replicaInit InitiateMultipartUploadResponse
if err = xmlDecoder(replNewRec.Body, &replicaInit, int64(replNewRec.Body.Len())); err != nil {
t.Fatal(err)
}
replPartReq, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, replicaInit.UploadID, "1"),
int64(len(rawPart)), bytes.NewReader(rawPart), replicator.AccessKey, replicator.SecretKey,
map[string]string{xhttp.MinIOSourceReplicationRequest: "true"})
if err != nil {
t.Fatal(err)
}
replPartRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replPartRec, replPartReq)
if replPartRec.Code != http.StatusOK {
t.Fatalf("replica PutPart status %d: %s", replPartRec.Code, replPartRec.Body.String())
}
replCompleteBody, err := xml.Marshal(CompleteMultipartUpload{Parts: []CompletePart{
{PartNumber: 1, ETag: canonicalizeETag(replPartRec.Header()[xhttp.ETag][0])},
}})
if err != nil {
t.Fatal(err)
}
replCompleteReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, replicaInit.UploadID),
int64(len(replCompleteBody)), bytes.NewReader(replCompleteBody), replicator.AccessKey, replicator.SecretKey,
map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.MinIOSourceMTime: sourceInfo.ModTime.Format(time.RFC3339Nano),
xhttp.MinIOSourceETag: sourceInfo.ETag,
xhttp.MinIOReplicationActualObjectSize: strconv.FormatInt(actualSize, 10),
})
if err != nil {
t.Fatal(err)
}
replCompleteRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replCompleteRec, replCompleteReq)
if replCompleteRec.Code != http.StatusOK {
t.Fatalf("replica Complete status %d: %s", replCompleteRec.Code, replCompleteRec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
reportedActual, aerr := info.GetActualSize()
t.Logf("replica mpu: stored size=%d compression=%q actual-size=%q GetActualSize=%d(err=%v)",
info.Size, info.UserDefined[ReservedMetadataPrefix+"compression"],
info.UserDefined[ReservedMetadataPrefix+"actual-size"], reportedActual, aerr)
assertSSECPlaintext(t, apiRouter, credentials, bucketName, object, sseHeaders, data)
})
}
// TestAPISSECCompressionProducerMatrix pins the scope of the exclusion across
// the PutObject and NewMultipartUpload producers: SSE-C is never compressed,
// while plaintext, SSE-S3 and SSE-KMS keep following allow_encryption.
func TestAPISSECCompressionProducerMatrix(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionProducerMatrix,
})
}
func testAPISSECCompressionProducerMatrix(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
ssecHeaders := ssecTestHeaders(0x5a)
sseS3Headers := map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}
sseKMSHeaders := map[string]string{
xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS,
xhttp.AmzServerSideEncryptionKmsID: "compressed-ssec-producer-matrix",
}
previousKMS := GlobalKMS
GlobalKMS = kms.NewStub("compressed-ssec-producer-matrix")
defer func() { GlobalKMS = previousKMS }()
big := bytes.Repeat([]byte("silo producer matrix payload "), 8192)
small := bytes.Repeat([]byte("s"), 1024) // below minCompressibleSize
for _, tc := range []struct {
name string
allowEncrypted bool
headers map[string]string
body []byte
wantCompressed bool
}{
// SSE-C is excluded from compression in both configurations, because the
// replication wire cannot carry the compression state.
{"ssec+allow_encryption-on+large", true, ssecHeaders, big, false},
{"ssec+allow_encryption-on+small", true, ssecHeaders, small, false},
{"ssec+allow_encryption-off+large", false, ssecHeaders, big, false},
// Plaintext still compresses in both configurations.
{"plain+allow_encryption-on+large", true, nil, big, true},
{"plain+allow_encryption-off+large", false, nil, big, true},
// SSE-S3 and SSE-KMS are the reason allow_encryption exists: the server
// owns the key, so the source decompresses before replicating.
{"sse-s3+allow_encryption-on+large", true, sseS3Headers, big, true},
{"sse-s3+allow_encryption-off+large", false, sseS3Headers, big, false},
{"sse-kms+allow_encryption-on+large", true, sseKMSHeaders, big, true},
{"sse-kms+allow_encryption-off+large", false, sseKMSHeaders, big, false},
} {
t.Run(instanceType+"/put/"+tc.name, func(t *testing.T) {
restore := setCopyChecksumCompression(tc.allowEncrypted)
defer restore()
object := "producer/" + tc.name + ".txt"
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(tc.body)), bytes.NewReader(tc.body), credentials.AccessKey, credentials.SecretKey, tc.headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("PUT status %d, want %d: %s", rec.Code, http.StatusOK, rec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
_, compressed := info.UserDefined[ReservedMetadataPrefix+"compression"]
if compressed != tc.wantCompressed {
t.Errorf("compressed=%v, want %v (userDefined=%v)", compressed, tc.wantCompressed, info.UserDefined)
}
})
}
// NewMultipartUpload has no size gate, so the exclusion turns on the SSE-C
// and allow_encryption combination alone.
for _, tc := range []struct {
name string
allowEncrypted bool
headers map[string]string
wantCompressed bool
}{
{"ssec+allow_encryption-on", true, ssecHeaders, false},
{"ssec+allow_encryption-off", false, ssecHeaders, false},
{"plain+allow_encryption-on", true, nil, true},
} {
t.Run(instanceType+"/mpu/"+tc.name, func(t *testing.T) {
restore := setCopyChecksumCompression(tc.allowEncrypted)
defer restore()
object := "producer/mpu-" + tc.name + ".txt"
req, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, tc.headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("NewMultipartUpload status %d, want %d: %s", rec.Code, http.StatusOK, rec.Body.String())
}
var init InitiateMultipartUploadResponse
if err = xmlDecoder(rec.Body, &init, int64(rec.Body.Len())); err != nil {
t.Fatal(err)
}
mi, err := obj.GetMultipartInfo(t.Context(), bucketName, object, init.UploadID, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
_, compressed := mi.UserDefined[ReservedMetadataPrefix+"compression"]
if compressed != tc.wantCompressed {
t.Errorf("upload compressed=%v, want %v", compressed, tc.wantCompressed)
}
})
}
}
// TestAPISSECCompressionSkippedOnCopyObject covers the third producer:
// CopyObjectHandler decides compression before it encrypts, so a copy with a
// destination customer key and allow_encryption=on used to store a compressed
// SSE-C object from a plaintext source. It also covers the reverse direction,
// where only copy-source customer headers are present and compression must
// still apply.
func TestAPISSECCompressionSkippedOnCopyObject(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionSkippedOnCopyObject,
endpoints: []string{"CopyObject", "PutObject", "GetObject"},
})
}
func testAPISSECCompressionSkippedOnCopyObject(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
ssecHeaders := ssecTestHeaders(0x7c)
data := bytes.Repeat([]byte("copy object compressed ssec payload "), 8192)
restore := setCopyChecksumCompression(true)
defer restore()
// An unencrypted source, stored compressed because compression is on. Only
// the copy adds encryption, so only the copy can change the decision.
src := "copysrc/plain.txt"
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, src),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("source PUT status %d: %s", rec.Code, rec.Body.String())
}
dst := "copydst/ssec.txt"
copyHeaders := map[string]string{"X-Amz-Copy-Source": SlashSeparator + bucketName + SlashSeparator + src}
for k, v := range ssecHeaders {
copyHeaders[k] = v
}
copyReq, err := newTestSignedRequestV4(http.MethodPut, getCopyObjectURL("", bucketName, dst),
0, nil, credentials.AccessKey, credentials.SecretKey, copyHeaders)
if err != nil {
t.Fatal(err)
}
copyRec := httptest.NewRecorder()
apiRouter.ServeHTTP(copyRec, copyReq)
if copyRec.Code != http.StatusOK {
t.Fatalf("CopyObject status %d: %s", copyRec.Code, copyRec.Body.String())
}
assertStoredSSECUncompressed(t, obj, bucketName, dst)
assertSSECPlaintext(t, apiRouter, credentials, bucketName, dst, ssecHeaders, data)
// The plaintext source is untouched by the copy and stays compressed.
srcInfo, err := obj.GetObjectInfo(t.Context(), bucketName, src, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if _, compressed := srcInfo.UserDefined[ReservedMetadataPrefix+"compression"]; !compressed {
t.Errorf("the plaintext copy source lost compression (userDefined=%v)", srcInfo.UserDefined)
}
// The reverse direction: a copy-source customer key is not a destination
// key, so copying the SSE-C object on to a plaintext destination still
// compresses. crypto.SSEC.IsRequested ignores the copy-source headers.
plain := "copydst/decrypted.txt"
decryptHeaders := map[string]string{
"X-Amz-Copy-Source": SlashSeparator + bucketName + SlashSeparator + dst,
xhttp.AmzServerSideEncryptionCopyCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCopyCustomerKey: ssecHeaders[xhttp.AmzServerSideEncryptionCustomerKey],
xhttp.AmzServerSideEncryptionCopyCustomerKeyMD5: ssecHeaders[xhttp.AmzServerSideEncryptionCustomerKeyMD5],
}
decryptReq, err := newTestSignedRequestV4(http.MethodPut, getCopyObjectURL("", bucketName, plain),
0, nil, credentials.AccessKey, credentials.SecretKey, decryptHeaders)
if err != nil {
t.Fatal(err)
}
decryptRec := httptest.NewRecorder()
apiRouter.ServeHTTP(decryptRec, decryptReq)
if decryptRec.Code != http.StatusOK {
t.Fatalf("CopyObject to a plaintext destination status %d: %s", decryptRec.Code, decryptRec.Body.String())
}
plainInfo, err := obj.GetObjectInfo(t.Context(), bucketName, plain, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if _, sealed := plainInfo.UserDefined[crypto.MetaSealedKeySSEC]; sealed {
t.Fatalf("%s is still SSE-C sealed, the fixture proves nothing", plain)
}
if _, compressed := plainInfo.UserDefined[ReservedMetadataPrefix+"compression"]; !compressed {
t.Errorf("a copy carrying only copy-source SSE-C headers was not compressed (userDefined=%v)", plainInfo.UserDefined)
}
}
// TestAPISSECCompressionSkippedOnSnowballExtract covers the fourth producer:
// PutObjectExtractHandler decides compression per entry before it encrypts, so
// a tar extract carrying customer key headers used to store every entry as a
// compressed SSE-C object.
func TestAPISSECCompressionSkippedOnSnowballExtract(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECCompressionSkippedOnSnowballExtract,
})
}
func testAPISSECCompressionSkippedOnSnowballExtract(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
entry := "extracted/entry.txt"
payload := bytes.Repeat([]byte("snowball compressed ssec entry "), 4096)
var body bytes.Buffer
tw := tar.NewWriter(&body)
if err := tw.WriteHeader(&tar.Header{Name: entry, Mode: 0o600, Size: int64(len(payload))}); err != nil {
t.Fatal(err)
}
if _, err := tw.Write(payload); err != nil {
t.Fatal(err)
}
if err := tw.Close(); err != nil {
t.Fatal(err)
}
restore := setCopyChecksumCompression(true)
defer restore()
ssecHeaders := ssecTestHeaders(0x2d)
headers := map[string]string{xhttp.AmzSnowballExtract: "true"}
for k, v := range ssecHeaders {
headers[k] = v
}
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, "snowball.tar"),
int64(body.Len()), bytes.NewReader(body.Bytes()), credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("snowball extract status %d: %s", rec.Code, rec.Body.String())
}
assertStoredSSECUncompressed(t, obj, bucketName, entry)
assertSSECPlaintext(t, apiRouter, credentials, bucketName, entry, ssecHeaders, payload)
}
// TestSSECBatchReplicationCannotRead is the control for the corruption path:
// batch replication reads without ReplicationRequest, so NoDecryption is never
// set and a non-empty SSE-C source fails at read time. Batch replication
// therefore cannot reach the replica shape in
// TestAPISSECCompressionReplicaStaysReadable; it cannot replicate a non-empty
// SSE-C object at all, compressed or not. A zero-byte object takes the reader
// shortcut, whose key check passes without a customer key.
func TestSSECBatchReplicationCannotRead(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testSSECBatchReplicationCannotRead,
})
}
func testSSECBatchReplicationCannotRead(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
ssecHeaders := ssecTestHeaders(0x6b)
data := bytes.Repeat([]byte("batch ssec payload "), 8192)
object := "batch/ssec-plain.txt"
restore := disableCompression()
defer restore()
req, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object),
int64(len(data)), bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, ssecHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("source PUT status %d: %s", rec.Code, rec.Body.String())
}
// The read shape used by BatchJobReplicateV1.ReplicateToTarget and
// writeAsArchive: no ReplicationRequest, so NoDecryption is never set.
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{}, ObjectOptions{})
if err == nil {
gr.Close()
t.Fatal("batch-shaped read of an SSE-C object unexpectedly succeeded")
}
t.Logf("batch-shaped read of an SSE-C object fails as expected: %v", err)
// Control: the replication worker's read shape succeeds and yields ciphertext.
gr2, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{ReplicationRequest: true})
if err != nil {
t.Fatalf("replication-shaped read failed: %v", err)
}
raw, err := io.ReadAll(gr2)
gr2.Close()
if err != nil {
t.Fatal(err)
}
if bytes.Equal(raw, data) {
t.Fatal("replication-shaped read returned plaintext")
}
t.Logf("replication-shaped read returns %d bytes of ciphertext (plaintext %d)", len(raw), len(data))
}
+4 -1
View File
@@ -227,7 +227,7 @@ func initHelp() {
},
config.HelpKV{
Key: config.ILMSubSys,
Description: "manage ILM settings for expiration and transition workers",
Description: "manage ILM settings for expiration, transition, and access-tier workers",
Optional: true,
},
}
@@ -704,6 +704,9 @@ func applyDynamicConfigForSubSys(ctx context.Context, objAPI ObjectLayer, s conf
if globalExpiryState != nil {
globalExpiryState.ResizeWorkers(ilmCfg.ExpirationWorkers)
}
if globalAccessTierState != nil {
globalAccessTierState.UpdateWorkers(ilmCfg.AccessWorkers)
}
globalILMConfig.update(ilmCfg)
}
}
+14
View File
@@ -899,6 +899,7 @@ type scannerItem struct {
objectName string // Only the object name without prefixes.
replication replicationConfig
lifeCycle *lifecycle.Lifecycle
poolIdx int
Typ fs.FileMode
heal struct {
enabled bool
@@ -909,6 +910,7 @@ type scannerItem struct {
type sizeSummary struct {
totalSize int64
hotTierSize int64
versions uint64
deleteMarkers uint64
replicatedSize int64
@@ -1159,6 +1161,14 @@ eventLoop:
globalExpiryState.enqueueNoncurrentVersions(i.bucket, toDel, noncurrentEvents)
}
i.alertExcessiveVersions(remainingVersions, cumulativeSize)
if globalILMConfig.accessTieringEnabled() {
for idx, oi := range objInfos {
if oi.IsLatest && events[idx].Action == lifecycle.NoneAction {
applyAccessTransition(ctx, i, oi)
break
}
}
}
}
func evalActionFromLifecycle(ctx context.Context, lc lifecycle.Lifecycle, lr lock.Retention, rcfg *replication.Config, obj ObjectInfo) lifecycle.Event {
@@ -1474,6 +1484,8 @@ const (
ILMFreeVersionDelete = "ilm:free-version-delete"
// ILMTransition - audit trail for ILM transitioning.
ILMTransition = " ilm:transition"
// ILMAccessTier - audit trail for moving objects between server pools.
ILMAccessTier = "ilm:access-tier"
)
func auditLogLifecycle(ctx context.Context, oi ObjectInfo, event string, tags map[string]string, traceFn func(event string, metadata map[string]string, err error)) {
@@ -1485,6 +1497,8 @@ func auditLogLifecycle(ctx context.Context, oi ObjectInfo, event string, tags ma
apiName = "ILMFreeVersionDelete"
case ILMTransition:
apiName = "ILMTransition"
case ILMAccessTier:
apiName = "ILMAccessTier"
}
auditLogInternal(ctx, AuditLogOptions{
Event: event,
+60 -5
View File
@@ -61,6 +61,7 @@ type dataUsageEntry struct {
Children dataUsageHashMap `msg:"ch"`
// These fields do no include any children.
Size int64 `msg:"sz"`
HotTierSize int64 `msg:"hts"`
Objects uint64 `msg:"os"`
Versions uint64 `msg:"vs"` // Versions that are not delete markers.
DeleteMarkers uint64 `msg:"dms"`
@@ -134,8 +135,8 @@ func (ts tierStats) add(u tierStats) tierStats {
}
}
//msgp:encode ignore dataUsageEntryV2 dataUsageEntryV3 dataUsageEntryV4 dataUsageEntryV5 dataUsageEntryV6 dataUsageEntryV7
//msgp:marshal ignore dataUsageEntryV2 dataUsageEntryV3 dataUsageEntryV4 dataUsageEntryV5 dataUsageEntryV6 dataUsageEntryV7
//msgp:encode ignore dataUsageEntryV2 dataUsageEntryV3 dataUsageEntryV4 dataUsageEntryV5 dataUsageEntryV6 dataUsageEntryV7 dataUsageEntryV8
//msgp:marshal ignore dataUsageEntryV2 dataUsageEntryV3 dataUsageEntryV4 dataUsageEntryV5 dataUsageEntryV6 dataUsageEntryV7 dataUsageEntryV8
//msgp:tuple dataUsageEntryV2
type dataUsageEntryV2 struct {
@@ -199,14 +200,30 @@ type dataUsageEntryV7 struct {
Compacted bool `msg:"c"`
}
// dataUsageEntryV8 is the on-disk shape before access-tier accounting was
// introduced. Keep it so caches written by the previous release decode
// without being discarded.
type dataUsageEntryV8 struct {
Children dataUsageHashMap `msg:"ch"`
// These fields do no include any children.
Size int64 `msg:"sz"`
Objects uint64 `msg:"os"`
Versions uint64 `msg:"vs"`
DeleteMarkers uint64 `msg:"dms"`
ObjSizes sizeHistogram `msg:"szs"`
ObjVersions versionsHistogram `msg:"vh"`
AllTierStats *allTierStats `msg:"ats,omitempty"`
Compacted bool `msg:"c"`
}
// dataUsageCache contains a cache of data usage entries latest version.
type dataUsageCache struct {
Info dataUsageCacheInfo
Cache map[string]dataUsageEntry
}
//msgp:encode ignore dataUsageCacheV2 dataUsageCacheV3 dataUsageCacheV4 dataUsageCacheV5 dataUsageCacheV6 dataUsageCacheV7
//msgp:marshal ignore dataUsageCacheV2 dataUsageCacheV3 dataUsageCacheV4 dataUsageCacheV5 dataUsageCacheV6 dataUsageCacheV7
//msgp:encode ignore dataUsageCacheV2 dataUsageCacheV3 dataUsageCacheV4 dataUsageCacheV5 dataUsageCacheV6 dataUsageCacheV7 dataUsageCacheV8
//msgp:marshal ignore dataUsageCacheV2 dataUsageCacheV3 dataUsageCacheV4 dataUsageCacheV5 dataUsageCacheV6 dataUsageCacheV7 dataUsageCacheV8
// dataUsageCacheV2 contains a cache of data usage entries version 2.
type dataUsageCacheV2 struct {
@@ -244,6 +261,12 @@ type dataUsageCacheV7 struct {
Cache map[string]dataUsageEntryV7
}
// dataUsageCacheV8 contains a cache of data usage entries version 8.
type dataUsageCacheV8 struct {
Info dataUsageCacheInfo
Cache map[string]dataUsageEntryV8
}
//msgp:ignore dataUsageEntryInfo
type dataUsageEntryInfo struct {
Name string
@@ -272,6 +295,7 @@ type dataUsageCacheInfo struct {
func (e *dataUsageEntry) addSizes(summary sizeSummary) {
e.Size += summary.totalSize
e.HotTierSize += summary.hotTierSize
e.Versions += summary.versions
e.DeleteMarkers += summary.deleteMarkers
e.ObjSizes.add(summary.totalSize)
@@ -291,6 +315,7 @@ func (e *dataUsageEntry) merge(other dataUsageEntry) {
e.Versions += other.Versions
e.DeleteMarkers += other.DeleteMarkers
e.Size += other.Size
e.HotTierSize += other.HotTierSize
for i, v := range other.ObjSizes[:] {
e.ObjSizes[i] += v
@@ -431,6 +456,7 @@ func (d *dataUsageCache) dui(path string, buckets []BucketInfo) DataUsageInfo {
flat := d.flatten(*e)
dui := DataUsageInfo{
LastUpdate: d.Info.LastUpdate,
ScannerCycle: d.Info.NextCycle,
ObjectsTotalCount: flat.Objects,
VersionsTotalCount: flat.Versions,
DeleteMarkersTotalCount: flat.DeleteMarkers,
@@ -781,6 +807,7 @@ func (d *dataUsageCache) bucketsUsageInfo(buckets []BucketInfo) map[string]Bucke
flat := d.flatten(*e)
bui := BucketUsageInfo{
Size: uint64(flat.Size),
HotTierSize: uint64(max(flat.HotTierSize, 0)),
VersionsCount: flat.Versions,
ObjectsCount: flat.Objects,
DeleteMarkersCount: flat.DeleteMarkers,
@@ -980,7 +1007,8 @@ func (d *dataUsageCache) save(ctx context.Context, store objectIO, name string)
// Bumping the cache version will drop data from previous versions
// and write new data with the new version.
const (
dataUsageCacheVerCurrent = 8
dataUsageCacheVerCurrent = 9
dataUsageCacheVerV8 = 8
dataUsageCacheVerV7 = 7
dataUsageCacheVerV6 = 6
dataUsageCacheVerV5 = 5
@@ -1181,6 +1209,33 @@ func (d *dataUsageCache) deserialize(r io.Reader) error {
}
}
return nil
case dataUsageCacheVerV8:
// Zstd compressed.
dec, err := zstd.NewReader(r, zstd.WithDecoderConcurrency(2))
if err != nil {
return err
}
defer dec.Close()
dold := &dataUsageCacheV8{}
if err = dold.DecodeMsg(msgp.NewReader(dec)); err != nil {
return err
}
d.Info = dold.Info
d.Cache = make(map[string]dataUsageEntry, len(dold.Cache))
for k, v := range dold.Cache {
d.Cache[k] = dataUsageEntry{
Children: v.Children,
Size: v.Size,
Objects: v.Objects,
Versions: v.Versions,
DeleteMarkers: v.DeleteMarkers,
ObjSizes: v.ObjSizes,
ObjVersions: v.ObjVersions,
AllTierStats: v.AllTierStats,
Compacted: v.Compacted,
}
}
return nil
case dataUsageCacheVerCurrent:
// Zstd compressed.
+605 -9
View File
@@ -1591,6 +1591,145 @@ func (z *dataUsageCacheV7) Msgsize() (s int) {
return
}
// DecodeMsg implements msgp.Decodable
func (z *dataUsageCacheV8) DecodeMsg(dc *msgp.Reader) (err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "Info":
err = z.Info.DecodeMsg(dc)
if err != nil {
err = msgp.WrapError(err, "Info")
return
}
case "Cache":
var zb0002 uint32
zb0002, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err, "Cache")
return
}
if z.Cache == nil {
z.Cache = make(map[string]dataUsageEntryV8, zb0002)
} else if len(z.Cache) > 0 {
clear(z.Cache)
}
for zb0002 > 0 {
zb0002--
var za0001 string
za0001, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "Cache")
return
}
var za0002 dataUsageEntryV8
err = za0002.DecodeMsg(dc)
if err != nil {
err = msgp.WrapError(err, "Cache", za0001)
return
}
z.Cache[za0001] = za0002
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
return
}
// UnmarshalMsg implements msgp.Unmarshaler
func (z *dataUsageCacheV8) UnmarshalMsg(bts []byte) (o []byte, err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "Info":
bts, err = z.Info.UnmarshalMsg(bts)
if err != nil {
err = msgp.WrapError(err, "Info")
return
}
case "Cache":
var zb0002 uint32
zb0002, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Cache")
return
}
if z.Cache == nil {
z.Cache = make(map[string]dataUsageEntryV8, zb0002)
} else if len(z.Cache) > 0 {
clear(z.Cache)
}
for zb0002 > 0 {
var za0002 dataUsageEntryV8
zb0002--
var za0001 string
za0001, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Cache")
return
}
bts, err = za0002.UnmarshalMsg(bts)
if err != nil {
err = msgp.WrapError(err, "Cache", za0001)
return
}
z.Cache[za0001] = za0002
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
o = bts
return
}
// Msgsize returns an upper bound estimate of the number of bytes occupied by the serialized message
func (z *dataUsageCacheV8) Msgsize() (s int) {
s = 1 + 5 + z.Info.Msgsize() + 6 + msgp.MapHeaderSize
if z.Cache != nil {
for za0001, za0002 := range z.Cache {
_ = za0002
s += msgp.StringPrefixSize + len(za0001) + za0002.Msgsize()
}
}
return
}
// DecodeMsg implements msgp.Decodable
func (z *dataUsageEntry) DecodeMsg(dc *msgp.Reader) (err error) {
var field []byte
@@ -1623,6 +1762,12 @@ func (z *dataUsageEntry) DecodeMsg(dc *msgp.Reader) (err error) {
err = msgp.WrapError(err, "Size")
return
}
case "hts":
z.HotTierSize, err = dc.ReadInt64()
if err != nil {
err = msgp.WrapError(err, "HotTierSize")
return
}
case "os":
z.Objects, err = dc.ReadUint64()
if err != nil {
@@ -1801,12 +1946,12 @@ func (z *dataUsageEntry) DecodeMsg(dc *msgp.Reader) (err error) {
// EncodeMsg implements msgp.Encodable
func (z *dataUsageEntry) EncodeMsg(en *msgp.Writer) (err error) {
// check for omitted fields
zb0001Len := uint32(9)
var zb0001Mask uint16 /* 9 bits */
zb0001Len := uint32(10)
var zb0001Mask uint16 /* 10 bits */
_ = zb0001Mask
if z.AllTierStats == nil {
zb0001Len--
zb0001Mask |= 0x80
zb0001Mask |= 0x100
}
// variable map header, size zb0001Len
err = en.Append(0x80 | uint8(zb0001Len))
@@ -1836,6 +1981,16 @@ func (z *dataUsageEntry) EncodeMsg(en *msgp.Writer) (err error) {
err = msgp.WrapError(err, "Size")
return
}
// write "hts"
err = en.Append(0xa3, 0x68, 0x74, 0x73)
if err != nil {
return
}
err = en.WriteInt64(z.HotTierSize)
if err != nil {
err = msgp.WrapError(err, "HotTierSize")
return
}
// write "os"
err = en.Append(0xa2, 0x6f, 0x73)
if err != nil {
@@ -1900,7 +2055,7 @@ func (z *dataUsageEntry) EncodeMsg(en *msgp.Writer) (err error) {
return
}
}
if (zb0001Mask & 0x80) == 0 { // if not omitted
if (zb0001Mask & 0x100) == 0 { // if not omitted
// write "ats"
err = en.Append(0xa3, 0x61, 0x74, 0x73)
if err != nil {
@@ -1981,12 +2136,12 @@ func (z *dataUsageEntry) EncodeMsg(en *msgp.Writer) (err error) {
func (z *dataUsageEntry) MarshalMsg(b []byte) (o []byte, err error) {
o = msgp.Require(b, z.Msgsize())
// check for omitted fields
zb0001Len := uint32(9)
var zb0001Mask uint16 /* 9 bits */
zb0001Len := uint32(10)
var zb0001Mask uint16 /* 10 bits */
_ = zb0001Mask
if z.AllTierStats == nil {
zb0001Len--
zb0001Mask |= 0x80
zb0001Mask |= 0x100
}
// variable map header, size zb0001Len
o = append(o, 0x80|uint8(zb0001Len))
@@ -2003,6 +2158,9 @@ func (z *dataUsageEntry) MarshalMsg(b []byte) (o []byte, err error) {
// string "sz"
o = append(o, 0xa2, 0x73, 0x7a)
o = msgp.AppendInt64(o, z.Size)
// string "hts"
o = append(o, 0xa3, 0x68, 0x74, 0x73)
o = msgp.AppendInt64(o, z.HotTierSize)
// string "os"
o = append(o, 0xa2, 0x6f, 0x73)
o = msgp.AppendUint64(o, z.Objects)
@@ -2024,7 +2182,7 @@ func (z *dataUsageEntry) MarshalMsg(b []byte) (o []byte, err error) {
for za0002 := range z.ObjVersions {
o = msgp.AppendUint64(o, z.ObjVersions[za0002])
}
if (zb0001Mask & 0x80) == 0 { // if not omitted
if (zb0001Mask & 0x100) == 0 { // if not omitted
// string "ats"
o = append(o, 0xa3, 0x61, 0x74, 0x73)
if z.AllTierStats == nil {
@@ -2088,6 +2246,12 @@ func (z *dataUsageEntry) UnmarshalMsg(bts []byte) (o []byte, err error) {
err = msgp.WrapError(err, "Size")
return
}
case "hts":
z.HotTierSize, bts, err = msgp.ReadInt64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "HotTierSize")
return
}
case "os":
z.Objects, bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
@@ -2265,7 +2429,7 @@ func (z *dataUsageEntry) UnmarshalMsg(bts []byte) (o []byte, err error) {
// Msgsize returns an upper bound estimate of the number of bytes occupied by the serialized message
func (z *dataUsageEntry) Msgsize() (s int) {
s = 1 + 3 + z.Children.Msgsize() + 3 + msgp.Int64Size + 3 + msgp.Uint64Size + 3 + msgp.Uint64Size + 4 + msgp.Uint64Size + 4 + msgp.ArrayHeaderSize + (dataUsageBucketLen * (msgp.Uint64Size)) + 3 + msgp.ArrayHeaderSize + (dataUsageVersionLen * (msgp.Uint64Size)) + 4
s = 1 + 3 + z.Children.Msgsize() + 3 + msgp.Int64Size + 4 + msgp.Int64Size + 3 + msgp.Uint64Size + 3 + msgp.Uint64Size + 4 + msgp.Uint64Size + 4 + msgp.ArrayHeaderSize + (dataUsageBucketLen * (msgp.Uint64Size)) + 3 + msgp.ArrayHeaderSize + (dataUsageVersionLen * (msgp.Uint64Size)) + 4
if z.AllTierStats == nil {
s += msgp.NilSize
} else {
@@ -3258,6 +3422,438 @@ func (z *dataUsageEntryV7) Msgsize() (s int) {
return
}
// DecodeMsg implements msgp.Decodable
func (z *dataUsageEntryV8) DecodeMsg(dc *msgp.Reader) (err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err)
return
}
var zb0001Mask uint8 /* 1 bits */
_ = zb0001Mask
for zb0001 > 0 {
zb0001--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "ch":
err = z.Children.DecodeMsg(dc)
if err != nil {
err = msgp.WrapError(err, "Children")
return
}
case "sz":
z.Size, err = dc.ReadInt64()
if err != nil {
err = msgp.WrapError(err, "Size")
return
}
case "os":
z.Objects, err = dc.ReadUint64()
if err != nil {
err = msgp.WrapError(err, "Objects")
return
}
case "vs":
z.Versions, err = dc.ReadUint64()
if err != nil {
err = msgp.WrapError(err, "Versions")
return
}
case "dms":
z.DeleteMarkers, err = dc.ReadUint64()
if err != nil {
err = msgp.WrapError(err, "DeleteMarkers")
return
}
case "szs":
var zb0002 uint32
zb0002, err = dc.ReadArrayHeader()
if err != nil {
err = msgp.WrapError(err, "ObjSizes")
return
}
if zb0002 != uint32(dataUsageBucketLen) {
err = msgp.ArrayError{Wanted: uint32(dataUsageBucketLen), Got: zb0002}
return
}
for za0001 := range z.ObjSizes {
z.ObjSizes[za0001], err = dc.ReadUint64()
if err != nil {
err = msgp.WrapError(err, "ObjSizes", za0001)
return
}
}
case "vh":
var zb0003 uint32
zb0003, err = dc.ReadArrayHeader()
if err != nil {
err = msgp.WrapError(err, "ObjVersions")
return
}
if zb0003 != uint32(dataUsageVersionLen) {
err = msgp.ArrayError{Wanted: uint32(dataUsageVersionLen), Got: zb0003}
return
}
for za0002 := range z.ObjVersions {
z.ObjVersions[za0002], err = dc.ReadUint64()
if err != nil {
err = msgp.WrapError(err, "ObjVersions", za0002)
return
}
}
case "ats":
if dc.IsNil() {
err = dc.ReadNil()
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
z.AllTierStats = nil
} else {
if z.AllTierStats == nil {
z.AllTierStats = new(allTierStats)
}
var zb0004 uint32
zb0004, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
for zb0004 > 0 {
zb0004--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
switch msgp.UnsafeString(field) {
case "ts":
var zb0005 uint32
zb0005, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers")
return
}
if z.AllTierStats.Tiers == nil {
z.AllTierStats.Tiers = make(map[string]tierStats, zb0005)
} else if len(z.AllTierStats.Tiers) > 0 {
clear(z.AllTierStats.Tiers)
}
for zb0005 > 0 {
zb0005--
var za0003 string
za0003, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers")
return
}
var za0004 tierStats
var zb0006 uint32
zb0006, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003)
return
}
for zb0006 > 0 {
zb0006--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003)
return
}
switch msgp.UnsafeString(field) {
case "ts":
za0004.TotalSize, err = dc.ReadUint64()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003, "TotalSize")
return
}
case "nv":
za0004.NumVersions, err = dc.ReadInt()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003, "NumVersions")
return
}
case "no":
za0004.NumObjects, err = dc.ReadInt()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003, "NumObjects")
return
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003)
return
}
}
}
z.AllTierStats.Tiers[za0003] = za0004
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
}
}
}
zb0001Mask |= 0x1
case "c":
z.Compacted, err = dc.ReadBool()
if err != nil {
err = msgp.WrapError(err, "Compacted")
return
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
// Clear omitted fields.
if (zb0001Mask & 0x1) == 0 {
z.AllTierStats = nil
}
return
}
// UnmarshalMsg implements msgp.Unmarshaler
func (z *dataUsageEntryV8) UnmarshalMsg(bts []byte) (o []byte, err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
var zb0001Mask uint8 /* 1 bits */
_ = zb0001Mask
for zb0001 > 0 {
zb0001--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "ch":
bts, err = z.Children.UnmarshalMsg(bts)
if err != nil {
err = msgp.WrapError(err, "Children")
return
}
case "sz":
z.Size, bts, err = msgp.ReadInt64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "Size")
return
}
case "os":
z.Objects, bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "Objects")
return
}
case "vs":
z.Versions, bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "Versions")
return
}
case "dms":
z.DeleteMarkers, bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "DeleteMarkers")
return
}
case "szs":
var zb0002 uint32
zb0002, bts, err = msgp.ReadArrayHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "ObjSizes")
return
}
if zb0002 != uint32(dataUsageBucketLen) {
err = msgp.ArrayError{Wanted: uint32(dataUsageBucketLen), Got: zb0002}
return
}
for za0001 := range z.ObjSizes {
z.ObjSizes[za0001], bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "ObjSizes", za0001)
return
}
}
case "vh":
var zb0003 uint32
zb0003, bts, err = msgp.ReadArrayHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "ObjVersions")
return
}
if zb0003 != uint32(dataUsageVersionLen) {
err = msgp.ArrayError{Wanted: uint32(dataUsageVersionLen), Got: zb0003}
return
}
for za0002 := range z.ObjVersions {
z.ObjVersions[za0002], bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "ObjVersions", za0002)
return
}
}
case "ats":
if msgp.IsNil(bts) {
bts, err = msgp.ReadNilBytes(bts)
if err != nil {
return
}
z.AllTierStats = nil
} else {
if z.AllTierStats == nil {
z.AllTierStats = new(allTierStats)
}
var zb0004 uint32
zb0004, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
for zb0004 > 0 {
zb0004--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
switch msgp.UnsafeString(field) {
case "ts":
var zb0005 uint32
zb0005, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers")
return
}
if z.AllTierStats.Tiers == nil {
z.AllTierStats.Tiers = make(map[string]tierStats, zb0005)
} else if len(z.AllTierStats.Tiers) > 0 {
clear(z.AllTierStats.Tiers)
}
for zb0005 > 0 {
var za0004 tierStats
zb0005--
var za0003 string
za0003, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers")
return
}
var zb0006 uint32
zb0006, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003)
return
}
for zb0006 > 0 {
zb0006--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003)
return
}
switch msgp.UnsafeString(field) {
case "ts":
za0004.TotalSize, bts, err = msgp.ReadUint64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003, "TotalSize")
return
}
case "nv":
za0004.NumVersions, bts, err = msgp.ReadIntBytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003, "NumVersions")
return
}
case "no":
za0004.NumObjects, bts, err = msgp.ReadIntBytes(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003, "NumObjects")
return
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats", "Tiers", za0003)
return
}
}
}
z.AllTierStats.Tiers[za0003] = za0004
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err, "AllTierStats")
return
}
}
}
}
zb0001Mask |= 0x1
case "c":
z.Compacted, bts, err = msgp.ReadBoolBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Compacted")
return
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
// Clear omitted fields.
if (zb0001Mask & 0x1) == 0 {
z.AllTierStats = nil
}
o = bts
return
}
// Msgsize returns an upper bound estimate of the number of bytes occupied by the serialized message
func (z *dataUsageEntryV8) Msgsize() (s int) {
s = 1 + 3 + z.Children.Msgsize() + 3 + msgp.Int64Size + 3 + msgp.Uint64Size + 3 + msgp.Uint64Size + 4 + msgp.Uint64Size + 4 + msgp.ArrayHeaderSize + (dataUsageBucketLen * (msgp.Uint64Size)) + 3 + msgp.ArrayHeaderSize + (dataUsageVersionLen * (msgp.Uint64Size)) + 4
if z.AllTierStats == nil {
s += msgp.NilSize
} else {
s += 1 + 3 + msgp.MapHeaderSize
if z.AllTierStats.Tiers != nil {
for za0003, za0004 := range z.AllTierStats.Tiers {
_ = za0004
s += msgp.StringPrefixSize + len(za0003) + 1 + 3 + msgp.Uint64Size + 3 + msgp.IntSize + 3 + msgp.IntSize
}
}
}
s += 2 + msgp.BoolSize
return
}
// DecodeMsg implements msgp.Decodable
func (z *dataUsageHash) DecodeMsg(dc *msgp.Reader) (err error) {
{
+6 -1
View File
@@ -46,7 +46,8 @@ type BucketTargetUsageInfo struct {
// - total objects in a bucket
// - object size histogram per bucket
type BucketUsageInfo struct {
Size uint64 `json:"size"`
Size uint64 `json:"size"`
HotTierSize uint64 `json:"hotTierSize,omitempty"`
// Following five fields suffixed with V1 are here for backward compatibility
// Total Size for objects that have not yet been replicated
ReplicationPendingSizeV1 uint64 `json:"objectsPendingReplicationTotalSize"`
@@ -78,6 +79,10 @@ type DataUsageInfo struct {
// LastUpdate is the timestamp of when the data usage info was last updated.
// This does not indicate a full scan.
LastUpdate time.Time `json:"lastUpdate"`
// ScannerCycle changes only after a complete scanner pass. Background
// consumers use it to distinguish a partial cache update from a baseline
// that has visited every bucket and server pool.
ScannerCycle uint32 `json:"scannerCycle,omitempty"`
// Objects total count across all buckets
ObjectsTotalCount uint64 `json:"objectsCount"`
+5
View File
@@ -1068,6 +1068,11 @@ func (er erasureObjects) HealObject(ctx context.Context, bucket, object, version
newReqInfo = logger.NewReqInfo("", "", globalDeploymentID(), "", "Heal", bucket, object)
}
healCtx := logger.SetReqInfo(GlobalContext, newReqInfo)
if opts.NoLock {
// The caller owns the namespace lock. Stop if that lock's context is
// canceled instead of continuing metadata writes after losing it.
healCtx = logger.SetReqInfo(ctx, newReqInfo)
}
// Healing directories handle it separately.
if HasSuffix(object, SlashSeparator) {
+790
View File
@@ -0,0 +1,790 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"crypto/md5"
"encoding/base64"
"encoding/xml"
"fmt"
"io"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
"github.com/minio/minio/internal/hash"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/sio"
)
// TestAPISSECReplicaPartNumberReads replicates a three-part SSE-C multipart
// object through the trusted-replication write path and compares what
// GET ?partNumber=N returns before and after the replica overwrite.
func TestAPISSECReplicaPartNumberReads(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPISSECReplicaPartNumberReads,
})
}
func testAPISSECReplicaPartNumberReads(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
replicator := newObjectAttributesAuthzUser(t, instanceType, bucketName, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
key := bytes.Repeat([]byte{0x42}, 32)
keyMD5 := md5.Sum(key)
sseHeaders := map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
const mib = 1024 * 1024
partLens := []int{5 * mib, 5 * mib, 1 * mib}
plaintext := make([]byte, 0, 11*mib)
partData := make([][]byte, len(partLens))
for i, n := range partLens {
b := make([]byte, n)
for j := range b {
// Distinct, position-dependent bytes so an off-by-N shift is visible.
b[j] = byte(i*7 + j%251)
}
partData[i] = b
plaintext = append(plaintext, b...)
}
object := "ssec-mp-3part"
// ---- 1. Build the source: a real three-part SSE-C multipart object. ----
newRec := httptest.NewRecorder()
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
t.Fatalf("source NewMultipart status %d: %s", newRec.Code, newRec.Body.String())
}
var srcInit InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &srcInit, int64(newRec.Body.Len())); err != nil {
t.Fatal(err)
}
srcParts := make([]CompletePart, len(partLens))
for i, b := range partData {
pn := strconv.Itoa(i + 1)
partReq, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, srcInit.UploadID, pn),
int64(len(b)), bytes.NewReader(b), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, partReq)
if rec.Code != http.StatusOK {
t.Fatalf("source PutPart %s status %d: %s", pn, rec.Code, rec.Body.String())
}
srcParts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
srcCompleteBody, err := xml.Marshal(CompleteMultipartUpload{Parts: srcParts})
if err != nil {
t.Fatal(err)
}
completeReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, srcInit.UploadID), int64(len(srcCompleteBody)),
bytes.NewReader(srcCompleteBody), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
completeRec := httptest.NewRecorder()
apiRouter.ServeHTTP(completeRec, completeReq)
if completeRec.Code != http.StatusOK {
t.Fatalf("source Complete status %d: %s", completeRec.Code, completeRec.Body.String())
}
// ---- 2. Record the source's per-part metadata and per-part GET answers. ----
srcOI, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] SOURCE parts:", instanceType)
for _, p := range srcOI.Parts {
t.Logf(" part %d Size=%d ActualSize=%d", p.Number, p.Size, p.ActualSize)
}
srcActual, err := srcOI.GetActualSize()
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] SOURCE object Size=%d GetActualSize=%d actual-size-meta=%q",
instanceType, srcOI.Size, srcActual, srcOI.UserDefined[ReservedMetadataPrefix+"actual-size"])
type getResult struct {
status int
clen string
crange string
body []byte
}
doGet := func(query string, extra map[string]string) getResult {
hdrs := map[string]string{}
for k, v := range sseHeaders {
hdrs[k] = v
}
for k, v := range extra {
hdrs[k] = v
}
u := getGetObjectURL("", bucketName, object) + query
req, err := newTestSignedRequestV4(http.MethodGet, u, 0, nil, credentials.AccessKey, credentials.SecretKey, hdrs)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return getResult{
status: rec.Code,
clen: rec.Header().Get(xhttp.ContentLength),
crange: rec.Header().Get(xhttp.ContentRange),
body: append([]byte(nil), rec.Body.Bytes()...),
}
}
doHead := func(query string) getResult {
u := getGetObjectURL("", bucketName, object) + query
req, err := newTestSignedRequestV4(http.MethodHead, u, 0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return getResult{
status: rec.Code,
clen: rec.Header().Get(xhttp.ContentLength),
crange: rec.Header().Get(xhttp.ContentRange),
}
}
queries := []string{"?partNumber=1", "?partNumber=2", "?partNumber=3"}
srcGets := make([]getResult, len(queries))
for i, q := range queries {
srcGets[i] = doGet(q, nil)
t.Logf("[%s] SOURCE GET %s -> status=%d Content-Length=%s Content-Range=%s len(body)=%d",
instanceType, q, srcGets[i].status, srcGets[i].clen, srcGets[i].crange, len(srcGets[i].body))
}
// Range GET crossing the part1/part2 boundary.
boundaryRange := fmt.Sprintf("bytes=%d-%d", 5*mib-16, 5*mib+15)
srcRange := doGet("", map[string]string{"Range": boundaryRange})
t.Logf("[%s] SOURCE GET Range %s -> status=%d Content-Length=%s len(body)=%d",
instanceType, boundaryRange, srcRange.status, srcRange.clen, len(srcRange.body))
// Sanity: the source must return exactly the part bytes.
for i := range partData {
if !bytes.Equal(srcGets[i].body, partData[i]) {
t.Fatalf("[%s] SOURCE partNumber=%d returned wrong bytes (len %d want %d)",
instanceType, i+1, len(srcGets[i].body), len(partData[i]))
}
}
// ---- 3. Read the raw ciphertext the replication worker would ship. ----
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{}, ObjectOptions{ReplicationRequest: true})
if err != nil {
t.Fatal(err)
}
sourceInfo := gr.ObjInfo
rawAll, err := io.ReadAll(gr)
gr.Close()
if err != nil {
t.Fatal(err)
}
if bytes.Equal(rawAll, plaintext) {
t.Fatal("source replication read did not return encrypted bytes")
}
rawParts := make([][]byte, len(sourceInfo.Parts))
off := int64(0)
for i, p := range sourceInfo.Parts {
rawParts[i] = rawAll[off : off+p.Size]
off += p.Size
}
if off != int64(len(rawAll)) {
t.Fatalf("raw ciphertext length %d != sum of part sizes %d", len(rawAll), off)
}
// ---- 4. Replicate onto the same key through the trusted write path. ----
replicationOpts, isMP, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
t.Fatal(err)
}
if !isMP {
t.Fatal("SSE-C multipart source was not recognized as multipart")
}
replicationOpts.Internal.SourceMTime = time.Time{}
replicationHeaders := make(map[string]string)
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
replicationHeaders[name] = values[0]
}
}
replNewReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, replicator.AccessKey, replicator.SecretKey, replicationHeaders)
if err != nil {
t.Fatal(err)
}
replNewRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replNewRec, replNewReq)
if replNewRec.Code != http.StatusOK {
t.Fatalf("replica NewMultipart status %d: %s", replNewRec.Code, replNewRec.Body.String())
}
var replInit InitiateMultipartUploadResponse
if err = xmlDecoder(replNewRec.Body, &replInit, int64(replNewRec.Body.Len())); err != nil {
t.Fatal(err)
}
replParts := make([]CompletePart, len(rawParts))
for i, raw := range rawParts {
pn := strconv.Itoa(i + 1)
req, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, replInit.UploadID, pn),
int64(len(raw)), bytes.NewReader(raw), replicator.AccessKey, replicator.SecretKey,
map[string]string{xhttp.MinIOSourceReplicationRequest: "true"})
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("replica PutPart %s status %d: %s", pn, rec.Code, rec.Body.String())
}
replParts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
replCompleteBody, err := xml.Marshal(CompleteMultipartUpload{Parts: replParts})
if err != nil {
t.Fatal(err)
}
srcActualSize, err := sourceInfo.GetActualSize()
if err != nil {
t.Fatal(err)
}
replCompleteHeaders := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.MinIOSourceMTime: sourceInfo.ModTime.Format(time.RFC3339Nano),
xhttp.MinIOSourceETag: sourceInfo.ETag,
xhttp.MinIOReplicationActualObjectSize: strconv.FormatInt(srcActualSize, 10),
}
replCompleteReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, replInit.UploadID), int64(len(replCompleteBody)),
bytes.NewReader(replCompleteBody), replicator.AccessKey, replicator.SecretKey, replCompleteHeaders)
if err != nil {
t.Fatal(err)
}
replCompleteRec := httptest.NewRecorder()
apiRouter.ServeHTTP(replCompleteRec, replCompleteReq)
if replCompleteRec.Code != http.StatusOK {
t.Fatalf("replica Complete status %d: %s", replCompleteRec.Code, replCompleteRec.Body.String())
}
// ---- 5. Same reads against the replica. ----
repOI, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] REPLICA parts:", instanceType)
for _, p := range repOI.Parts {
t.Logf(" part %d Size=%d ActualSize=%d", p.Number, p.Size, p.ActualSize)
}
repActual, err := repOI.GetActualSize()
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] REPLICA object Size=%d GetActualSize=%d actual-size-meta=%q",
instanceType, repOI.Size, repActual, repOI.UserDefined[ReservedMetadataPrefix+"actual-size"])
// Whole-object GET must still be byte-identical.
whole := doGet("", nil)
if whole.status != http.StatusOK || !bytes.Equal(whole.body, plaintext) {
t.Errorf("[%s] REPLICA whole-object GET: status=%d len=%d want %d, equal=%v",
instanceType, whole.status, len(whole.body), len(plaintext), bytes.Equal(whole.body, plaintext))
} else {
t.Logf("[%s] REPLICA whole-object GET: OK, %d bytes identical", instanceType, len(whole.body))
}
// Fresh replica parts must record the plaintext lengths, not the
// ciphertext lengths the sender shipped.
if len(repOI.Parts) != len(partLens) {
t.Fatalf("[%s] REPLICA has %d parts, want %d", instanceType, len(repOI.Parts), len(partLens))
}
for i, p := range repOI.Parts {
if p.ActualSize != int64(partLens[i]) {
t.Errorf("[%s] REPLICA part %d ActualSize=%d, want the uploaded length %d",
instanceType, p.Number, p.ActualSize, partLens[i])
}
}
// Expected framing from independent prefix sums of the uploaded lengths.
total := len(plaintext)
start := 0
for i, q := range queries {
wantLen := partLens[i]
wantRange := fmt.Sprintf("bytes %d-%d/%d", start, start+wantLen-1, total)
start += wantLen
got := doGet(q, nil)
want := srcGets[i]
if got.status != want.status || got.status != http.StatusPartialContent {
t.Errorf("[%s] REPLICA GET %s status=%d, source=%d, want 206", instanceType, q, got.status, want.status)
}
if got.clen != strconv.Itoa(wantLen) || got.crange != wantRange {
t.Errorf("[%s] REPLICA GET %s Content-Length=%s Content-Range=%s, want %d and %q",
instanceType, q, got.clen, got.crange, wantLen, wantRange)
}
if want.clen != strconv.Itoa(wantLen) || want.crange != wantRange {
t.Errorf("[%s] SOURCE GET %s Content-Length=%s Content-Range=%s, want %d and %q",
instanceType, q, want.clen, want.crange, wantLen, wantRange)
}
if len(got.body) != wantLen {
t.Errorf("[%s] REPLICA GET %s body is %d bytes, want %d", instanceType, q, len(got.body), wantLen)
}
if !bytes.Equal(got.body, want.body) {
firstDiff := -1
for k := 0; k < len(got.body) && k < len(want.body); k++ {
if got.body[k] != want.body[k] {
firstDiff = k
break
}
}
t.Errorf("[%s] REPLICA partNumber=%d returned DIFFERENT bytes than the source: got %d bytes (Content-Range %q), want %d bytes (Content-Range %q), first differing byte at %d",
instanceType, i+1, len(got.body), got.crange, len(want.body), want.crange, firstDiff)
}
head := doHead(q)
if head.status != http.StatusPartialContent || head.clen != strconv.Itoa(wantLen) || head.crange != wantRange {
t.Errorf("[%s] REPLICA HEAD %s status=%d Content-Length=%s Content-Range=%s, want 206, %d and %q",
instanceType, q, head.status, head.clen, head.crange, wantLen, wantRange)
}
}
repRange := doGet("", map[string]string{"Range": boundaryRange})
sameRange := bytes.Equal(repRange.body, srcRange.body)
t.Logf("[%s] REPLICA GET Range %s -> status=%d Content-Length=%s len(body)=%d | source len=%d | bytes-equal=%v",
instanceType, boundaryRange, repRange.status, repRange.clen, len(repRange.body), len(srcRange.body), sameRange)
if !sameRange {
t.Errorf("[%s] REPLICA boundary Range GET returned different bytes", instanceType)
}
}
// TestSSECReplicaPartActualSizeDataMovement reproduces what a decommission or
// rebalance does to a replica whose parts already carry the ciphertext length in
// ActualSize: it replays the object through the object layer exactly the way
// decommissionObject does (cmd/erasure-server-pool-decom.go:605-667), passing the
// stale ActualSize to PutObjectPart and completing without ReplicationRequest, so
// CompleteMultipartUpload recomputes the object-level actual-size from the sum of
// part ActualSizes (cmd/erasure-multipart.go:1365,1440).
func TestSSECReplicaPartActualSizeDataMovement(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testSSECReplicaPartActualSizeDataMovement,
})
}
func testSSECReplicaPartActualSizeDataMovement(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
key := bytes.Repeat([]byte{0x37}, 32)
keyMD5 := md5.Sum(key)
sseHeaders := map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
const mib = 1024 * 1024
partLens := []int{5 * mib, 5 * mib, 1 * mib}
plaintext := make([]byte, 0, 11*mib)
partData := make([][]byte, len(partLens))
for i, n := range partLens {
b := make([]byte, n)
for j := range b {
b[j] = byte(i*13 + j%241)
}
partData[i] = b
plaintext = append(plaintext, b...)
}
object := "ssec-mp-datamovement"
newRec := httptest.NewRecorder()
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
t.Fatalf("NewMultipart status %d: %s", newRec.Code, newRec.Body.String())
}
var init InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &init, int64(newRec.Body.Len())); err != nil {
t.Fatal(err)
}
srcParts := make([]CompletePart, len(partLens))
for i, b := range partData {
req, err := newTestSignedRequestV4(http.MethodPut,
getPutObjectPartURL("", bucketName, object, init.UploadID, strconv.Itoa(i+1)),
int64(len(b)), bytes.NewReader(b), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("PutPart %d status %d: %s", i+1, rec.Code, rec.Body.String())
}
srcParts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
body, err := xml.Marshal(CompleteMultipartUpload{Parts: srcParts})
if err != nil {
t.Fatal(err)
}
cReq, err := newTestSignedRequestV4(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, init.UploadID), int64(len(body)),
bytes.NewReader(body), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
cRec := httptest.NewRecorder()
apiRouter.ServeHTTP(cRec, cReq)
if cRec.Code != http.StatusOK {
t.Fatalf("Complete status %d: %s", cRec.Code, cRec.Body.String())
}
// Replay it the way decommissionObject does, but hand PutObjectPart the STALE
// ActualSize an already-written SSE-C replica carries: the ciphertext length.
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{},
ObjectOptions{NoDecryption: true, NoLock: true, NoAuditLog: true})
if err != nil {
t.Fatal(err)
}
oi := gr.ObjInfo
res, err := obj.NewMultipartUpload(t.Context(), bucketName, object, ObjectOptions{
UserDefined: oi.UserDefined,
DataMovement: true,
NoAuditLog: true,
})
if err != nil {
gr.Close()
t.Fatal(err)
}
moved := make([]CompletePart, len(oi.Parts))
for i, part := range oi.Parts {
staleActual := part.Size // what a bad replica records
hr, herr := hash.NewReader(t.Context(), io.LimitReader(gr, part.Size), part.Size, "", "", staleActual)
if herr != nil {
gr.Close()
t.Fatal(herr)
}
pi, perr := obj.PutObjectPart(t.Context(), bucketName, object, res.UploadID, part.Number,
NewPutObjReader(hr), ObjectOptions{
PreserveETag: part.ETag,
IndexCB: func() []byte { return part.Index },
NoAuditLog: true,
})
if perr != nil {
gr.Close()
t.Fatalf("data-movement PutObjectPart part %d: %v", part.Number, perr)
}
moved[i] = CompletePart{ETag: pi.ETag, PartNumber: pi.PartNumber}
}
gr.Close()
// decommissionObject/rebalanceObject complete WITHOUT ReplicationRequest, so
// the object-level actual-size is recomputed from the part ActualSizes.
if _, err = obj.CompleteMultipartUpload(t.Context(), bucketName, object, res.UploadID, moved,
ObjectOptions{DataMovement: true, MTime: oi.ModTime, NoAuditLog: true}); err != nil {
t.Fatalf("data-movement CompleteMultipartUpload: %v", err)
}
after, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
t.Logf("[%s] AFTER DATA MOVEMENT parts:", instanceType)
for _, p := range after.Parts {
t.Logf(" part %d Size=%d ActualSize=%d", p.Number, p.Size, p.ActualSize)
}
gotActual, err := after.GetActualSize()
if err != nil {
t.Fatalf("GetActualSize after data movement: %v", err)
}
t.Logf("[%s] AFTER DATA MOVEMENT object Size=%d GetActualSize=%d actual-size-meta=%q",
instanceType, after.Size, gotActual, after.UserDefined[ReservedMetadataPrefix+"actual-size"])
wantActual := int64(len(plaintext))
if gotActual != wantActual {
t.Errorf("[%s] object-level actual size after data movement = %d, want %d",
instanceType, gotActual, wantActual)
}
for i, p := range after.Parts {
if p.ActualSize != int64(partLens[i]) {
t.Errorf("[%s] part %d ActualSize after data movement = %d, want %d",
instanceType, p.Number, p.ActualSize, partLens[i])
}
}
getReq, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
getRec := httptest.NewRecorder()
apiRouter.ServeHTTP(getRec, getReq)
if getRec.Code != http.StatusOK || !bytes.Equal(getRec.Body.Bytes(), plaintext) {
t.Errorf("[%s] whole-object GET after data movement: status=%d len=%d want %d equal=%v",
instanceType, getRec.Code, getRec.Body.Len(), len(plaintext), bytes.Equal(getRec.Body.Bytes(), plaintext))
}
// The advertised length must be the body length: a poisoned object-level
// actual-size shows up here as a Content-Length larger than the body.
if clen := getRec.Header().Get(xhttp.ContentLength); clen != strconv.Itoa(len(plaintext)) || clen != strconv.Itoa(getRec.Body.Len()) {
t.Errorf("[%s] whole-object GET after data movement advertises Content-Length=%s for a %d-byte body (plaintext %d)",
instanceType, clen, getRec.Body.Len(), len(plaintext))
}
}
// TestPartNumberToRangeSpecEncryptedParts pins the read-side repair: for an
// encrypted, uncompressed object the part range is derived from the stored
// ciphertext length, so a replica whose parts still record the ciphertext
// length in ActualSize reads correctly, while plaintext and compressed objects
// keep using ActualSize, and a part whose length cannot be a valid encrypted
// stream is reported as tampered by both callers. See pgsty/silo#119.
func TestPartNumberToRangeSpecEncryptedParts(t *testing.T) {
const mib = 1024 * 1024
plain := []int64{5 * mib, 5 * mib, 1024}
cipher := make([]int64, len(plain))
for i, n := range plain {
c, err := sio.EncryptedSize(uint64(n))
if err != nil {
t.Fatal(err)
}
cipher[i] = int64(c)
}
sum := func(v []int64) (s int64) {
for _, n := range v {
s += n
}
return s
}
encMeta := map[string]string{
crypto.MetaSealedKeySSEC: "sealed-key",
crypto.MetaIV: "iv",
crypto.MetaAlgorithm: crypto.InsecureSealAlgorithm,
}
compressedEncMeta := map[string]string{ReservedMetadataPrefix + "compression": compressionAlgorithmV2}
for k, v := range encMeta {
compressedEncMeta[k] = v
}
mkParts := func(sizes, actual []int64) []ObjectPartInfo {
parts := make([]ObjectPartInfo, len(sizes))
for i := range sizes {
parts[i] = ObjectPartInfo{Number: i + 1, Size: sizes[i], ActualSize: actual[i]}
}
return parts
}
// A compressed part's ActualSize is the uploaded length before compression,
// which bears no relation to the ciphertext length: use lengths whose
// decrypted size differs from ActualSize so that dropping the compression
// exclusion is detectable.
uploaded := []int64{2 * plain[0], 2 * plain[1], 2 * plain[2]}
wantRangeOf := func(lens []int64, pn int) (start, end int64) {
for i := 0; i < pn-1; i++ {
start += lens[i]
}
return start, start + lens[pn-1] - 1
}
for _, tc := range []struct {
name string
oi ObjectInfo
lens []int64
}{
{"encrypted, stale ciphertext ActualSize", ObjectInfo{Size: sum(cipher), UserDefined: encMeta, Parts: mkParts(cipher, cipher)}, plain},
{"encrypted, correct ActualSize", ObjectInfo{Size: sum(cipher), UserDefined: encMeta, Parts: mkParts(cipher, plain)}, plain},
{"plaintext", ObjectInfo{Size: sum(plain), UserDefined: map[string]string{}, Parts: mkParts(plain, plain)}, plain},
{"compressed and encrypted keeps ActualSize", ObjectInfo{Size: sum(cipher), UserDefined: compressedEncMeta, Parts: mkParts(cipher, uploaded)}, uploaded},
} {
for pn := 1; pn <= len(plain); pn++ {
rs, err := partNumberToRangeSpec(tc.oi, pn)
if err != nil {
t.Fatalf("%s: partNumber=%d: %v", tc.name, pn, err)
}
start, end := wantRangeOf(tc.lens, pn)
if rs == nil || rs.Start != start || rs.End != end {
t.Errorf("%s: partNumber=%d range %+v, want %d-%d", tc.name, pn, rs, start, end)
}
}
}
// A 31-byte part cannot be a sio stream: both callers report it as tampered.
bad := ObjectInfo{
Size: 31 + cipher[1] + cipher[2],
UserDefined: encMeta,
Parts: mkParts([]int64{31, cipher[1], cipher[2]}, []int64{31, plain[1], plain[2]}),
}
for pn := 1; pn <= 2; pn++ {
if _, err := partNumberToRangeSpec(bad, pn); err != errObjectTampered {
t.Errorf("malformed part: partNumber=%d err=%v, want errObjectTampered", pn, err)
}
}
if _, _, _, err := NewGetObjectReader(nil, bad, ObjectOptions{PartNumber: 1}, http.Header{}); err != errObjectTampered {
t.Errorf("NewGetObjectReader on a malformed part: err=%v, want errObjectTampered", err)
}
if err := setObjectHeaders(t.Context(), httptest.NewRecorder(), bad, nil, ObjectOptions{PartNumber: 1}); err != errObjectTampered {
t.Errorf("setObjectHeaders on a malformed part: err=%v, want errObjectTampered", err)
}
// A zero-length trailing part is a valid (empty) stream and stays accepted.
zero := ObjectInfo{Size: cipher[0], UserDefined: encMeta, Parts: mkParts([]int64{cipher[0], 0}, []int64{cipher[0], 0})}
rs, err := partNumberToRangeSpec(zero, 2)
if err != nil || rs == nil || rs.Start != plain[0] {
t.Errorf("zero-length trailing part: range %+v err=%v, want start %d", rs, err, plain[0])
}
}
// TestAPISSECReplicaMalformedPartIsRejected asserts that a trusted SSE-C replica
// part whose ciphertext length cannot be a valid encrypted stream is rejected
// as tampered before the part is committed. See pgsty/silo#119.
func TestAPISSECReplicaMalformedPartIsRejected(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPISSECReplicaMalformedPartIsRejected})
}
func testAPISSECReplicaMalformedPartIsRejected(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
replicator := newObjectAttributesAuthzUser(t, instanceType, bucketName, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
key := bytes.Repeat([]byte{0x45}, 32)
keyMD5 := md5.Sum(key)
sseHeaders := map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(keyMD5[:]),
}
object := "ssec-replica-malformed"
data := bytes.Repeat([]byte("malformed-part-"), 1024)
putReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectURL("", bucketName, object), int64(len(data)),
bytes.NewReader(data), credentials.AccessKey, credentials.SecretKey, sseHeaders)
if err != nil {
t.Fatal(err)
}
putRec := httptest.NewRecorder()
apiRouter.ServeHTTP(putRec, putReq)
if putRec.Code != http.StatusOK {
t.Fatalf("%s: source PUT %d: %s", instanceType, putRec.Code, putRec.Body.String())
}
gr, err := obj.GetObjectNInfo(t.Context(), bucketName, object, nil, http.Header{}, ObjectOptions{ReplicationRequest: true})
if err != nil {
t.Fatal(err)
}
sourceInfo := gr.ObjInfo
raw, err := io.ReadAll(gr)
gr.Close()
if err != nil {
t.Fatal(err)
}
replicationOpts, _, err := putReplicationOpts(t.Context(), "", sourceInfo)
if err != nil {
t.Fatal(err)
}
replicationOpts.Internal.SourceMTime = time.Time{}
replicationHeaders := make(map[string]string)
for name, values := range replicationOpts.Header() {
if len(values) > 0 {
replicationHeaders[name] = values[0]
}
}
newReq, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", bucketName, object),
0, nil, replicator.AccessKey, replicator.SecretKey, replicationHeaders)
if err != nil {
t.Fatal(err)
}
newRec := httptest.NewRecorder()
apiRouter.ServeHTTP(newRec, newReq)
if newRec.Code != http.StatusOK {
t.Fatalf("%s: replica NewMultipart %d: %s", instanceType, newRec.Code, newRec.Body.String())
}
var init InitiateMultipartUploadResponse
if err = xmlDecoder(newRec.Body, &init, int64(newRec.Body.Len())); err != nil {
t.Fatal(err)
}
partHeaders := map[string]string{xhttp.MinIOSourceReplicationRequest: "true"}
// 31 bytes cannot be a sio stream (the package header plus its
// authentication tag occupy 32 bytes).
badReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectPartURL("", bucketName, object, init.UploadID, "1"),
31, bytes.NewReader(raw[:31]), replicator.AccessKey, replicator.SecretKey, partHeaders)
if err != nil {
t.Fatal(err)
}
badRec := httptest.NewRecorder()
apiRouter.ServeHTTP(badRec, badReq)
if badRec.Code == http.StatusOK || !strings.Contains(badRec.Body.String(), "XMinioObjectTampered") {
t.Fatalf("%s: malformed replica part answered %d: %s", instanceType, badRec.Code, badRec.Body.String())
}
lpi, err := obj.ListObjectParts(t.Context(), bucketName, object, init.UploadID, 0, 10, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if len(lpi.Parts) != 0 {
t.Fatalf("%s: malformed replica part was committed: %+v", instanceType, lpi.Parts)
}
// Control: the real ciphertext still uploads on the same upload.
goodReq, err := newTestSignedRequestV4(http.MethodPut, getPutObjectPartURL("", bucketName, object, init.UploadID, "1"),
int64(len(raw)), bytes.NewReader(raw), replicator.AccessKey, replicator.SecretKey, partHeaders)
if err != nil {
t.Fatal(err)
}
goodRec := httptest.NewRecorder()
apiRouter.ServeHTTP(goodRec, goodReq)
if goodRec.Code != http.StatusOK {
t.Fatalf("%s: valid replica part answered %d: %s", instanceType, goodRec.Code, goodRec.Body.String())
}
lpi, err = obj.ListObjectParts(t.Context(), bucketName, object, init.UploadID, 0, 10, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if len(lpi.Parts) != 1 || lpi.Parts[0].ActualSize != int64(len(data)) {
t.Fatalf("%s: valid replica part recorded %+v, want one part with ActualSize %d", instanceType, lpi.Parts, len(data))
}
}
@@ -157,6 +157,13 @@ func copyPartWithoutChecksumHTTP(t *testing.T, apiRouter http.Handler, creds aut
if sourceRange != "" {
req.Header.Set(xhttp.AmzCopySourceRange, sourceRange)
}
// Re-sign so the copy-source x-amz-* headers are covered by the signature,
// as real S3 clients send them; the verifier rejects unsigned x-amz-*.
if creds.AccessKey != "" && creds.SecretKey != "" {
if err := signRequestV4(req, creds.AccessKey, creds.SecretKey); err != nil {
t.Fatalf("failed to re-sign UploadPartCopy request: %v", err)
}
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
+72 -26
View File
@@ -710,20 +710,26 @@ func (er erasureObjects) PutObjectPart(ctx context.Context, bucket, object, uplo
}
actualSize := data.ActualSize()
if actualSize < 0 {
_, encrypted := crypto.IsEncrypted(fi.Metadata)
compressed := fi.IsCompressed()
switch {
case compressed:
// ... nothing changes for compressed stream.
// if actualSize is -1 we have no known way to
// determine what is the actualSize.
case encrypted:
decSize, err := sio.DecryptedSize(uint64(n))
if err == nil {
actualSize = int64(decSize)
}
default:
_, encrypted := crypto.IsEncrypted(fi.Metadata)
compressed := fi.IsCompressed()
switch {
case compressed:
// ... nothing changes for compressed stream.
// if actualSize is -1 we have no known way to
// determine what is the actualSize.
case encrypted:
// The uploaded length of an encrypted part is always derivable from the
// bytes just written, and the caller's value cannot be trusted: trusted
// SSE-C replication and the data movement paths hand over the ciphertext
// length. Derive it with the arithmetic the read path applies to
// part.Size, so the stored value matches how the part is read back.
decSize, err := sio.DecryptedSize(uint64(n))
if err != nil {
return pi, toObjectErr(errObjectTampered, bucket, object, uploadID)
}
actualSize = int64(decSize)
default:
if actualSize < 0 {
actualSize = n
}
}
@@ -1108,7 +1114,7 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
auditObjectErasureSet(ctx, "CompleteMultipartUpload", object, &er)
}
if opts.CheckPrecondFn != nil {
if opts.CheckPrecondFn != nil || opts.ReplicaLockReconcile {
if !opts.NoLock {
ns := er.NewNSLock(bucket, object)
lkctx, err := ns.GetLock(ctx, globalOperationTimeout)
@@ -1120,18 +1126,24 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
opts.NoLock = true
}
obj, err := er.getObjectInfo(ctx, bucket, object, opts)
if err == nil && opts.CheckPrecondFn(obj) {
return ObjectInfo{}, PreConditionFailed{}
}
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
return ObjectInfo{}, err
}
// The Object Lock reconcile below needs the version being committed, read
// after checkUploadIDExists, so only the precondition read happens here;
// both run under this same write lock, held until the version is renamed
// into place.
if opts.CheckPrecondFn != nil {
obj, err := er.getObjectInfo(ctx, bucket, object, opts)
if err == nil && opts.CheckPrecondFn(obj) {
return ObjectInfo{}, PreConditionFailed{}
}
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
return ObjectInfo{}, err
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return ObjectInfo{}, err
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return ObjectInfo{}, err
}
}
}
@@ -1143,6 +1155,40 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
return oi, toObjectErr(err, bucket, object, uploadID)
}
// Reconcile against the upload's persisted version, under the object lock.
// Multi-pool callers supply a resolver spanning all pools, including those
// draining their contents; the upload itself stays in its original pool.
if opts.ReplicaLockReconcile {
// A persisted upload records the null version as an empty VersionID; look
// it up as the null version so the reconcile reads the addressed version's
// stored lock, not the latest version's.
lookupVersionID := fi.VersionID
if lookupVersionID == "" {
lookupVersionID = nullVersionID
}
getObjectInfo := er.getObjectInfo
if opts.replicaObjectInfo != nil {
getObjectInfo = opts.replicaObjectInfo
}
curr, gerr := getObjectInfo(ctx, bucket, object, ObjectOptions{
VersionID: lookupVersionID,
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoLock: true,
})
switch {
case gerr == nil:
reconcileStoredObjectLock(fi.Metadata, storedObjectLockState(curr.UserDefined))
reconcileStoredObjectTags(fi.Metadata, curr.UserDefined)
case isErrVersionNotFound(gerr) || isErrObjectNotFound(gerr):
// No existing version to order against: keep the upload's own accepted
// lock, including a pre-upgrade upload that persisted values without
// their ordering timestamps.
default:
return oi, toObjectErr(gerr, bucket, object)
}
}
uploadIDPath := er.getUploadIDDir(bucket, object, uploadID)
onlineDisks := er.getDisks()
writeQuorum := fi.WriteQuorum(er.defaultWQuorum())
@@ -0,0 +1,295 @@
// Copyright (c) 2015-2025 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"net/http"
"testing"
)
// TestDeleteObjectConditional verifies that a conditional DeleteObject
// (If-Match, wired through opts.CheckPrecondFn) is evaluated atomically at the
// object layer: a non-matching ETag must fail with PreConditionFailed and leave
// the object intact, a matching ETag must delete it, and an If-Match against a
// missing object must return a not-found error rather than silently succeeding.
func TestDeleteObjectConditional(t *testing.T) {
ctx := context.Background()
obj, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatal(err)
}
defer obj.Shutdown(context.Background())
defer removeRoots(fsDirs)
bucket := "test-bucket"
object := "test-object"
if err = obj.MakeBucket(ctx, bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
if _, err = obj.PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewReader([]byte("test-value")),
int64(len("test-value")), "", ""), ObjectOptions{}); err != nil {
t.Fatal(err)
}
objInfo, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
existingETag := objInfo.ETag
// If-Match with a wrong ETag must fail and preserve the object.
t.Run("wrong-etag-precondition-failed", func(t *testing.T) {
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, "wrong-etag")
},
}
if _, err := obj.DeleteObject(ctx, bucket, object, opts); !isErrPreconditionFailed(err) {
t.Errorf("expected PreConditionFailed, got: %v", err)
}
if _, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{}); err != nil {
t.Errorf("object must still exist after a failed conditional delete, got: %v", err)
}
})
// If-Match against a missing object must return a not-found error.
t.Run("missing-object-not-found", func(t *testing.T) {
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, existingETag)
},
}
_, err := obj.DeleteObject(ctx, bucket, "does-not-exist", opts)
if !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
t.Errorf("expected ObjectNotFound/VersionNotFound, got: %v", err)
}
})
// If-Match with the correct ETag must delete the object (run last).
t.Run("correct-etag-succeeds", func(t *testing.T) {
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, existingETag)
},
}
if _, err := obj.DeleteObject(ctx, bucket, object, opts); err != nil {
t.Errorf("expected a successful delete with matching ETag, got: %v", err)
}
if _, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{}); !isErrObjectNotFound(err) {
t.Errorf("object must be removed after a matching conditional delete, got: %v", err)
}
})
}
// TestDeleteObjectConditionalWithReadQuorumFailure verifies that a conditional
// (If-Match) DeleteObject does NOT proceed when the object's current state
// cannot be read due to read-quorum loss: without a verified ETag the delete
// must fail rather than remove the object blindly.
func TestDeleteObjectConditionalWithReadQuorumFailure(t *testing.T) {
ctx := context.Background()
obj, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatal(err)
}
defer obj.Shutdown(context.Background())
defer removeRoots(fsDirs)
z := obj.(*erasureServerPools)
xl := z.serverPools[0].sets[0]
bucket := "test-bucket"
object := "test-object"
if err = obj.MakeBucket(ctx, bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
if _, err = obj.PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewReader([]byte("test-value")),
int64(len("test-value")), "", ""), ObjectOptions{}); err != nil {
t.Fatal(err)
}
objInfo, err := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
existingETag := objInfo.ETag
// Simulate read-quorum loss by taking 8 of 16 disks offline (EC 8+8).
erasureDisks := xl.getDisks()
z.serverPools[0].erasureDisksMu.Lock()
xl.getDisks = func() []StorageAPI {
for i := range erasureDisks[:8] {
erasureDisks[i] = nil
}
return erasureDisks
}
z.serverPools[0].erasureDisksMu.Unlock()
// Even with the correct ETag we must not delete: the current state (hence the
// ETag) cannot be verified under read-quorum loss.
opts := ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
return !isETagEqual(oi.ETag, existingETag)
},
}
if _, err := obj.DeleteObject(ctx, bucket, object, opts); err == nil {
t.Error("expected an error for a conditional delete under read-quorum loss, got nil (object may have been deleted without ETag verification)")
}
}
// TestDeleteObjectConditionalVersioned verifies conditional DeleteObject on a
// versioned bucket, where the precondition is evaluated at the server-pool layer
// against the version that will actually be removed:
// - If-Match "*" when the latest version is a delete marker must fail (412),
// because there is no live object to match.
// - An explicit versionId If-Match is evaluated against the addressed version,
// not the latest one (match deletes it, mismatch is refused).
// - An If-Match against a missing version returns VersionNotFound, which the
// handler maps to NoSuchVersion.
func TestDeleteObjectConditionalVersioned(t *testing.T) {
ctx := context.Background()
obj, fsDirs, err := prepareErasure16(ctx)
if err != nil {
t.Fatal(err)
}
defer obj.Shutdown(context.Background())
defer removeRoots(fsDirs)
bucket := "test-bucket"
if err = obj.MakeBucket(ctx, bucket, MakeBucketOptions{VersioningEnabled: true}); err != nil {
t.Fatal(err)
}
versioned := globalBucketVersioningSys.PrefixEnabled(bucket, "any")
if !versioned {
t.Fatalf("expected versioning to be enabled on %q", bucket)
}
put := func(object, content string) ObjectInfo {
oi, perr := obj.PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewReader([]byte(content)), int64(len(content)), "", ""),
ObjectOptions{Versioned: versioned})
if perr != nil {
t.Fatalf("put %q: %v", object, perr)
}
return oi
}
ifMatch := func(value string) CheckPreconditionFn {
return func(oi ObjectInfo) bool {
return deleteIfMatchPreconditionFailed(http.Header{}, value, oi)
}
}
// If-Match "*" against a delete-marker-latest must fail with 412.
t.Run("wildcard-on-delete-marker-latest", func(t *testing.T) {
object := "dm-object"
put(object, "v1")
// Create a delete marker (unconditional), making the latest a delete marker.
if _, derr := obj.DeleteObject(ctx, bucket, object, ObjectOptions{Versioned: versioned}); derr != nil {
t.Fatalf("create delete marker: %v", derr)
}
opts := ObjectOptions{Versioned: versioned, HasIfMatch: true, CheckPrecondFn: ifMatch("*")}
if _, derr := obj.DeleteObject(ctx, bucket, object, opts); !isErrPreconditionFailed(derr) {
t.Errorf("expected PreConditionFailed for If-Match:* on a delete-marker-latest, got: %v", derr)
}
})
// Explicit versionId is evaluated against the addressed (older) version.
t.Run("explicit-version-selection", func(t *testing.T) {
object := "ver-object"
v1 := put(object, "first")
v2 := put(object, "second-longer") // v2 is now the latest with a different ETag
if v1.ETag == v2.ETag {
t.Fatalf("test setup: versions must have distinct ETags")
}
// Mismatch: delete v2 with v1's ETag must be refused, v2 preserved.
mismatch := ObjectOptions{Versioned: versioned, VersionID: v2.VersionID, HasIfMatch: true, CheckPrecondFn: ifMatch(v1.ETag)}
if _, derr := obj.DeleteObject(ctx, bucket, object, mismatch); !isErrPreconditionFailed(derr) {
t.Errorf("expected PreConditionFailed deleting v2 with v1 ETag, got: %v", derr)
}
if _, gerr := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: v2.VersionID}); gerr != nil {
t.Errorf("v2 must still exist after a refused conditional delete, got: %v", gerr)
}
// Match: delete v1 with v1's ETag must succeed even though v1 is not latest.
match := ObjectOptions{Versioned: versioned, VersionID: v1.VersionID, HasIfMatch: true, CheckPrecondFn: ifMatch(v1.ETag)}
if _, derr := obj.DeleteObject(ctx, bucket, object, match); derr != nil {
t.Errorf("expected the addressed version to be deleted, got: %v", derr)
}
if _, gerr := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: v1.VersionID}); !isErrVersionNotFound(gerr) {
t.Errorf("v1 must be gone after a matching conditional delete, got: %v", gerr)
}
if _, gerr := obj.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: v2.VersionID}); gerr != nil {
t.Errorf("v2 must remain after deleting v1, got: %v", gerr)
}
})
// If-Match against a missing version on an EXISTING key returns VersionNotFound.
t.Run("missing-version", func(t *testing.T) {
object := "missing-version-object"
put(object, "only")
opts := ObjectOptions{Versioned: versioned, VersionID: mustGetUUID(), HasIfMatch: true, CheckPrecondFn: ifMatch("anything")}
if _, derr := obj.DeleteObject(ctx, bucket, object, opts); !isErrVersionNotFound(derr) {
t.Errorf("expected VersionNotFound for If-Match on a missing version, got: %v", derr)
}
})
// If-Match against a missing version on an ABSENT key must also return
// VersionNotFound (NoSuchVersion), not NoSuchKey: the request addresses a
// specific version, which does not exist regardless of the key.
t.Run("missing-version-absent-key", func(t *testing.T) {
opts := ObjectOptions{Versioned: versioned, VersionID: mustGetUUID(), HasIfMatch: true, CheckPrecondFn: ifMatch("anything")}
if _, derr := obj.DeleteObject(ctx, bucket, "never-existed", opts); !isErrVersionNotFound(derr) {
t.Errorf("expected VersionNotFound for If-Match on a version of an absent key, got: %v", derr)
}
})
// If-Match "*" addressing a delete-marker VERSION by id must fail with 412,
// not 405: a delete marker has no entity-tag to match. getObjectInfo returns
// the marker alongside MethodNotAllowed; the precondition runs on the marker.
t.Run("wildcard-on-explicit-delete-marker-version", func(t *testing.T) {
object := "explicit-dm-object"
put(object, "live")
dm, derr := obj.DeleteObject(ctx, bucket, object, ObjectOptions{Versioned: versioned})
if derr != nil {
t.Fatalf("create delete marker: %v", derr)
}
if !dm.DeleteMarker || dm.VersionID == "" {
t.Fatalf("expected a delete-marker version, got DeleteMarker=%v VersionID=%q", dm.DeleteMarker, dm.VersionID)
}
opts := ObjectOptions{Versioned: versioned, VersionID: dm.VersionID, HasIfMatch: true, CheckPrecondFn: ifMatch("*")}
if _, derr := obj.DeleteObject(ctx, bucket, object, opts); !isErrPreconditionFailed(derr) {
t.Errorf("expected PreConditionFailed for If-Match:* on an addressed delete-marker version, got: %v", derr)
}
})
}
+28 -8
View File
@@ -133,6 +133,11 @@ func (er erasureObjects) CopyObject(ctx context.Context, srcBucket, srcObject, d
return fi.ToObjectInfo(srcBucket, srcObject, srcOpts.Versioned || srcOpts.VersionSuspended), toObjectErr(errMethodNotAllowed, srcBucket, srcObject)
}
if dstOpts.ReplicaLockReconcile {
reconcileStoredObjectLock(srcInfo.UserDefined, storedObjectLockState(fi.Metadata))
reconcileStoredObjectTags(srcInfo.UserDefined, fi.Metadata)
}
filterOnlineDisksInplace(fi, metaArr, onlineDisks)
versionID := srcInfo.VersionID
@@ -1268,7 +1273,7 @@ func (er erasureObjects) putObject(ctx context.Context, bucket string, object st
data := r.Reader
if opts.CheckPrecondFn != nil {
if opts.CheckPrecondFn != nil || opts.ReplicaLockReconcile {
if !opts.NoLock {
ns := er.NewNSLock(bucket, object)
lkctx, err := ns.GetLock(ctx, globalOperationTimeout)
@@ -1280,18 +1285,33 @@ func (er erasureObjects) putObject(ctx context.Context, bucket string, object st
opts.NoLock = true
}
obj, err := er.getObjectInfo(ctx, bucket, object, opts)
if err == nil && opts.CheckPrecondFn(obj) {
return objInfo, PreConditionFailed{}
getObjectInfo := er.getObjectInfo
if opts.replicaObjectInfo != nil {
getObjectInfo = opts.replicaObjectInfo
}
obj, err := getObjectInfo(ctx, bucket, object, opts)
// A destination read that fails for a reason other than not-found must not
// be taken as a passed precondition or as absent lock state.
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
return objInfo, err
}
if opts.CheckPrecondFn != nil {
if err == nil && opts.CheckPrecondFn(obj) {
return objInfo, PreConditionFailed{}
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return objInfo, err
}
}
// if object doesn't exist return error for If-Match conditional requests
// If-None-Match should be allowed to proceed for non-existent objects
if err != nil && opts.HasIfMatch && (isErrObjectNotFound(err) || isErrVersionNotFound(err)) {
return objInfo, err
// Re-read under the write lock, using the pools-layer resolver when
// present. A missing version keeps the incoming accepted state; an
// existing version contributes independently ordered lock and tags.
if opts.ReplicaLockReconcile && err == nil {
reconcileStoredObjectLock(opts.UserDefined, storedObjectLockState(obj.UserDefined))
reconcileStoredObjectTags(opts.UserDefined, obj.UserDefined)
}
}
+350
View File
@@ -0,0 +1,350 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"context"
"maps"
"sort"
"strings"
"time"
madmin "github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/bucket/lifecycle"
xhttp "github.com/minio/minio/internal/http"
)
// objectPoolInfos reads the addressed version in every pool, including pools
// draining their contents. The caller must hold the pools-layer object lock.
// An unreadable pool may contain a newer version or metadata; it is not absence.
func (z *erasureServerPools) objectPoolInfos(ctx context.Context, bucket, object string, opts ObjectOptions) ([]PoolObjInfo, error) {
opts.NoLock = true
opts.CheckPrecondFn = nil
var copies []PoolObjInfo
for i, pool := range z.serverPools {
oi, err := pool.GetObjectInfo(ctx, bucket, object, opts)
if err == nil || (oi.DeleteMarker && (isErrObjectNotFound(err) || isErrMethodNotAllowed(err))) {
copies = append(copies, PoolObjInfo{Index: i, ObjInfo: oi})
continue
}
if !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
return nil, err
}
}
sort.Slice(copies, func(i, j int) bool {
a, b := copies[i], copies[j]
if a.ObjInfo.ModTime.Equal(b.ObjInfo.ModTime) {
return a.Index < b.Index
}
return a.ObjInfo.ModTime.After(b.ObjInfo.ModTime)
})
if len(copies) == 0 {
if opts.VersionID != "" {
return nil, VersionNotFound{Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID}
}
return nil, ObjectNotFound{Bucket: bucket, Object: decodeDirObject(object)}
}
return copies, nil
}
// Healing also commits object metadata. Both queued healing and drive healing
// must use the same namespace as pooled writes, rather than racing them under
// a destination set's independent namespace.
func (z *erasureServerPools) healObjectInPool(ctx context.Context, er *erasureObjects, bucket, object, versionID string, opts madmin.HealOpts) (madmin.HealResultItem, error) {
if !z.SinglePool() && !opts.NoLock {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return madmin.HealResultItem{}, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
opts.NoLock = true
}
return er.HealObject(ctx, bucket, object, versionID, opts)
}
// mergePoolLockState resolves each independently ordered field. An ordered
// removal is a value in its own right; an absent unordered field is not one.
func mergePoolLockState(copies []PoolObjInfo) objectLockState {
state := storedObjectLockState(copies[0].ObjInfo.UserDefined)
for _, copy := range copies[1:] {
next := storedObjectLockState(copy.ObjInfo.UserDefined)
retentionTime, _ := time.Parse(time.RFC3339Nano, next.retentionTimestamp)
if state.retentionIsOlderThan(retentionTime) || (state.retentionTimestamp == "" && state.mode == "" && next.mode != "") {
state.mode, state.retainUntil, state.retentionTimestamp = next.mode, next.retainUntil, next.retentionTimestamp
}
holdTime, _ := time.Parse(time.RFC3339Nano, next.legalHoldTimestamp)
if state.legalHoldIsOlderThan(holdTime) || (state.legalHoldTimestamp == "" && state.legalHold == "" && next.legalHold != "") {
state.legalHold, state.legalHoldTimestamp = next.legalHold, next.legalHoldTimestamp
}
}
return state
}
func replaceObjectLockMetadata(metadata map[string]string, state objectLockState) {
for _, key := range []string{
strings.ToLower(xhttp.AmzObjectLockMode), strings.ToLower(xhttp.AmzObjectLockRetainUntilDate),
strings.ToLower(xhttp.AmzObjectLockLegalHold),
ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp,
ReservedMetadataPrefixLower + ObjectLockLegalHoldTimestamp,
} {
// Metadata updates use maps.Copy at the set layer. Empty values also
// clear a previous value there, including timestamp-only removals.
if _, exists := metadata[key]; exists {
metadata[key] = ""
}
}
state.restoreRetention(metadata)
state.restoreLegalHold(metadata)
}
func mergedPoolObjectInfo(copies []PoolObjInfo) ObjectInfo {
oi := copies[0].ObjInfo
oi.UserDefined = maps.Clone(oi.UserDefined)
if oi.UserDefined == nil {
oi.UserDefined = make(map[string]string)
}
replaceObjectLockMetadata(oi.UserDefined, mergePoolLockState(copies))
for _, copy := range copies[1:] {
ts := copy.ObjInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]
stamp, _ := time.Parse(time.RFC3339Nano, ts)
if olderThan(oi.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], stamp) {
oi.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = ts
oi.UserDefined[xhttp.AmzObjectTagging] = copy.ObjInfo.UserDefined[xhttp.AmzObjectTagging]
}
}
oi.UserTags = oi.UserDefined[xhttp.AmzObjectTagging]
return oi
}
func (z *erasureServerPools) replicaObjectInfo(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
copies, err := z.objectPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
if copies[0].ObjInfo.DeleteMarker {
return copies[0].ObjInfo, MethodNotAllowed{Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID}
}
return mergedPoolObjectInfo(copies), nil
}
// metadataPoolInfos resolves an unqualified metadata request to one logical
// version before collecting its copies, rather than mixing different versions.
func (z *erasureServerPools) metadataPoolInfos(ctx context.Context, bucket, object string, opts ObjectOptions) ([]PoolObjInfo, error) {
copies, err := z.objectPoolInfos(ctx, bucket, object, opts)
if err != nil {
return nil, err
}
if copies[0].ObjInfo.DeleteMarker {
return nil, MethodNotAllowed{Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID}
}
if opts.VersionID == "" {
opts.VersionID = copies[0].ObjInfo.VersionID
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
return z.objectPoolInfos(ctx, bucket, object, opts)
}
return copies, nil
}
func (z *erasureServerPools) updatePoolMetadata(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
copies, err := z.metadataPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
updated := mergedPoolObjectInfo(copies)
before := maps.Clone(updated.UserDefined)
if opts.EvalMetadataFn != nil {
if _, err := opts.EvalMetadataFn(&updated, nil); err != nil {
return ObjectInfo{}, err
}
}
changes := make(map[string]string)
for key, value := range updated.UserDefined {
if prior, ok := before[key]; !ok || prior != value {
changes[key] = value
}
}
for key := range before {
if _, ok := updated.UserDefined[key]; !ok {
changes[key] = ""
}
}
state := storedObjectLockState(updated.UserDefined)
opts.VersionID = updated.VersionID
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
opts.NoLock = true
opts.EvalMetadataFn = func(oi *ObjectInfo, _ error) (ReplicateDecision, error) {
maps.Copy(oi.UserDefined, changes)
replaceObjectLockMetadata(oi.UserDefined, state)
for _, key := range []string{xhttp.AmzObjectTagging, ReservedMetadataPrefixLower + TaggingTimestamp} {
value, exists := updated.UserDefined[key]
if exists || oi.UserDefined[key] != "" {
oi.UserDefined[key] = value
}
}
return ReplicateDecision{}, nil
}
var primary ObjectInfo
for _, copy := range copies {
oi, err := z.serverPools[copy.Index].PutObjectMetadata(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
if copy.Index == copies[0].Index {
primary = oi
}
}
return primary, nil
}
func reconcileStoredObjectTags(metadata, stored map[string]string) {
key := ReservedMetadataPrefixLower + TaggingTimestamp
stamp, err := time.Parse(time.RFC3339Nano, stored[key])
if err != nil {
return
}
incoming, err := time.Parse(time.RFC3339Nano, metadata[key])
if err != nil || !stamp.Before(incoming) {
metadata[key] = stored[key]
metadata[xhttp.AmzObjectTagging] = stored[xhttp.AmzObjectTagging]
}
}
// A restored version still owns its tier reference even while IsRemote is
// false. Only the last copy of a reference may schedule its contents for GC.
func sharesTierObject(oi ObjectInfo, copies []PoolObjInfo) bool {
ref := oi.TransitionedObject
if ref.Status != lifecycle.TransitionComplete {
return false
}
for _, copy := range copies {
other := copy.ObjInfo.TransitionedObject
if other.Status == lifecycle.TransitionComplete && ref.Tier == other.Tier && ref.Name == other.Name && ref.VersionID == other.VersionID {
return true
}
}
return false
}
// retireReplicaCopies runs only after committing a replacement. Failures are
// returned to the caller, so a stale copy cannot be hidden behind a successful
// response. Data movement owns its source cleanup and does not use this helper.
func (z *erasureServerPools) retireReplicaCopies(ctx context.Context, bucket, object string, keep int, oi ObjectInfo) error {
versionID := oi.VersionID
if versionID == "" {
versionID = nullVersionID
}
copies, err := z.objectPoolInfos(ctx, bucket, object, ObjectOptions{VersionID: versionID})
if err != nil {
return err
}
var retained []PoolObjInfo
for _, copy := range copies {
if copy.Index == keep {
retained = []PoolObjInfo{copy}
break
}
}
if len(retained) == 0 {
return VersionNotFound{Bucket: bucket, Object: decodeDirObject(object), VersionID: versionID}
}
for i, copy := range copies {
if copy.Index == keep {
continue
}
_, err := z.serverPools[copy.Index].DeleteObject(ctx, bucket, object,
ObjectOptions{
VersionID: versionID, NoLock: true, NoAuditLog: true,
SkipFreeVersion: sharesTierObject(copy.ObjInfo, retained) || sharesTierObject(copy.ObjInfo, copies[i+1:]),
})
if err != nil && !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
return err
}
}
return nil
}
// deleteObjectConditional evaluates the condition once against the logical
// version, then removes all its copies under the same lock as pooled writers.
func (z *erasureServerPools) deleteObjectConditional(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
copies, err := z.objectPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
primary := copies[0]
if opts.CheckPrecondFn(primary.ObjInfo) {
return ObjectInfo{}, PreConditionFailed{}
}
opts.CheckPrecondFn = nil
opts.NoLock = true
if opts.EvalRetentionBypassFn != nil || opts.EvalMetadataFn != nil {
versions, err := z.metadataPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
logical := mergedPoolObjectInfo(versions)
if opts.EvalRetentionBypassFn != nil {
if err := opts.EvalRetentionBypassFn(logical, nil); err != nil {
return ObjectInfo{}, err
}
opts.EvalRetentionBypassFn = nil
}
if opts.EvalMetadataFn != nil {
decision, err := opts.EvalMetadataFn(&logical, nil)
if err != nil {
return ObjectInfo{}, err
}
if decision.ReplicateAny() {
opts.SetDeleteReplicationState(decision, opts.VersionID)
}
opts.EvalMetadataFn = nil
}
}
if opts.VersionID == "" && (opts.Versioned || opts.VersionSuspended) {
// A single new delete marker hides the current version. Older versions
// remain history and must not be removed by an unqualified DELETE.
oi, err := z.serverPools[primary.Index].DeleteObject(ctx, bucket, object, opts)
oi.Name = decodeDirObject(object)
oi.replicationDecision = opts.DeleteReplication.ReplicateDecisionStr
return oi, err
}
// Retire non-authoritative copies first. If cleanup fails, retain the
// authoritative version and report the error instead of acknowledging a
// deletion that would expose an older copy.
for i := 1; i < len(copies); i++ {
candidate := copies[i]
deleteOpts := opts
deleteOpts.SkipFreeVersion = opts.SkipFreeVersion || sharesTierObject(candidate.ObjInfo, copies[:1]) || sharesTierObject(candidate.ObjInfo, copies[i+1:])
_, err := z.serverPools[candidate.Index].DeleteObject(ctx, bucket, object, deleteOpts)
if err != nil && !isErrObjectNotFound(err) && !isErrVersionNotFound(err) {
return ObjectInfo{}, err
}
}
oi, err := z.serverPools[primary.Index].DeleteObject(ctx, bucket, object, opts)
oi.Name = decodeDirObject(object)
oi.replicationDecision = opts.DeleteReplication.ReplicateDecisionStr
return oi, err
}
+682
View File
@@ -0,0 +1,682 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"errors"
"fmt"
"io"
"maps"
"strings"
"sync"
"testing"
"time"
madmin "github.com/minio/madmin-go/v3"
xhttp "github.com/minio/minio/internal/http"
)
func consistencyPools(t *testing.T) (*erasureServerPools, string) {
t.Helper()
ctx, cancel := context.WithCancel(t.Context())
obj, dirs, err := prepareErasurePoolsWithContext(ctx)
if err != nil {
cancel()
t.Fatal(err)
}
z := obj.(*erasureServerPools)
previous := newObjectLayerFn()
setObjectLayer(z)
t.Cleanup(func() { cancel(); z.Shutdown(context.Background()); removeRoots(dirs); setObjectLayer(previous) })
bucket := "pool-consistency"
if err := z.MakeBucket(t.Context(), bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
return z, bucket
}
func putConsistencyObject(t *testing.T, z *erasureServerPools, bucket, object string, pool int, body string, opts ObjectOptions) ObjectInfo {
t.Helper()
oi, err := z.serverPools[pool].PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString(body), int64(len(body)), "", ""), opts)
if err != nil {
t.Fatal(err)
}
return oi
}
func TestPoolsConditionalDeleteVersionSelection(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "split-versions"
old := putConsistencyObject(t, z, bucket, object, 1, "old", ObjectOptions{Versioned: true, MTime: time.Now().Add(-time.Hour)})
latest := putConsistencyObject(t, z, bucket, object, 0, "latest", ObjectOptions{Versioned: true})
_, err := z.DeleteObject(t.Context(), bucket, object, ObjectOptions{
Versioned: true, VersionID: old.VersionID, HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != old.ETag },
})
if err != nil {
t.Fatalf("delete addressed version in pool 1: %v", err)
}
if _, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: old.VersionID}); !isErrVersionNotFound(err) {
t.Errorf("addressed version survived: %v", err)
}
if oi, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); err != nil || oi.VersionID != latest.VersionID {
t.Errorf("delete changed the latest version: %+v, %v", oi, err)
}
}
func TestPoolsConditionalDeleteDuplicateVersion(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "duplicate-version"
oi := putConsistencyObject(t, z, bucket, object, 0, "payload", ObjectOptions{Versioned: true})
putConsistencyObject(t, z, bucket, object, 1, "payload", ObjectOptions{Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime})
_, err := z.DeleteObject(t.Context(), bucket, object, ObjectOptions{
Versioned: true, VersionID: oi.VersionID, HasIfMatch: true,
CheckPrecondFn: func(current ObjectInfo) bool { return current.ETag != oi.ETag },
})
if err != nil {
t.Fatal(err)
}
if _, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID}); !isErrVersionNotFound(err) {
t.Fatalf("success left a readable duplicate version: %v", err)
}
}
func TestPoolsConditionalDeleteReportsOtherPoolFailure(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "delete-failure"
putConsistencyObject(t, z, bucket, object, 1, "older", ObjectOptions{MTime: time.Now().Add(-time.Hour)})
oi := putConsistencyObject(t, z, bucket, object, 0, "latest", ObjectOptions{})
set := z.serverPools[1].getHashedSet(object)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = accessMoveDeleteFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: object}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
_, err := z.DeleteObject(t.Context(), bucket, object, ObjectOptions{
HasIfMatch: true, CheckPrecondFn: func(current ObjectInfo) bool { return current.ETag != oi.ETag },
})
if err == nil {
t.Fatal("delete reported success despite the other pool's write failure")
}
}
func TestPoolsConditionalDeleteSerializesPut(t *testing.T) {
testPoolsConditionalDeleteWriter(t, false)
}
func TestPoolsConditionalDeleteSerializesCompletion(t *testing.T) {
testPoolsConditionalDeleteWriter(t, true)
}
func testPoolsConditionalDeleteWriter(t *testing.T, multipart bool) {
z, bucket := consistencyPools(t)
const object = "concurrent-put"
initial := putConsistencyObject(t, z, bucket, object, 1, "before", ObjectOptions{})
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
reader := mustGetPutObjReader(t, bytes.NewBufferString("after"), 5, "", "")
destination := 1
write := func() error {
_, err := z.PutObject(ctx, bucket, object, reader, ObjectOptions{DataMovement: true, SrcPoolIdx: 0, DstPoolIdx: &destination})
return err
}
if multipart {
mp, err := z.serverPools[1].NewMultipartUpload(ctx, bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
part, err := z.serverPools[1].PutObjectPart(ctx, bucket, object, mp.UploadID, 1, reader, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
write = func() error {
_, err := z.CompleteMultipartUpload(ctx, bucket, object, mp.UploadID,
[]CompletePart{{PartNumber: 1, ETag: part.ETag}}, ObjectOptions{})
return err
}
}
checked, resume := make(chan struct{}), make(chan struct{})
var resumeOnce sync.Once
release := func() { resumeOnce.Do(func() { close(resume) }) }
defer release()
deleted := make(chan error, 1)
go func() {
_, err := z.DeleteObject(ctx, bucket, object, ObjectOptions{
HasIfMatch: true,
CheckPrecondFn: func(oi ObjectInfo) bool {
close(checked)
<-resume
return oi.ETag != initial.ETag
},
})
deleted <- err
}()
select {
case <-checked:
case err := <-deleted:
t.Fatalf("delete did not evaluate the precondition: %v", err)
case <-ctx.Done():
t.Fatal(ctx.Err())
}
written := make(chan error, 1)
go func() { written <- write() }()
var early bool
select {
case err := <-written:
early = true
if err != nil {
t.Errorf("concurrent PUT: %v", err)
}
case <-time.After(250 * time.Millisecond):
}
release()
if err := <-deleted; err != nil {
t.Fatal(err)
}
if !early {
if err := <-written; err != nil {
t.Fatal(err)
}
}
if early {
t.Error("PUT committed while DELETE was between its comparison and removal")
}
if _, err := z.GetObjectInfo(ctx, bucket, object, ObjectOptions{}); err != nil {
t.Errorf("matching the old ETag removed the concurrent replacement: %v", err)
}
}
type consistencyGateReader struct {
io.Reader
entered, resume chan struct{}
once sync.Once
release sync.Once
}
func (r *consistencyGateReader) Read(p []byte) (int, error) {
r.once.Do(func() { close(r.entered); <-r.resume })
return r.Reader.Read(p)
}
func TestPoolsReplicaSerializesMetadataAndHealing(t *testing.T) {
for _, heal := range []bool{false, true} {
t.Run(fmt.Sprintf("heal=%t", heal), func(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "metadata-race"
old, recent := "2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"
meta := poolLockMetadata("GOVERNANCE", "OFF", old, old)
oi := putConsistencyObject(t, z, bucket, object, 1, "data", ObjectOptions{Versioned: true, UserDefined: meta})
z.poolMetaMutex.Lock()
z.poolMeta.Pools[0].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
gate := &consistencyGateReader{Reader: strings.NewReader("data"), entered: make(chan struct{}), resume: make(chan struct{})}
release := func() { gate.release.Do(func() { close(gate.resume) }) }
defer release()
reader := mustGetPutObjReader(t, gate, 4, "", "")
written := make(chan error, 1)
go func() {
_, err := z.PutObject(ctx, bucket, object, reader, ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime,
ReplicaLockReconcile: true, UserDefined: maps.Clone(meta),
})
written <- err
}()
select {
case <-gate.entered:
case err := <-written:
t.Fatalf("write failed before consuming the body: %v", err)
case <-ctx.Done():
t.Fatal(ctx.Err())
}
mutated := make(chan error, 1)
go func() {
var err error
if heal {
_, err = z.healObjectInPool(ctx, z.serverPools[1].getHashedSet(object), bucket, object, oi.VersionID, madmin.HealOpts{})
} else {
_, err = z.PutObjectMetadata(ctx, bucket, object, ObjectOptions{
VersionID: oi.VersionID, MTime: oi.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
current.UserDefined[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = "ON"
current.UserDefined[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = recent
return ReplicateDecision{}, nil
},
})
}
mutated <- err
}()
var early bool
select {
case err := <-mutated:
early = true
t.Errorf("mutation completed during the paused replica write: %v", err)
case <-time.After(250 * time.Millisecond):
}
release()
if err := <-written; err != nil {
t.Fatal(err)
}
if !early {
if err := <-mutated; err != nil {
t.Fatal(err)
}
}
got, err := z.GetObjectInfo(ctx, bucket, object, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
if !heal && storedObjectLockState(got.UserDefined).legalHold != "ON" {
t.Error("replica write rolled back the newer metadata update")
}
})
}
}
func TestPoolsConditionalDeletePreservesVersionHistory(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "conditional-marker"
older := putConsistencyObject(t, z, bucket, object, 1, "older", ObjectOptions{Versioned: true, MTime: time.Now().Add(-time.Hour)})
latest := putConsistencyObject(t, z, bucket, object, 0, "latest", ObjectOptions{Versioned: true})
_, err := z.DeleteObject(t.Context(), bucket, object, ObjectOptions{
Versioned: true,
CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != older.ETag },
})
if !isErrPreconditionFailed(err) {
t.Fatalf("wrong latest ETag: %v", err)
}
marker, err := z.DeleteObject(t.Context(), bucket, object, ObjectOptions{
Versioned: true,
CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != latest.ETag },
})
if err != nil || !marker.DeleteMarker {
t.Fatalf("logical delete did not create a marker: %+v, %v", marker, err)
}
if got, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); !isErrObjectNotFound(err) || !got.DeleteMarker {
t.Errorf("delete marker did not hide the split object: %+v, %v", got, err)
}
for _, version := range []string{older.VersionID, latest.VersionID} {
if _, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: version}); err != nil {
t.Errorf("conditional marker removed history %s: %v", version, err)
}
}
}
func poolLockMetadata(retention, hold, retentionTime, holdTime string) map[string]string {
return map[string]string{
strings.ToLower(xhttp.AmzObjectLockMode): retention,
strings.ToLower(xhttp.AmzObjectLockRetainUntilDate): "2030-01-01T00:00:00Z",
strings.ToLower(xhttp.AmzObjectLockLegalHold): hold,
ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp: retentionTime,
ReservedMetadataPrefixLower + ObjectLockLegalHoldTimestamp: holdTime,
}
}
func TestPoolsReplicaIndependentLockWinners(t *testing.T) {
for _, tc := range []struct {
name string
multipart, rebalance bool
}{
{"put-decommission", false, false},
{"put-rebalance", false, true},
{"multipart-decommission", true, false},
{"multipart-rebalance", true, true},
} {
t.Run(tc.name, func(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "split-lock-state"
old, recent := "2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"
first := poolLockMetadata("GOVERNANCE", "OFF", recent, old)
second := poolLockMetadata("", "ON", old, recent)
oi := putConsistencyObject(t, z, bucket, object, 0, "data", ObjectOptions{Versioned: true, UserDefined: first})
putConsistencyObject(t, z, bucket, object, 1, "data", ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, UserDefined: second,
})
opts := ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime,
ReplicaLockReconcile: true, UserDefined: poolLockMetadata("", "OFF", old, old),
}
var uploadID string
var parts []CompletePart
if tc.multipart {
mp, err := z.serverPools[1].NewMultipartUpload(t.Context(), bucket, object, opts)
if err != nil {
t.Fatal(err)
}
uploadID = mp.UploadID
part, err := z.serverPools[1].PutObjectPart(t.Context(), bucket, object, uploadID, 1,
mustGetPutObjReader(t, bytes.NewBufferString("data"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
parts = []CompletePart{{PartNumber: 1, ETag: part.ETag}}
}
// The upload and both versions predate the routing change. Retention's
// winner remains in the draining pool; legal hold's winner is in pool 1.
if tc.rebalance {
z.rebalMu.Lock()
z.rebalMeta = &rebalanceMeta{PoolStats: []*rebalanceStats{
{Participating: true, Info: rebalanceInfo{Status: rebalStarted}}, {},
}}
z.rebalMu.Unlock()
} else {
z.poolMetaMutex.Lock()
z.poolMeta.Pools[0].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
}
var err error
if tc.multipart {
_, err = z.CompleteMultipartUpload(t.Context(), bucket, object, uploadID, parts,
ObjectOptions{Versioned: true, MTime: oi.ModTime, ReplicaLockReconcile: true})
} else {
_, err = z.PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString("data"), 4, "", ""), opts)
}
if err != nil {
t.Fatal(err)
}
got, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
state := storedObjectLockState(got.UserDefined)
if state.mode != "GOVERNANCE" || state.retentionTimestamp != recent || state.legalHold != "ON" || state.legalHoldTimestamp != recent {
t.Errorf("pooled read lost independently ordered lock state: %+v", state)
}
if _, err := z.serverPools[0].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID}); !isErrVersionNotFound(err) {
t.Errorf("successful replacement left its stale competing copy: %v", err)
}
})
}
}
func TestPoolsReplicaSoleDrainingOwner(t *testing.T) {
for _, rebalance := range []bool{false, true} {
for _, null := range []bool{false, true} {
t.Run(fmt.Sprintf("rebalance=%t/null=%t", rebalance, null), func(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "sole-owner"
old, recent := "2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"
meta := poolLockMetadata("", "ON", recent, recent)
delete(meta, strings.ToLower(xhttp.AmzObjectLockRetainUntilDate))
oi := putConsistencyObject(t, z, bucket, object, 0, "data", ObjectOptions{Versioned: !null, UserDefined: meta})
versionID := oi.VersionID
if null {
versionID = nullVersionID
// A different latest version must not donate its lock to the null version.
putConsistencyObject(t, z, bucket, object, 1, "other-version", ObjectOptions{Versioned: true})
}
if rebalance {
z.rebalMu.Lock()
z.rebalMeta = &rebalanceMeta{PoolStats: []*rebalanceStats{
{Participating: true, Info: rebalanceInfo{Status: rebalStarted}}, {},
}}
z.rebalMu.Unlock()
} else {
z.poolMetaMutex.Lock()
z.poolMeta.Pools[0].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
}
_, err := z.PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString("data"), 4, "", ""), ObjectOptions{
Versioned: !null, VersionSuspended: null, VersionID: versionID, MTime: oi.ModTime,
ReplicaLockReconcile: true, UserDefined: poolLockMetadata("GOVERNANCE", "OFF", old, old),
})
if err != nil {
t.Fatal(err)
}
got, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: versionID})
if err != nil {
t.Fatal(err)
}
state := storedObjectLockState(got.UserDefined)
if state.mode != "" || state.retainUntil != "" || state.retentionTimestamp != recent || state.legalHold != "ON" {
t.Errorf("lost a removal or legal hold from the sole draining owner: %+v", state)
}
})
}
}
}
func TestPoolsMetadataUpdateUsesMergedVersion(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "metadata-update"
old, recent := "2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"
oi := putConsistencyObject(t, z, bucket, object, 0, "data", ObjectOptions{
Versioned: true, UserDefined: poolLockMetadata("GOVERNANCE", "OFF", recent, old),
})
putConsistencyObject(t, z, bucket, object, 1, "data", ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, UserDefined: poolLockMetadata("", "ON", old, recent),
})
latest := putConsistencyObject(t, z, bucket, object, 1, "latest", ObjectOptions{Versioned: true})
called := 0
_, err := z.PutObjectMetadata(t.Context(), bucket, object, ObjectOptions{
VersionID: oi.VersionID, MTime: oi.ModTime,
EvalMetadataFn: func(current *ObjectInfo, _ error) (ReplicateDecision, error) {
called++
state := storedObjectLockState(current.UserDefined)
if state.mode != "GOVERNANCE" || state.legalHold != "ON" {
return ReplicateDecision{}, fmt.Errorf("metadata policy evaluated stale state: %+v", state)
}
current.UserDefined["custom-update"] = "preserved"
return ReplicateDecision{}, nil
},
})
if err != nil || called != 1 {
t.Fatalf("metadata update: %v; callback count %d", err, called)
}
for i, pool := range z.serverPools {
got, err := pool.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
state := storedObjectLockState(got.UserDefined)
if state.mode != "GOVERNANCE" || state.legalHold != "ON" || got.UserDefined["custom-update"] != "preserved" {
t.Errorf("pool %d did not receive the merged update: %+v", i, got.UserDefined)
}
}
if current, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); err != nil || current.VersionID != latest.VersionID {
t.Errorf("updating an older version changed the latest: %+v, %v", current, err)
}
}
func TestPoolsReplicaMetadataCopyReconcilesLockAndTags(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "replica-metadata-copy"
old, recent := "2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"
meta := poolLockMetadata("GOVERNANCE", "OFF", recent, old)
meta[xhttp.AmzObjectTagging] = "key=new"
meta[ReservedMetadataPrefixLower+TaggingTimestamp] = recent
oi := putConsistencyObject(t, z, bucket, object, 0, "data", ObjectOptions{Versioned: true, UserDefined: meta})
putConsistencyObject(t, z, bucket, object, 1, "data", ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, UserDefined: poolLockMetadata("", "ON", old, recent),
})
stale := oi
stale.metadataOnly = true
stale.UserDefined = maps.Clone(oi.UserDefined)
stale.UserDefined[xhttp.AmzObjectTagging] = "key=old"
stale.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = old
_, err := z.CopyObject(t.Context(), bucket, object, bucket, object, stale,
ObjectOptions{VersionID: oi.VersionID}, ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, ReplicaLockReconcile: true,
})
if err != nil {
t.Fatal(err)
}
got, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
state := storedObjectLockState(got.UserDefined)
if state.mode != "GOVERNANCE" || state.legalHold != "ON" || got.UserTags != "key=new" {
t.Errorf("metadata replication rolled back a newer field: %+v", got.UserDefined)
}
}
func TestPoolsReplicaCleanupFailureCanRetry(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "replica-cleanup-failure"
old, recent := "2026-09-09T09:00:00Z", "2026-09-09T10:00:00Z"
oi := putConsistencyObject(t, z, bucket, object, 0, "data", ObjectOptions{
Versioned: true, UserDefined: poolLockMetadata("GOVERNANCE", "ON", recent, recent),
})
z.poolMetaMutex.Lock()
z.poolMeta.Pools[0].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
set := z.serverPools[0].getHashedSet(object)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = accessMoveDeleteFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: object, version: oi.VersionID}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
write := func() error {
_, err := z.PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString("data"), 4, "", ""), ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, ReplicaLockReconcile: true,
UserDefined: poolLockMetadata("", "OFF", old, old),
})
return err
}
if err := write(); err == nil {
t.Fatal("replica replacement hid a competing-copy cleanup failure")
}
if got, err := z.serverPools[1].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID}); err != nil || storedObjectLockState(got.UserDefined).legalHold != "ON" {
t.Fatalf("cleanup failure lost the committed reconciled version: %+v, %v", got, err)
}
set.getDisks = getDisks
if err := write(); err != nil {
t.Fatalf("retry could not finish cleanup: %v", err)
}
if _, err := z.serverPools[0].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID}); !isErrVersionNotFound(err) {
t.Errorf("retry left the competing version: %v", err)
}
}
func TestPoolsRetiringCopyPreservesSharedTierObject(t *testing.T) {
for _, test := range []struct {
name string
deleting bool
failPrimaryDelete bool
differentRemote bool
restored bool
}{
{name: "metadata-copy"},
{name: "restored-metadata-copy", restored: true},
{name: "metadata-copy-distinct-reference", differentRemote: true},
{name: "failed-primary-delete", deleting: true, failPrimaryDelete: true},
{name: "failed-primary-delete-distinct-reference", deleting: true, failPrimaryDelete: true, differentRemote: true},
{name: "successful-delete", deleting: true},
} {
t.Run(test.name, func(t *testing.T) {
z, bucket := consistencyPools(t)
const object = "shared-tier-object"
metadata := map[string]string{
ReservedMetadataPrefixLower + TransitionStatus: "complete",
ReservedMetadataPrefixLower + TransitionTier: "TEST-TIER",
ReservedMetadataPrefixLower + TransitionedObjectName: "shared-remote-object",
ReservedMetadataPrefixLower + TransitionedVersionID: "shared-remote-version",
}
if test.restored {
metadata[xhttp.AmzRestore] = completedRestoreObj(time.Now().Add(time.Hour)).String()
}
oi := putConsistencyObject(t, z, bucket, object, 0, "data", ObjectOptions{Versioned: true, UserDefined: metadata})
secondaryMetadata := maps.Clone(metadata)
if test.differentRemote {
secondaryMetadata[ReservedMetadataPrefixLower+TransitionedObjectName] = "other-remote-object"
}
putConsistencyObject(t, z, bucket, object, 1, "data", ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, UserDefined: secondaryMetadata,
})
current, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID})
if err != nil || current.TransitionedObject.Status != "complete" || current.IsRemote() == test.restored {
t.Fatalf("fixture did not persist the tier reference: %+v, %v", current.TransitionedObject, err)
}
if test.failPrimaryDelete {
// The authoritative copy remains readable if its deletion fails.
// Retiring a secondary copy must not schedule its shared remote
// contents for garbage collection in that case.
set := z.serverPools[0].getHashedSet(object)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
for i := range faulty {
faulty[i] = accessMoveDeleteFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: object, version: oi.VersionID}
}
set.getDisks = func() []StorageAPI { return faulty }
defer func() { set.getDisks = getDisks }()
}
if test.deleting {
_, err := z.DeleteObject(t.Context(), bucket, object, ObjectOptions{
Versioned: true, VersionID: oi.VersionID,
CheckPrecondFn: func(info ObjectInfo) bool { return info.ETag != current.ETag },
})
if (err != nil) != test.failPrimaryDelete {
t.Fatalf("unexpected authoritative delete result: %v", err)
}
} else {
current.metadataOnly = true
current.UserDefined["metadata-update"] = "new"
_, err := z.CopyObject(t.Context(), bucket, object, bucket, object, current,
ObjectOptions{VersionID: oi.VersionID}, ObjectOptions{
Versioned: true, VersionID: oi.VersionID, MTime: oi.ModTime, ReplicaLockReconcile: true,
})
if err != nil {
t.Fatal(err)
}
}
_, err = z.serverPools[0].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: oi.VersionID})
primaryDeleted := test.deleting && !test.failPrimaryDelete
if primaryDeleted {
if !isErrVersionNotFound(err) {
t.Fatalf("authoritative copy survived successful delete: %v", err)
}
} else if err != nil {
t.Fatalf("lost retained authoritative copy: %v", err)
}
for pool, wantFree := range []bool{primaryDeleted, test.differentRemote} {
for _, disk := range z.serverPools[pool].getHashedSet(object).getDisks() {
data, err := disk.ReadAll(t.Context(), bucket, pathJoin(object, xlStorageFormatFile))
if errors.Is(err, errFileNotFound) && !wantFree {
continue
}
if err != nil {
t.Fatal(err)
}
versions, err := getFileInfoVersions(data, bucket, object, false)
if err != nil {
t.Fatal(err)
}
wantCount := 0
if wantFree {
wantCount = 1
}
if len(versions.FreeVersions) != wantCount {
t.Fatalf("pool %d has %d tier GC markers, want %d", pool, len(versions.FreeVersions), wantCount)
}
}
}
})
}
}
+5 -1
View File
@@ -23,6 +23,10 @@ import (
)
func prepareErasurePools() (ObjectLayer, []string, error) {
return prepareErasurePoolsWithContext(context.Background())
}
func prepareErasurePoolsWithContext(ctx context.Context) (ObjectLayer, []string, error) {
nDisks := 32
fsDirs, err := getRandomDisks(nDisks)
if err != nil {
@@ -32,7 +36,7 @@ func prepareErasurePools() (ObjectLayer, []string, error) {
pools := mustGetPoolEndpoints(0, fsDirs[:16]...)
pools = append(pools, mustGetPoolEndpoints(1, fsDirs[16:]...)...)
objLayer, _, err := initObjectLayer(context.Background(), pools)
objLayer, _, err := initObjectLayer(ctx, pools)
if err != nil {
removeRoots(fsDirs)
return nil, nil, err
+268 -72
View File
@@ -615,7 +615,17 @@ func (z *erasureServerPools) getPoolIdxExistingNoLock(ctx context.Context, bucke
})
}
func (z *erasureServerPools) getPoolIdxNoLock(ctx context.Context, bucket, object string, size int64) (idx int, err error) {
func (z *erasureServerPools) getPoolIdxNoLock(ctx context.Context, bucket, object string, size int64, dstPoolIdx *int) (idx int, err error) {
if dstPoolIdx != nil {
if *dstPoolIdx < 0 || *dstPoolIdx >= len(z.serverPools) {
return -1, errInvalidArgument
}
if z.IsSuspended(*dstPoolIdx) || z.IsPoolRebalancing(*dstPoolIdx) {
return -1, toObjectErr(errDiskFull)
}
return *dstPoolIdx, nil
}
idx, err = z.getPoolIdxExistingNoLock(ctx, bucket, object)
if err != nil && !isErrObjectNotFound(err) {
return idx, err
@@ -634,8 +644,25 @@ func (z *erasureServerPools) getPoolIdxNoLock(ctx context.Context, bucket, objec
// getPoolIdx returns the found previous object and its corresponding pool idx,
// if none are found falls back to most available space pool, this function is
// designed to be only used by PutObject, CopyObject (newObject creation) and NewMultipartUpload.
func (z *erasureServerPools) getPoolIdx(ctx context.Context, bucket, object string, size int64) (idx int, err error) {
func (z *erasureServerPools) getPoolIdx(ctx context.Context, bucket, object string, size int64, dstPoolIdx *int) (idx int, err error) {
return z.getWritePoolIdx(ctx, bucket, object, size, dstPoolIdx, false)
}
// getWritePoolIdx keeps the write-allocation policy when the caller already
// holds the pools-layer object lock and must not reacquire a set read lock.
func (z *erasureServerPools) getWritePoolIdx(ctx context.Context, bucket, object string, size int64, dstPoolIdx *int, noLock bool) (idx int, err error) {
if dstPoolIdx != nil {
if *dstPoolIdx < 0 || *dstPoolIdx >= len(z.serverPools) {
return -1, errInvalidArgument
}
if z.IsSuspended(*dstPoolIdx) || z.IsPoolRebalancing(*dstPoolIdx) {
return -1, toObjectErr(errDiskFull)
}
return *dstPoolIdx, nil
}
pinfo, _, err := z.getPoolInfoExistingWithOpts(ctx, bucket, object, ObjectOptions{
NoLock: noLock,
SkipDecommissioned: true,
SkipRebalancing: true,
})
@@ -656,6 +683,13 @@ func (z *erasureServerPools) getPoolIdx(ctx context.Context, bucket, object stri
return idx, nil
}
func dataMovementDstPool(opts ObjectOptions) *int {
if !opts.DataMovement {
return nil
}
return opts.DstPoolIdx
}
func (z *erasureServerPools) Shutdown(ctx context.Context) error {
g := errgroup.WithNErrs(len(z.serverPools))
@@ -1124,14 +1158,31 @@ func (z *erasureServerPools) PutObject(ctx context.Context, bucket string, objec
object = encodeDirObject(object)
if z.SinglePool() {
_, err := z.getPoolIdx(ctx, bucket, object, data.Size())
idx, err := z.getPoolIdx(ctx, bucket, object, data.Size(), dataMovementDstPool(opts))
if err != nil {
return ObjectInfo{}, err
}
if dataMovementDstPool(opts) != nil && idx == opts.SrcPoolIdx {
return ObjectInfo{}, DataMovementOverwriteErr{
Bucket: bucket, Object: object, VersionID: opts.VersionID,
Err: errDataMovementSrcDstPoolSame,
}
}
return z.serverPools[0].PutObject(ctx, bucket, object, data, opts)
}
idx, err := z.getPoolIdx(ctx, bucket, object, data.Size())
if !opts.NoLock {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return ObjectInfo{}, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
}
opts.NoLock = true
idx, err := z.getWritePoolIdx(ctx, bucket, object, data.Size(), dataMovementDstPool(opts), true)
if err != nil {
return ObjectInfo{}, err
}
@@ -1145,7 +1196,15 @@ func (z *erasureServerPools) PutObject(ctx context.Context, bucket string, objec
}
}
return z.serverPools[idx].PutObject(ctx, bucket, object, data, opts)
if opts.ReplicaLockReconcile || opts.DataMovement {
opts.ReplicaLockReconcile = true
opts.replicaObjectInfo = z.replicaObjectInfo
}
oi, err := z.serverPools[idx].PutObject(ctx, bucket, object, data, opts)
if err == nil && opts.ReplicaLockReconcile && !opts.DataMovement {
err = z.retireReplicaCopies(ctx, bucket, object, idx, oi)
}
return oi, err
}
func (z *erasureServerPools) deletePrefix(ctx context.Context, bucket string, prefix string) error {
@@ -1166,7 +1225,7 @@ func (z *erasureServerPools) DeleteObject(ctx context.Context, bucket string, ob
object = encodeDirObject(object)
}
// Acquire a write lock before deleting the object.
// Serialize logical deletion with pooled writes and metadata updates.
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalDeleteOperationTimeout)
if err != nil {
@@ -1179,6 +1238,32 @@ func (z *erasureServerPools) DeleteObject(ctx context.Context, bucket string, ob
return ObjectInfo{}, z.deletePrefix(ctx, bucket, object)
}
// Access-tier moves must recreate delete markers on the explicitly
// selected destination. The regular data-movement path discovers a pool
// from existing object state, which is ambiguous while both source and
// destination temporarily contain the version stack.
if dstPoolIdx := dataMovementDstPool(opts); dstPoolIdx != nil {
if *dstPoolIdx < 0 || *dstPoolIdx >= len(z.serverPools) {
return ObjectInfo{}, errInvalidArgument
}
if *dstPoolIdx == opts.SrcPoolIdx {
return ObjectInfo{}, DataMovementOverwriteErr{
Bucket: bucket, Object: decodeDirObject(object), VersionID: opts.VersionID,
Err: errDataMovementSrcDstPoolSame,
}
}
if z.IsSuspended(*dstPoolIdx) || z.IsPoolRebalancing(*dstPoolIdx) {
return ObjectInfo{}, toObjectErr(errDiskFull)
}
objInfo, err = z.serverPools[*dstPoolIdx].DeleteObject(ctx, bucket, object, opts)
objInfo.Name = decodeDirObject(object)
return objInfo, err
}
if !z.SinglePool() && opts.CheckPrecondFn != nil {
return z.deleteObjectConditional(ctx, bucket, object, opts)
}
gopts := opts
gopts.NoLock = true
@@ -1187,9 +1272,55 @@ func (z *erasureServerPools) DeleteObject(ctx context.Context, bucket string, ob
if _, ok := err.(InsufficientReadQuorum); ok {
return objInfo, InsufficientWriteQuorum{}
}
// A conditional (If-Match) delete addressing a specific version treats an
// absent key as an absent version. getPoolInfoExistingWithOpts strips
// VersionID, so a missing key surfaces ObjectNotFound here even for a
// version-scoped delete; normalize it to VersionNotFound (NoSuchVersion),
// matching this function's tail. The unconditional path is unchanged.
if opts.CheckPrecondFn != nil && opts.VersionID != "" && isErrObjectNotFound(err) {
return objInfo, VersionNotFound{Bucket: bucket, Object: object, VersionID: opts.VersionID}
}
return objInfo, err
}
// Evaluate the conditional (If-Match) precondition while the write lock
// acquired above is held, before the delete-marker short-circuit and before
// any version is removed, so the object cannot change between the check and
// the delete. This is scoped to a single erasure set (see the note at the
// lock above): only there do the delete lock and the write path share the
// same lock namespace, making the check-then-delete atomic.
if opts.CheckPrecondFn != nil {
// pinfo.ObjInfo is the current latest version. getPoolInfoExistingWithOpts
// intentionally strips VersionID, so for a version-scoped delete read the
// specifically addressed version and evaluate the precondition against it.
checkInfo := pinfo.ObjInfo
if opts.VersionID != "" {
vopts := opts
vopts.NoLock = true // delete lock already held above
vopts.CheckPrecondFn = nil
vi, verr := z.serverPools[pinfo.Index].GetObjectInfo(ctx, bucket, object, vopts)
if verr != nil && (!isErrMethodNotAllowed(verr) || !vi.DeleteMarker) {
// Genuine read failure for the addressed version: a missing
// version -> VersionNotFound (NoSuchVersion), read-quorum loss, etc.
return objInfo, verr
}
// verr is nil for a live version, or MethodNotAllowed with a populated
// delete-marker ObjectInfo when the addressed version is a delete
// marker. In the latter case evaluate the precondition against the
// marker, which fails any If-Match (-> 412), rather than surfacing 405.
checkInfo = vi
} else if checkInfo.Name == "" {
// The current state could not be read (e.g. read-quorum loss); refuse
// the conditional delete rather than act on an unverified precondition.
return objInfo, InsufficientReadQuorum{}
}
if opts.CheckPrecondFn(checkInfo) {
return objInfo, PreConditionFailed{}
}
// Precondition satisfied; lower layers must not re-evaluate it.
opts.CheckPrecondFn = nil
}
// Delete marker already present we are not going to create new delete markers.
if pinfo.ObjInfo.DeleteMarker && opts.VersionID == "" {
pinfo.ObjInfo.Name = decodeDirObject(object)
@@ -1268,8 +1399,12 @@ func (z *erasureServerPools) deleteObjectFromAllPools(ctx context.Context, bucke
// the delete call tries to clean up other pools during DeleteObject call.
objInfo = dobjects[0]
objInfo.Name = decodeDirObject(object)
err = derrs[0]
return objInfo, err
for _, derr := range derrs {
if derr != nil && !isErrObjectNotFound(derr) && !isErrVersionNotFound(derr) {
return objInfo, derr
}
}
return objInfo, derrs[0]
}
func (z *erasureServerPools) DeleteObjects(ctx context.Context, bucket string, objects []ObjectToDelete, opts ObjectOptions) ([]DeletedObject, []error) {
@@ -1357,10 +1492,43 @@ func (z *erasureServerPools) CopyObject(ctx context.Context, srcBucket, srcObjec
dstOpts.NoLock = true
}
poolIdx, err := z.getPoolIdxNoLock(ctx, dstBucket, dstObject, srcInfo.Size)
if !z.SinglePool() && cpSrcDstSame && srcInfo.metadataOnly && dstOpts.ReplicaLockReconcile && srcOpts.VersionID == dstOpts.VersionID {
copies, err := z.metadataPoolInfos(ctx, dstBucket, dstObject, dstOpts)
if err != nil {
return ObjectInfo{}, err
}
stored := mergedPoolObjectInfo(copies)
reconcileStoredObjectLock(srcInfo.UserDefined, storedObjectLockState(stored.UserDefined))
reconcileStoredObjectTags(srcInfo.UserDefined, stored.UserDefined)
idx := copies[0].Index
oi, err := z.serverPools[idx].CopyObject(ctx, srcBucket, srcObject, dstBucket, dstObject, srcInfo, srcOpts, dstOpts)
if err == nil {
err = z.retireReplicaCopies(ctx, dstBucket, dstObject, idx, oi)
}
return oi, err
}
poolIdx, err := z.getPoolIdxNoLock(ctx, dstBucket, dstObject, srcInfo.Size, dataMovementDstPool(dstOpts))
if err != nil {
return objInfo, err
}
if dataMovementDstPool(dstOpts) != nil && poolIdx == dstOpts.SrcPoolIdx {
return ObjectInfo{}, DataMovementOverwriteErr{
Bucket: dstBucket, Object: dstObject, VersionID: dstOpts.VersionID,
Err: errDataMovementSrcDstPoolSame,
}
}
if !z.SinglePool() && dstOpts.ReplicaLockReconcile {
dstOpts.replicaObjectInfo = z.replicaObjectInfo
if !dstOpts.DataMovement {
defer func() {
if err == nil {
err = z.retireReplicaCopies(ctx, dstBucket, dstObject, poolIdx, objInfo)
}
}()
}
}
// CopyObjectHandler predicts the outcome of this decision in
// copyRewritesObjectData(); keep the two in sync.
@@ -1396,6 +1564,8 @@ func (z *erasureServerPools) CopyObject(ctx context.Context, srcBucket, srcObjec
EncryptFn: dstOpts.EncryptFn,
WantChecksum: dstOpts.WantChecksum,
WantServerSideChecksumType: dstOpts.WantServerSideChecksumType,
ReplicaLockReconcile: dstOpts.ReplicaLockReconcile,
replicaObjectInfo: dstOpts.replicaObjectInfo,
}
return z.serverPools[poolIdx].PutObject(ctx, dstBucket, dstObject, srcInfo.PutObjReader, putOpts)
@@ -1807,29 +1977,41 @@ func (z *erasureServerPools) NewMultipartUpload(ctx context.Context, bucket, obj
}()
if z.SinglePool() {
return z.serverPools[0].NewMultipartUpload(ctx, bucket, object, opts)
}
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) || z.IsPoolRebalancing(idx) {
continue
}
result, err := pool.ListMultipartUploads(ctx, bucket, object, "", "", "", maxUploadsList)
idx, err := z.getPoolIdx(ctx, bucket, object, -1, dataMovementDstPool(opts))
if err != nil {
return nil, err
}
// If there is a multipart upload with the same bucket/object name,
// create the new multipart in the same pool, this will avoid
// creating two multiparts uploads in two different pools
if len(result.Uploads) != 0 {
return z.serverPools[idx].NewMultipartUpload(ctx, bucket, object, opts)
if dataMovementDstPool(opts) != nil && idx == opts.SrcPoolIdx {
return nil, DataMovementOverwriteErr{
Bucket: bucket, Object: object, VersionID: opts.VersionID,
Err: errDataMovementSrcDstPoolSame,
}
}
return z.serverPools[0].NewMultipartUpload(ctx, bucket, object, opts)
}
if dataMovementDstPool(opts) == nil {
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) || z.IsPoolRebalancing(idx) {
continue
}
result, err := pool.ListMultipartUploads(ctx, bucket, object, "", "", "", maxUploadsList)
if err != nil {
return nil, err
}
// If there is a multipart upload with the same bucket/object name,
// create the new multipart in the same pool, this will avoid
// creating two multiparts uploads in two different pools.
if len(result.Uploads) != 0 {
return z.serverPools[idx].NewMultipartUpload(ctx, bucket, object, opts)
}
}
}
// any parallel writes on the object will block for this poolIdx
// to return since this holds a read lock on the namespace.
idx, err := z.getPoolIdx(ctx, bucket, object, -1)
idx, err := z.getPoolIdx(ctx, bucket, object, -1, dataMovementDstPool(opts))
if err != nil {
return nil, err
}
@@ -1863,7 +2045,7 @@ func (z *erasureServerPools) PutObjectPart(ctx context.Context, bucket, object,
}
if z.SinglePool() {
_, err := z.getPoolIdx(ctx, bucket, object, data.Size())
_, err := z.getPoolIdx(ctx, bucket, object, data.Size(), dataMovementDstPool(opts))
if err != nil {
return PartInfo{}, err
}
@@ -2031,6 +2213,19 @@ func (z *erasureServerPools) CompleteMultipartUpload(ctx context.Context, bucket
}
}()
if !z.SinglePool() {
if !opts.NoLock {
lk := z.NewNSLock(bucket, encodeDirObject(object))
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return objInfo, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
}
opts.NoLock = true
}
// Hold write locks to verify uploaded parts, also disallows any
// parallel PutObjectPart() requests.
uploadIDLock := z.NewNSLock(bucket, pathJoin(object, uploadID))
@@ -2045,13 +2240,21 @@ func (z *erasureServerPools) CompleteMultipartUpload(ctx context.Context, bucket
return z.serverPools[0].CompleteMultipartUpload(ctx, bucket, object, uploadID, uploadedParts, opts)
}
if opts.ReplicaLockReconcile || opts.DataMovement {
opts.ReplicaLockReconcile = true
opts.replicaObjectInfo = z.replicaObjectInfo
}
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) {
continue
}
objInfo, err = pool.CompleteMultipartUpload(ctx, bucket, object, uploadID, uploadedParts, opts)
if err == nil {
return objInfo, nil
if opts.ReplicaLockReconcile && !opts.DataMovement {
err = z.retireReplicaCopies(ctx, bucket, encodeDirObject(object), idx, objInfo)
}
return objInfo, err
}
if _, ok := err.(InvalidUploadID); ok {
// upload id not found move to next pool
@@ -2073,6 +2276,11 @@ func (z *erasureServerPools) GetBucketInfo(ctx context.Context, bucket string, o
if err != nil {
return bucketInfo, toObjectErr(err, bucket)
}
// Physical existence/creation probes must not be overwritten by cached
// metadata, which can legitimately lack Created on an unmigrated bucket.
if opts.NoMetadata {
return bucketInfo, nil
}
meta, err := globalBucketMetadataSys.Get(bucket)
if err == nil {
@@ -2141,7 +2349,15 @@ func (z *erasureServerPools) DeleteBucket(ctx context.Context, bucket string, op
opts.Force = true
}
err := z.s3Peer.DeleteBucket(ctx, bucket, opts)
// Take the metadata writer lock before deleting anything. Failure or
// cancellation must leave both the bucket and its metadata intact.
ctx, unlock, err := lockBucketMetadata(ctx, z, bucket)
if err != nil {
return toObjectErr(err, bucket)
}
defer unlock()
err = z.s3Peer.DeleteBucket(ctx, bucket, opts)
if err == nil || isErrBucketNotFound(err) {
// If site replication is configured, hold on to deleted bucket state until sites sync
if opts.SRDeleteOp == MarkDelete {
@@ -2150,8 +2366,9 @@ func (z *erasureServerPools) DeleteBucket(ctx context.Context, bucket string, op
}
if err == nil {
// Purge the entire bucket metadata entirely.
z.deleteAll(context.Background(), minioMetaBucket, pathJoin(bucketMetaPrefix, bucket))
// Finish cleanup after a committed delete even if the client disconnects.
// Both the bucket-name and metadata locks remain held until return.
z.deleteAll(context.WithoutCancel(ctx), minioMetaBucket, pathJoin(bucketMetaPrefix, bucket))
}
return toObjectErr(err, bucket)
@@ -2569,7 +2786,7 @@ func (z *erasureServerPools) HealObject(ctx context.Context, bucket, object, ver
wg.Add(1)
go func(idx int, pool *erasureSets) {
defer wg.Done()
result, err := pool.HealObject(ctx, bucket, object, versionID, opts)
result, err := z.healObjectInPool(ctx, pool.getHashedSet(object), bucket, object, versionID, opts)
result.Object = decodeDirObject(result.Object)
errs[idx] = err
results[idx] = result
@@ -2838,15 +3055,7 @@ func (z *erasureServerPools) PutObjectMetadata(ctx context.Context, bucket, obje
defer lk.Unlock(lkctx)
}
opts.MetadataChg = true
opts.NoLock = true
// We don't know the size here set 1GiB at least.
idx, err := z.getPoolIdxExistingWithOpts(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
return z.serverPools[idx].PutObjectMetadata(ctx, bucket, object, opts)
return z.updatePoolMetadata(ctx, bucket, object, opts)
}
// PutObjectTags - replace or add tags to an existing object
@@ -2867,44 +3076,31 @@ func (z *erasureServerPools) PutObjectTags(ctx context.Context, bucket, object s
defer lk.Unlock(lkctx)
}
opts.MetadataChg = true
opts.NoLock = true
// We don't know the size here set 1GiB at least.
idx, err := z.getPoolIdxExistingWithOpts(ctx, bucket, object, opts)
copies, err := z.metadataPoolInfos(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
return z.serverPools[idx].PutObjectTags(ctx, bucket, object, tags, opts)
opts.NoLock = true
opts.VersionID = copies[0].ObjInfo.VersionID
if opts.VersionID == "" {
opts.VersionID = nullVersionID
}
var primary ObjectInfo
for _, copy := range copies {
oi, err := z.serverPools[copy.Index].PutObjectTags(ctx, bucket, object, tags, opts)
if err != nil {
return ObjectInfo{}, err
}
if copy.Index == copies[0].Index {
primary = oi
}
}
return primary, nil
}
// DeleteObjectTags - delete object tags from an existing object
func (z *erasureServerPools) DeleteObjectTags(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
object = encodeDirObject(object)
if z.SinglePool() {
return z.serverPools[0].DeleteObjectTags(ctx, bucket, object, opts)
}
if !opts.NoLock {
// Lock the object before deleting tags.
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
return ObjectInfo{}, err
}
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
}
opts.MetadataChg = true
opts.NoLock = true
idx, err := z.getPoolIdxExistingWithOpts(ctx, bucket, object, opts)
if err != nil {
return ObjectInfo{}, err
}
return z.serverPools[idx].DeleteObjectTags(ctx, bucket, object, opts)
return z.PutObjectTags(ctx, bucket, object, "", opts)
}
// GetObjectTags - get object tags from an existing object
@@ -2914,7 +3110,7 @@ func (z *erasureServerPools) GetObjectTags(ctx context.Context, bucket, object s
return z.serverPools[0].GetObjectTags(ctx, bucket, object, opts)
}
oi, _, err := z.getLatestObjectInfoWithIdx(ctx, bucket, object, opts)
oi, err := z.GetObjectInfo(ctx, bucket, object, opts)
if err != nil {
return nil, err
}
@@ -3023,7 +3219,7 @@ func (z *erasureServerPools) DecomTieredObject(ctx context.Context, bucket, obje
defer ns.Unlock(lkctx)
opts.NoLock = true
}
idx, err := z.getPoolIdxNoLock(ctx, bucket, object, fi.Size)
idx, err := z.getPoolIdxNoLock(ctx, bucket, object, fi.Size, dataMovementDstPool(opts))
if err != nil {
return err
}
+8 -2
View File
@@ -166,6 +166,12 @@ func (er *erasureObjects) healErasureSet(ctx context.Context, buckets []string,
if objAPI == nil {
return errServerNotInitialized
}
healInPool := er.HealObject
if z, ok := objAPI.(*erasureServerPools); ok && !z.SinglePool() {
healInPool = func(ctx context.Context, bucket, object, versionID string, opts madmin.HealOpts) (madmin.HealResultItem, error) {
return z.healObjectInPool(ctx, er, bucket, object, versionID, opts)
}
}
started := tracker.Started
if started.IsZero() || started.Equal(timeSentinel) {
@@ -419,7 +425,7 @@ func (er *erasureObjects) healErasureSet(ctx context.Context, buckets []string,
var result healEntryResult
fivs, err := entry.fileInfoVersions(bucket)
if err != nil {
res, err := er.HealObject(ctx, bucket, encodedEntryName, "",
res, err := healInPool(ctx, bucket, encodedEntryName, "",
madmin.HealOpts{
ScanMode: scanMode,
Remove: healDeleteDangling,
@@ -455,7 +461,7 @@ func (er *erasureObjects) healErasureSet(ctx context.Context, buckets []string,
continue
}
res, err := er.HealObject(ctx, bucket, encodedEntryName,
res, err := healInPool(ctx, bucket, encodedEntryName,
version.VersionID, madmin.HealOpts{
ScanMode: scanMode,
Remove: healDeleteDangling,
+4 -6
View File
@@ -51,9 +51,8 @@ func initGlobalGrid(ctx context.Context, eps EndpointServerPools) error {
grid.ContextDialer(xhttp.DialContextWithLookupHost(lookupHost, xhttp.NewInternodeDialContext(rest.DefaultTimeout, globalTCPOptions.ForWebsocket()))),
newCachedAuthToken(),
&tls.Config{
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
CurvePreferences: crypto.TLSCurveIDs(),
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
}),
Local: local,
Hosts: hosts,
@@ -84,9 +83,8 @@ func initGlobalLockGrid(ctx context.Context, eps EndpointServerPools) error {
grid.ContextDialer(xhttp.DialContextWithLookupHost(lookupHost, xhttp.NewInternodeDialContext(rest.DefaultTimeout, globalTCPOptions.ForWebsocket()))),
newCachedAuthToken(),
&tls.Config{
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
CurvePreferences: crypto.TLSCurveIDs(),
RootCAs: globalRootCAs,
CipherSuites: crypto.TLSCiphers(),
}, grid.RouteLockPath),
Local: local,
Hosts: hosts,
+529
View File
@@ -0,0 +1,529 @@
// Copyright 2026 PGSTY contributors.
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"bytes"
"context"
"crypto/md5"
"encoding/base64"
"encoding/xml"
"errors"
"fmt"
"io"
"maps"
"net/http"
"net/http/httptest"
"sync"
"testing"
"time"
"github.com/minio/minio/internal/dsync"
"github.com/minio/minio/internal/hash"
xhttp "github.com/minio/minio/internal/http"
)
func accessMovePools(t *testing.T) (*erasureServerPools, string) {
t.Helper()
ctx, cancel := context.WithCancel(t.Context())
dirs, err := getRandomDisks(32)
if err != nil {
t.Fatal(err)
}
endpoints := mustGetPoolEndpoints(0, dirs[:16]...)
endpoints = append(endpoints, mustGetPoolEndpoints(1, dirs[16:]...)...)
obj, _, err := initObjectLayer(ctx, endpoints)
if err != nil {
cancel()
removeRoots(dirs)
t.Fatal(err)
}
z := obj.(*erasureServerPools)
previous := newObjectLayerFn()
setObjectLayer(z)
t.Cleanup(func() { cancel(); z.Shutdown(context.Background()); removeRoots(dirs); setObjectLayer(previous) })
bucket := "access-move-test"
if err := z.MakeBucket(ctx, bucket, MakeBucketOptions{VersioningEnabled: true}); err != nil {
t.Fatal(err)
}
return z, bucket
}
func putAccessMoveVersion(t *testing.T, z *erasureServerPools, bucket, object string, pool int, body string, moved bool) ObjectInfo {
t.Helper()
metadata := map[string]string{"test-value": body}
if moved {
metadata[accessTierMetadataKey] = accessTierStamp(pool, time.Now().UnixNano())
}
cs := hash.NewChecksumFromData(hash.ChecksumCRC32C, []byte(body))
oi, err := z.serverPools[pool].PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString(body), int64(len(body)), "", ""),
ObjectOptions{Versioned: true, UserDefined: metadata, WantChecksum: cs})
if err != nil {
t.Fatal(err)
}
return oi
}
func assertAccessMoveVersion(t *testing.T, z *erasureServerPools, bucket, object string, pool int, want ObjectInfo, body string) {
t.Helper()
gr, err := z.serverPools[pool].GetObjectNInfo(t.Context(), bucket, object, nil, http.Header{}, ObjectOptions{VersionID: want.VersionID})
if err != nil {
t.Errorf("pool %d version %s: %v", pool, want.VersionID, err)
return
}
got, err := io.ReadAll(gr)
gr.Close()
if err != nil || string(got) != body {
t.Errorf("pool %d payload = %q, err = %v, want %q", pool, got, err, body)
}
if gr.ObjInfo.ETag != want.ETag || !gr.ObjInfo.ModTime.Equal(want.ModTime) {
t.Error("move changed ETag or modification time")
}
if len(want.Checksum) == 0 {
t.Fatal("fixture has no checksum")
}
if !bytes.Equal(gr.ObjInfo.Checksum, want.Checksum) {
t.Error("move changed or dropped the stored checksum")
}
if gr.ObjInfo.UserDefined["test-value"] != body {
t.Error("move changed user metadata")
}
}
func TestAccessMoveVersionStack(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "versions"
old := putAccessMoveVersion(t, z, bucket, object, 1, "old", false)
dm, err := z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
latest := putAccessMoveVersion(t, z, bucket, object, 1, "new", false)
for _, pair := range [][2]int{{1, 0}, {0, 1}} {
if _, err := moveObjectPool(t.Context(), z, bucket, object, pair[0], pair[1], nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, pair[1], old, "old")
assertAccessMoveVersion(t, z, bucket, object, pair[1], latest, "new")
versions, err := accessObjectVersions(t.Context(), z, pair[1], bucket, object)
if err != nil {
t.Fatal(err)
}
found := false
for _, version := range versions {
if version.VersionID == dm.VersionID && version.Deleted {
found = true
}
}
if !found || len(versions) != 3 {
t.Errorf("move lost a version or delete marker: %+v", versions)
}
if _, err := z.serverPools[pair[0]].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); !isErrObjectNotFound(err) {
t.Errorf("source remains after completed move: %v", err)
}
}
}
func TestAccessMovePreservesNestedObject(t *testing.T) {
z, bucket := accessMovePools(t)
parent := putAccessMoveVersion(t, z, bucket, "parent", 1, "parent", false)
child := putAccessMoveVersion(t, z, bucket, "parent/child", 1, "child", false)
if _, err := moveObjectPool(t.Context(), z, bucket, "parent", 1, 0, nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, "parent", 0, parent, "parent")
assertAccessMoveVersion(t, z, bucket, "parent/child", 1, child, "child")
}
func TestAccessMoveNullVersion(t *testing.T) {
z, bucket := accessMovePools(t)
const object, body = "null-version", "unversioned"
// A null version can remain in a bucket after versioning is enabled.
oi, err := z.serverPools[1].PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString(body), int64(len(body)), "", ""), ObjectOptions{
UserDefined: map[string]string{"test-value": body},
WantChecksum: hash.NewChecksumFromData(hash.ChecksumCRC32C, []byte(body)),
})
if err != nil || oi.VersionID != "" {
t.Fatalf("null version fixture failed: %v %q", err, oi.VersionID)
}
if _, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, oi, body)
if _, err := z.serverPools[1].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); !isErrObjectNotFound(err) {
t.Fatalf("null version source was not removed: %v", err)
}
}
// Source removal can partially commit before a process or disk fails. The
// destination can therefore hold the only remaining copy of an older version.
func TestAccessMoveRetryPreservesDestinationVersions(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "retry"
onlyAtDestination := putAccessMoveVersion(t, z, bucket, object, 0, "unique-old", true)
source := putAccessMoveVersion(t, z, bucket, object, 1, "source-new", false)
if _, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, onlyAtDestination, "unique-old")
assertAccessMoveVersion(t, z, bucket, object, 0, source, "source-new")
}
type accessMoveFaultDisk struct {
StorageAPI
failVersion string
}
func (d accessMoveFaultDisk) RenameData(ctx context.Context, srcVolume, srcPath string, fi FileInfo, dstVolume, dstPath string, opts RenameOptions) (RenameDataResp, error) {
if fi.VersionID == d.failVersion {
return RenameDataResp{}, errDiskFull
}
return d.StorageAPI.RenameData(ctx, srcVolume, srcPath, fi, dstVolume, dstPath, opts)
}
func TestAccessMoveWriteFailureCanResume(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "write-failure"
old := putAccessMoveVersion(t, z, bucket, object, 1, "old", false)
latest := putAccessMoveVersion(t, z, bucket, object, 1, "new", false)
existing := putAccessMoveVersion(t, z, bucket, object, 0, "unique-destination", true)
set := z.serverPools[0].getHashedSet(object)
getDisks := set.getDisks
disks := getDisks()
faulty := make([]StorageAPI, len(disks))
for i, disk := range disks {
faulty[i] = accessMoveFaultDisk{StorageAPI: disk, failVersion: latest.VersionID}
}
set.getDisks = func() []StorageAPI { return faulty }
_, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil)
set.getDisks = getDisks
if err == nil {
t.Fatal("injected destination write failure was ignored")
}
assertAccessMoveVersion(t, z, bucket, object, 1, old, "old")
assertAccessMoveVersion(t, z, bucket, object, 1, latest, "new")
assertAccessMoveVersion(t, z, bucket, object, 0, existing, "unique-destination")
if _, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, old, "old")
assertAccessMoveVersion(t, z, bucket, object, 0, latest, "new")
assertAccessMoveVersion(t, z, bucket, object, 0, existing, "unique-destination")
}
func TestAccessMoveExcludesSourceWriter(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "concurrent"
old := putAccessMoveVersion(t, z, bucket, object, 1, "old", false)
_, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, func(ObjectInfo, uint64) error {
ctx, cancel := context.WithTimeout(t.Context(), 200*time.Millisecond)
defer cancel()
_, err := z.serverPools[1].PutObject(ctx, bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString("racing-write"), 12, "", ""), ObjectOptions{Versioned: true})
if err == nil {
t.Error("source writer committed while move was in progress")
}
if ctx.Err() == nil {
return errors.New("writer did not wait for the move lock")
}
return nil
})
if err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, old, "old")
}
func TestAccessMoveResumesPartialCopy(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "partial-copy"
old := putAccessMoveVersion(t, z, bucket, object, 1, "old", false)
latest := putAccessMoveVersion(t, z, bucket, object, 1, "new", false)
gr, err := z.serverPools[1].GetObjectNInfo(t.Context(), bucket, object, nil, http.Header{}, ObjectOptions{VersionID: old.VersionID, NoDecryption: true})
if err != nil {
t.Fatal(err)
}
// Model a process stopping after the first copy commits, before cleanup.
if err := moveAccessTierVersion(t.Context(), z, 1, 0, bucket, gr, time.Now().UnixNano()); err != nil {
t.Fatal(err)
}
if _, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, old, "old")
assertAccessMoveVersion(t, z, bucket, object, 0, latest, "new")
}
func TestAccessMoveRefusesConflictingVersion(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "conflicting-version"
source := putAccessMoveVersion(t, z, bucket, object, 1, "source", false)
destination, err := z.serverPools[0].PutObject(t.Context(), bucket, object,
mustGetPutObjReader(t, bytes.NewBufferString("target"), 6, "", ""), ObjectOptions{
VersionID: source.VersionID, MTime: source.ModTime,
UserDefined: map[string]string{"test-value": "target"},
WantChecksum: hash.NewChecksumFromData(hash.ChecksumCRC32C, []byte("target")),
})
if err != nil {
t.Fatal(err)
}
if _, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil); !errors.Is(err, errAccessTierNotEligible) {
t.Fatalf("conflicting version should be left intact: %v", err)
}
assertAccessMoveVersion(t, z, bucket, object, 1, source, "source")
assertAccessMoveVersion(t, z, bucket, object, 0, destination, "target")
}
type accessMoveDeleteFaultDisk struct {
StorageAPI
bucket, object, version string
}
func (d accessMoveDeleteFaultDisk) DeleteVersion(ctx context.Context, volume, path string, fi FileInfo, forceDelMarker bool, opts DeleteOptions) error {
if volume == d.bucket && path == d.object && fi.VersionID == d.version {
return errDiskFull
}
return d.StorageAPI.DeleteVersion(ctx, volume, path, fi, forceDelMarker, opts)
}
func TestAccessMoveSourceDeleteFailureCanResume(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "source-delete-failure"
old := putAccessMoveVersion(t, z, bucket, object, 1, "old", false)
latest := putAccessMoveVersion(t, z, bucket, object, 1, "new", false)
set := z.serverPools[1].getHashedSet(object)
getDisks := set.getDisks
faulty := append([]StorageAPI(nil), getDisks()...)
// Commit removal of the old version, then reject the latest version on
// every disk. This fixes the retry boundary without relying on how a
// partially deleted erasure quorum is subsequently resolved or healed.
for i := range faulty {
faulty[i] = accessMoveDeleteFaultDisk{StorageAPI: faulty[i], bucket: bucket, object: object, version: latest.VersionID}
}
set.getDisks = func() []StorageAPI { return faulty }
_, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil)
set.getDisks = getDisks
if err == nil {
t.Fatal("injected partial source purge was ignored")
}
assertAccessMoveVersion(t, z, bucket, object, 0, old, "old")
assertAccessMoveVersion(t, z, bucket, object, 0, latest, "new")
if _, err := z.serverPools[1].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{VersionID: old.VersionID}); !isErrVersionNotFound(err) {
t.Fatalf("old source version should already be removed: %v", err)
}
assertAccessMoveVersion(t, z, bucket, object, 1, latest, "new")
if _, err := moveObjectPool(t.Context(), z, bucket, object, 1, 0, nil); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, old, "old")
assertAccessMoveVersion(t, z, bucket, object, 0, latest, "new")
}
// Exercise the public multipart API and compare whole-object and part reads
// across both directions of a real pool move, including raw SSE-C ciphertext.
func TestAccessMoveMultipartChecksums(t *testing.T) {
z, _ := accessMovePools(t)
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
bucket, router, err := initAPIHandlerTest(ctx, z, nil, MakeBucketOptions{VersioningEnabled: true})
if err != nil {
t.Fatal(err)
}
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
partData, full := multipartChecksumTestData()
for _, tc := range []struct {
typ hash.ChecksumType
ssec bool
}{{hash.ChecksumCRC32, false}, {hash.ChecksumCRC32C, false}, {hash.ChecksumCRC64NVME, false}, {hash.ChecksumCRC32C, true}} {
t.Run(fmt.Sprintf("%s/ssec=%t", tc.typ, tc.ssec), func(t *testing.T) {
object := getRandomObjectName()
headers := map[string]string{}
if tc.ssec {
key := bytes.Repeat([]byte{0x42}, 32)
keyMD5 := md5.Sum(key)
headers[xhttp.AmzServerSideEncryptionCustomerAlgorithm] = xhttp.AmzEncryptionAES
headers[xhttp.AmzServerSideEncryptionCustomerKey] = base64.StdEncoding.EncodeToString(key)
headers[xhttp.AmzServerSideEncryptionCustomerKeyMD5] = base64.StdEncoding.EncodeToString(keyMD5[:])
}
do := func(method, url string, body []byte, extra map[string]string) *httptest.ResponseRecorder {
t.Helper()
h := maps.Clone(headers)
maps.Copy(h, extra)
req, err := newTestSignedRequestV4(method, url, int64(len(body)), bytes.NewReader(body), globalActiveCred.AccessKey, globalActiveCred.SecretKey, h)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != http.StatusOK && rec.Code != http.StatusPartialContent {
t.Fatalf("%s %s: %d %s", method, url, rec.Code, rec.Body.String())
}
return rec
}
init := do(http.MethodPost, getNewMultipartURL("", bucket, object), nil, map[string]string{
xhttp.AmzChecksumAlgo: tc.typ.String(), xhttp.AmzChecksumType: xhttp.AmzChecksumTypeFullObject,
})
var upload InitiateMultipartUploadResponse
if err := xml.Unmarshal(init.Body.Bytes(), &upload); err != nil {
t.Fatal(err)
}
parts := make([]CompletePart, len(partData))
for i, data := range partData {
rec := do(http.MethodPut, getPutObjectPartURL("", bucket, object, upload.UploadID, fmt.Sprint(i+1)), data,
map[string]string{tc.typ.Key(): mustChecksum(t, tc.typ, data)})
parts[i] = CompletePart{PartNumber: i + 1, ETag: canonicalizeETag(rec.Header()[xhttp.ETag][0])}
}
body, err := xml.Marshal(CompleteMultipartUpload{Parts: parts})
if err != nil {
t.Fatal(err)
}
do(http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, upload.UploadID), body,
map[string]string{tc.typ.Key(): mustChecksum(t, tc.typ, full), xhttp.AmzChecksumType: xhttp.AmzChecksumTypeFullObject})
source, err := z.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil || len(source.Checksum) == 0 {
t.Fatalf("source checksum missing: %v", err)
}
src := 0
if _, err := z.serverPools[src].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); isErrObjectNotFound(err) {
src = 1
}
for i := 0; i < 2; i++ {
dst := 1 - src
if _, err := moveObjectPool(t.Context(), z, bucket, object, src, dst, nil); err != nil {
t.Fatal(err)
}
got, err := z.serverPools[dst].GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil || !bytes.Equal(got.Checksum, source.Checksum) || got.ETag != source.ETag || got.VersionID != source.VersionID || !got.ModTime.Equal(source.ModTime) {
t.Fatalf("move changed version metadata: %v checksumEqual=%t ETag=%q/%q version=%q/%q modTimeEqual=%t", err, bytes.Equal(got.Checksum, source.Checksum), got.ETag, source.ETag, got.VersionID, source.VersionID, got.ModTime.Equal(source.ModTime))
}
rec := do(http.MethodGet, getGetObjectURL("", bucket, object), nil, map[string]string{xhttp.AmzChecksumMode: "ENABLED"})
if !bytes.Equal(rec.Body.Bytes(), full) || rec.Header().Get(tc.typ.Key()) != mustChecksum(t, tc.typ, full) {
t.Fatal("GET payload or checksum changed after move")
}
for j, data := range partData {
rec := do(http.MethodGet, getGetObjectURL("", bucket, object)+fmt.Sprintf("?partNumber=%d", j+1), nil, nil)
if !bytes.Equal(rec.Body.Bytes(), data) {
t.Fatalf("part %d changed after move", j+1)
}
}
src = dst
}
})
}
}
type accessMoveNamedLocker struct {
*localLocker
address string
}
func (l accessMoveNamedLocker) String() string { return l.address }
func TestAccessMoveSharedDistributedLockers(t *testing.T) {
// Three overlapping sets of real dsync lock servers. Taking independent
// quorum locks for the same object would contend with our own earlier lock.
peers := make([]dsync.NetLocker, 5)
for i := range peers {
peers[i] = accessMoveNamedLocker{localLocker: newLocker(), address: fmt.Sprint(i)}
}
z := &erasureServerPools{}
for i := 0; i < 3; i++ {
setPeers := peers[i : i+3]
set := &erasureObjects{nsMutex: &nsLockMap{isDistErasure: true}, getLockers: func() ([]dsync.NetLocker, string) {
return setPeers, "access-move-fixture"
}}
z.serverPools = append(z.serverPools, &erasureSets{sets: []*erasureObjects{set}, distributionAlgo: formatErasureVersionV3DistributionAlgoV3})
}
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
_, unlock, err := lockAccessTierObject(ctx, z, "bucket", "object", 1, 2)
if err != nil {
t.Fatal(err)
}
release := sync.OnceFunc(unlock)
defer release()
for i, pool := range z.serverPools {
ctx, cancel := context.WithTimeout(t.Context(), 100*time.Millisecond)
lock := pool.NewNSLock("bucket", "object")
lc, err := lock.GetLock(ctx, globalOperationTimeout)
cancel()
if err == nil {
lock.Unlock(lc)
t.Fatalf("pool %d writer acquired its quorum during the move", i)
}
}
release()
// Every acquired peer lock is released, so ordinary writes can resume.
for _, peer := range peers {
// Distributed Unlock sends releases asynchronously.
lock := (&nsLockMap{isDistErasure: true}).NewNSLock(func() ([]dsync.NetLocker, string) {
return []dsync.NetLocker{peer}, "access-move-fixture"
}, "bucket", "object")
ctx, cancel := context.WithTimeout(t.Context(), 2*time.Second)
lc, err := lock.GetLock(ctx, globalOperationTimeout)
cancel()
if err != nil {
t.Fatalf("peer %s leaked a lock: %v", peer.String(), err)
}
lock.Unlock(lc)
}
}
func TestAccessMoveConfiguredPromotionAndDemotion(t *testing.T) {
z, bucket := accessMovePools(t)
const object = "configured-move"
version := putAccessMoveVersion(t, z, bucket, object, 1, "configured", false)
oldCfg, oldTracker := globalILMConfig.accessCfg(), globalAccessTracker
defer func() { globalILMConfig.update(oldCfg); globalAccessTracker = oldTracker }()
cfg := oldCfg
cfg.AccessTiering, cfg.AccessPools = true, []int{0, 1}
cfg.AccessMinResidency, cfg.AccessPromoteWatermark = 0, 99
globalILMConfig.update(cfg)
globalAccessTracker = newAccessTracker()
lc := accessLifecycleForTest(t)
data, err := xml.Marshal(lc)
if err != nil {
t.Fatal(err)
}
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketLifecycleConfig, data); err != nil {
t.Fatal(err)
}
state := newAccessTierState(t.Context())
state.z, state.objAPI = z, z
task := accessTierTask{ctx: t.Context(), bucket: bucket, object: object, src: 1, dst: 0, direction: accessTierPromote, bytes: uint64(version.Size)}
// A queued move is rechecked when configuration is disabled dynamically.
cfg.AccessTiering = false
globalILMConfig.update(cfg)
if err := state.processTask(task); !errors.Is(err, errAccessTierNotEligible) {
t.Fatalf("disabled mover = %v", err)
}
cfg.AccessTiering = true
globalILMConfig.update(cfg)
globalAccessTracker.merged.Store(&mergedAccess{binWidth: 60, entries: map[string]accessEntry{
accessKey(bucket, object): {Bins: []uint32{100}, LastAt: time.Now().Unix()},
}})
if reason, ok := state.reservePromotion(task, cfg, 0); !ok {
t.Fatalf("promotion reservation failed: %s", reason)
}
if err := state.processTask(task); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 0, version, "configured")
// Advance the counter view beyond the rule's idle period.
globalAccessTracker.merged.Store(&mergedAccess{binWidth: 60, entries: map[string]accessEntry{
accessKey(bucket, object): {Bins: []uint32{0}, LastAt: time.Now().Add(-2 * time.Hour).Unix()},
}})
task.src, task.dst, task.direction, task.bytes = 0, 1, accessTierDemote, 0
if err := state.processTask(task); err != nil {
t.Fatal(err)
}
assertAccessMoveVersion(t, z, bucket, object, 1, version, "configured")
if state.promotions.Load() != 1 || state.demotions.Load() != 1 || state.bytesMoved.Load() != 2*uint64(version.Size) || len(state.pending) != 0 || len(state.reserved) != 0 {
t.Fatal("promotion/demotion metrics or reservation cleanup are incorrect")
}
}
File diff suppressed because it is too large Load Diff
+263
View File
@@ -0,0 +1,263 @@
// Copyright (c) 2015-2026 MinIO, Inc.
package cmd
import (
"bytes"
"context"
"errors"
"strconv"
"strings"
"testing"
"time"
"github.com/klauspost/compress/zstd"
"github.com/minio/minio/internal/bucket/lifecycle"
"github.com/minio/minio/internal/config/ilm"
"github.com/tinylib/msgp/msgp"
)
func accessLifecycleForTest(t *testing.T) *lifecycle.Lifecycle {
t.Helper()
xml := "<LifecycleConfiguration>" +
"<Rule><ID>access</ID><Status>Enabled</Status>" +
"<AccessTransition><Window>10m</Window><PromoteAfterAccesses>100</PromoteAfterAccesses>" +
"<DemoteAfterAccesses>5</DemoteAfterAccesses><DemoteAfterIdle>1h</DemoteAfterIdle></AccessTransition>" +
"</Rule></LifecycleConfiguration>"
lc, err := lifecycle.ParseLifecycleConfig(strings.NewReader(xml))
if err != nil {
t.Fatal(err)
}
return lc
}
func TestAccessTierStampAndDemoteEligibility(t *testing.T) {
oldTracker := globalAccessTracker
globalAccessTracker = newAccessTracker()
t.Cleanup(func() { globalAccessTracker = oldTracker })
now := time.Now()
cfg := ilm.Config{
AccessTiering: true, AccessPools: []int{0, 1},
AccessBinWidth: time.Minute, AccessBins: 12,
AccessMinResidency: 30 * time.Minute,
}
oi := ObjectInfo{
Bucket: "bucket", Name: "object", Size: 10, IsLatest: true,
UserDefined: map[string]string{
accessTierMetadataKey: "0:" + strconv.FormatInt(now.Add(-2*time.Hour).UnixNano(), 10),
},
}
if !accessDemoteEligible(accessLifecycleForTest(t), oi, 0, cfg, now) {
t.Fatal("cold, resident object should be demotion eligible")
}
oi.UserDefined[accessTierMetadataKey] = "0:" + strconv.FormatInt(now.Add(-time.Minute).UnixNano(), 10)
if accessDemoteEligible(accessLifecycleForTest(t), oi, 0, cfg, now) {
t.Fatal("object inside minimum residency was eligible")
}
oi.UserDefined[accessTierMetadataKey] = "broken"
if accessDemoteEligible(accessLifecycleForTest(t), oi, 0, cfg, now) {
t.Fatal("object with malformed marker was eligible")
}
}
func TestAccessTierReservationsEnforceCaps(t *testing.T) {
state := newAccessTierState(t.Context())
state.usageReady = true
state.baseUsage["a"] = 80
state.baseTotal = 180
cfg := ilm.Config{AccessMaxSize: 200}
task := accessTierTask{ctx: t.Context(), bucket: "a", object: "one", bytes: 30}
if reason, ok := state.reservePromotion(task, cfg, 0); ok || reason != "max-size" {
t.Fatalf("max-size reserve = %q/%v", reason, ok)
}
cfg.AccessMaxSize = 0
if reason, ok := state.reservePromotion(task, cfg, 100); ok || reason != "quota" {
t.Fatalf("quota reserve = %q/%v", reason, ok)
}
task.bytes = 20
if reason, ok := state.reservePromotion(task, cfg, 100); !ok || reason != "" {
t.Fatalf("valid reserve = %q/%v", reason, ok)
}
if got := state.bucketUsageLocked("a"); got != 100 {
t.Fatalf("reserved bucket usage = %d, want 100", got)
}
state.releasePending(task, true)
}
func TestDataMovementDestinationIsGated(t *testing.T) {
dst := 0
if got := dataMovementDstPool(ObjectOptions{DstPoolIdx: &dst}); got != nil {
t.Fatal("destination honored without DataMovement")
}
if got := dataMovementDstPool(ObjectOptions{DataMovement: true, DstPoolIdx: &dst}); got == nil || *got != 0 {
t.Fatalf("destination = %v", got)
}
}
func TestForcedDestinationPoolSelection(t *testing.T) {
z := &erasureServerPools{serverPools: make([]*erasureSets, 2)}
dst := 1
if got, err := z.getPoolIdx(context.Background(), "bucket", "object", 1, &dst); err != nil || got != dst {
t.Fatalf("destination = %d, err = %v", got, err)
}
bad := 2
if _, err := z.getPoolIdx(context.Background(), "bucket", "object", 1, &bad); !errors.Is(err, errInvalidArgument) {
t.Fatalf("out-of-range error = %v", err)
}
z.poolMeta.Pools = make([]PoolStatus, 2)
z.poolMeta.Pools[1].Decommission = &PoolDecommissionInfo{}
if _, err := z.getPoolIdx(context.Background(), "bucket", "object", 1, &dst); err == nil {
t.Fatal("suspended destination was accepted")
}
}
func TestHotTierAccounting(t *testing.T) {
entry := dataUsageEntry{}
entry.addSizes(sizeSummary{totalSize: 100, hotTierSize: 60, versions: 1})
entry.merge(dataUsageEntry{Size: 50, HotTierSize: 25})
if entry.Size != 150 || entry.HotTierSize != 85 {
t.Fatalf("usage = size:%d hot:%d", entry.Size, entry.HotTierSize)
}
}
func TestScannerCyclesApart(t *testing.T) {
tests := []struct {
current, previous uint32
want uint32
}{
{current: 12, previous: 10, want: 2},
{current: 0, previous: ^uint32(0), want: 1},
// A cycle counter reset is not evidence that a complete pass covered
// recent moves, so it must not release conservative deltas.
{current: 1, previous: 100, want: 0},
}
for _, tt := range tests {
if got := scannerCyclesApart(tt.current, tt.previous); got != tt.want {
t.Fatalf("scannerCyclesApart(%d, %d) = %d, want %d", tt.current, tt.previous, got, tt.want)
}
}
}
func TestDataUsageCacheV8Migration(t *testing.T) {
var encoded bytes.Buffer
if err := encoded.WriteByte(dataUsageCacheVerV8); err != nil {
t.Fatal(err)
}
zw, err := zstd.NewWriter(&encoded)
if err != nil {
t.Fatal(err)
}
mw := msgp.NewWriter(zw)
if err = mw.WriteMapHeader(2); err != nil {
t.Fatal(err)
}
if err = mw.WriteString("Info"); err != nil {
t.Fatal(err)
}
info := dataUsageCacheInfo{Name: dataUsageRoot, NextCycle: 17, LastUpdate: time.Now().UTC()}
if err = info.EncodeMsg(mw); err != nil {
t.Fatal(err)
}
if err = mw.WriteString("Cache"); err != nil {
t.Fatal(err)
}
if err = mw.WriteMapHeader(1); err != nil {
t.Fatal(err)
}
if err = mw.WriteString("entry"); err != nil {
t.Fatal(err)
}
// A v8 entry has no hts field. Encoding only populated fields also
// verifies that its map decoder retains normal msgpack compatibility.
if err = mw.WriteMapHeader(2); err != nil {
t.Fatal(err)
}
if err = mw.WriteString("sz"); err != nil {
t.Fatal(err)
}
if err = mw.WriteInt64(123); err != nil {
t.Fatal(err)
}
if err = mw.WriteString("os"); err != nil {
t.Fatal(err)
}
if err = mw.WriteUint64(2); err != nil {
t.Fatal(err)
}
if err = mw.Flush(); err != nil {
t.Fatal(err)
}
if err = zw.Close(); err != nil {
t.Fatal(err)
}
var migrated dataUsageCache
if err = migrated.deserialize(bytes.NewReader(encoded.Bytes())); err != nil {
t.Fatal(err)
}
entry := migrated.Cache["entry"]
if entry.Size != 123 || entry.Objects != 2 || entry.HotTierSize != 0 {
t.Fatalf("migrated entry = size:%d objects:%d hot:%d", entry.Size, entry.Objects, entry.HotTierSize)
}
if migrated.Info.NextCycle != 17 || !migrated.Info.LastUpdate.Equal(info.LastUpdate) {
t.Fatalf("migrated cache info = %+v", migrated.Info)
}
}
// The stamp written on every moved version and the stamp the rollback path
// matches against must agree, otherwise rollback silently skips its own work.
func TestAccessTierStampRoundTrip(t *testing.T) {
movedAt := time.Now().UnixNano()
stamp := accessTierStamp(3, movedAt)
pool, at, ok := parseAccessTierStamp(map[string]string{accessTierMetadataKey: stamp})
if !ok {
t.Fatalf("stamp %q did not parse", stamp)
}
if pool != 3 {
t.Fatalf("pool = %d, want 3", pool)
}
if at.UnixNano() != movedAt {
t.Fatalf("movedAt = %d, want %d", at.UnixNano(), movedAt)
}
// A different move of the same object must not match, so a concurrent
// client overwrite is never mistaken for our own copy.
if stamp == accessTierStamp(3, movedAt+1) {
t.Fatal("stamps from distinct moves collided")
}
if stamp == accessTierStamp(4, movedAt) {
t.Fatal("stamps from distinct pools collided")
}
}
func TestAccessHitsMightPromote(t *testing.T) {
lc := accessLifecycleForTest(t)
hot := accessEntry{Bins: []uint32{100}, HeadAt: 0}
cold := accessEntry{Bins: []uint32{1}, HeadAt: 0}
if !accessHitsMightPromote(lc, "object", hot, 60) {
t.Fatal("object above promote threshold was rejected")
}
if accessHitsMightPromote(lc, "object", cold, 60) {
t.Fatal("object below promote threshold was accepted")
}
xml := "<LifecycleConfiguration><Rule><ID>logs</ID><Status>Enabled</Status>" +
"<Filter><Prefix>logs/</Prefix></Filter>" +
"<AccessTransition><Window>10m</Window><PromoteAfterAccesses>100</PromoteAfterAccesses>" +
"<DemoteAfterAccesses>5</DemoteAfterAccesses><DemoteAfterIdle>1h</DemoteAfterIdle></AccessTransition>" +
"</Rule></LifecycleConfiguration>"
prefixed, err := lifecycle.ParseLifecycleConfig(strings.NewReader(xml))
if err != nil {
t.Fatal(err)
}
if accessHitsMightPromote(prefixed, "data/object", hot, 60) {
t.Fatal("object outside the rule prefix was accepted")
}
if !accessHitsMightPromote(prefixed, "logs/object", hot, 60) {
t.Fatal("object inside the rule prefix was rejected")
}
}
+598
View File
@@ -0,0 +1,598 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"cmp"
"context"
"fmt"
"slices"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/minio/minio/internal/config/ilm"
"github.com/zeebo/xxh3"
)
//go:generate msgp -file=$GOFILE -unexported
//msgp:ignore accessTracker mergedAccess
const (
// accessTrackerPrefix is where each node publishes its own counters.
// One object per node, merged by every node on a timer.
accessTrackerPrefix = minioConfigPrefix + "/ilm/access"
// accessQueueSize bounds the GET -> tracker handoff. Overflow drops
// samples rather than slowing down reads.
accessQueueSize = 100000
// accessMaxDemoteCandidates bounds what the scanner may hand over in a
// single flush interval.
accessMaxDemoteCandidates = 100000
// accessShardStaleFactor multiplies the flush interval to decide when a
// peer's counters are too old to trust, e.g. after a node is removed.
accessShardStaleFactor = 5
)
func accessShardFresh(now int64, shard accessShard, stale int64) bool {
if shard.UpdatedAt <= 0 {
return false
}
if stale <= 0 {
return true
}
age := now - shard.UpdatedAt
return age >= -stale && age <= stale
}
// accessEntry is a rolling hit counter for one object. Bins[0] is the current
// bin and each subsequent bin is one bin-width older, so a rule asking for
// "100 hits in 10 minutes" sums the newest ceil(10m/binWidth) bins.
//
// A fixed window is used rather than an exponentially decayed score because
// the rule is stated to operators in exactly those terms.
type accessEntry struct {
Bins []uint32 `msg:"b"`
HeadAt int64 `msg:"h"` // unix seconds at the start of Bins[0]
LastAt int64 `msg:"l"` // unix seconds of the most recent hit
}
// demoteCandidate is an object the scanner found sitting on a fast pool with
// no recent reads. Candidates ride the node's own counter shard so they reach
// the leader without a new peer RPC.
type demoteCandidate struct {
Bucket string `msg:"b"`
Object string `msg:"o"`
Pool int `msg:"p"`
}
// accessShard is what one node publishes. BinWidth is carried so a peer that
// has not yet picked up a configuration change is ignored rather than merged
// with mismatched bins.
type accessShard struct {
UpdatedAt int64 `msg:"u"`
BinWidth int64 `msg:"bw"`
Entries map[string]accessEntry `msg:"e"`
Demote []demoteCandidate `msg:"d"`
}
// binStart truncates a unix timestamp to the start of its bin.
func binStart(now, binWidth int64) int64 {
if binWidth <= 0 {
return now
}
return now - now%binWidth
}
// rollTo advances the counter to now, zeroing the bins that elapsed since the
// last update and resizing if the configured bin count changed.
func (e *accessEntry) rollTo(now, binWidth int64, nbins int) {
if nbins <= 0 || binWidth <= 0 {
return
}
if len(e.Bins) != nbins {
resized := make([]uint32, nbins)
copy(resized, e.Bins)
e.Bins = resized
}
head := binStart(now, binWidth)
if e.HeadAt == 0 {
e.HeadAt = head
return
}
steps := (head - e.HeadAt) / binWidth
if steps <= 0 {
return
}
if steps >= int64(nbins) {
clear(e.Bins)
} else {
copy(e.Bins[steps:], e.Bins[:nbins-int(steps)])
clear(e.Bins[:steps])
}
e.HeadAt = head
}
// hits returns the number of accesses recorded over the newest bins covering
// window. A window longer than the configured history is clamped to it.
//
// The newest bin is partial, so the covered span is between window-binWidth
// and window. Operators tune resolution with ilm access_bin_width.
func (e accessEntry) hits(window time.Duration, binWidth int64) uint64 {
if binWidth <= 0 || len(e.Bins) == 0 {
return 0
}
n := int((int64(window/time.Second) + binWidth - 1) / binWidth)
if n < 1 {
n = 1
}
if n > len(e.Bins) {
n = len(e.Bins)
}
var total uint64
for _, v := range e.Bins[:n] {
total += uint64(v)
}
return total
}
// total is the whole retained history, used to decide what to evict.
func (e accessEntry) total() uint64 {
var t uint64
for _, v := range e.Bins {
t += uint64(v)
}
return t
}
// mergeFrom adds another node's counters for the same object. Both sides must
// already be rolled to the same head.
func (e *accessEntry) mergeFrom(o accessEntry) {
for i := range e.Bins {
if i < len(o.Bins) {
total := uint64(e.Bins[i]) + uint64(o.Bins[i])
if total > uint64(^uint32(0)) {
total = uint64(^uint32(0))
}
e.Bins[i] = uint32(total)
}
}
if o.LastAt > e.LastAt {
e.LastAt = o.LastAt
}
}
// mergedAccess is an immutable cluster-wide snapshot. Readers take it from an
// atomic pointer, so the hot scanner and sweep paths never take a lock.
type mergedAccess struct {
entries map[string]accessEntry
binWidth int64
at int64
}
func (m *mergedAccess) hits(key string, window time.Duration) uint64 {
if m == nil {
return 0
}
e, ok := m.entries[key]
if !ok {
return 0
}
return e.hits(window, m.binWidth)
}
func (m *mergedAccess) lastAccess(key string) int64 {
if m == nil {
return 0
}
return m.entries[key].LastAt
}
// accessTracker records how often each object is read.
//
// Ownership is deliberately narrow: the live counter map is touched only by
// run()'s goroutine, so it needs no lock. Everything read from elsewhere goes
// through the immutable merged snapshot.
type accessTracker struct {
ch chan string
enabled atomic.Bool
merged atomic.Pointer[mergedAccess]
// Demote candidates arrive from scanner goroutines, so this one does
// need a lock. It is small: only objects we previously promoted.
demoteMu sync.Mutex
demote map[string]demoteCandidate
dropped atomic.Uint64
}
var globalAccessTracker = newAccessTracker()
func newAccessTracker() *accessTracker {
return &accessTracker{
ch: make(chan string, accessQueueSize),
demote: make(map[string]demoteCandidate),
}
}
// accessKey is the tracker's map key. Bucket names cannot contain '/', so the
// join is unambiguous.
func accessKey(bucket, object string) string {
return bucket + "/" + object
}
func splitAccessKey(key string) (bucket, object string, ok bool) {
bucket, object, ok = strings.Cut(key, "/")
if !ok || bucket == "" || object == "" {
return "", "", false
}
return bucket, object, true
}
// note records one read. It is called from the GET path and must never block
// or allocate meaningfully: on a full queue the sample is dropped.
func (t *accessTracker) note(bucket, object string) {
if t == nil || !t.enabled.Load() {
return
}
select {
case t.ch <- accessKey(bucket, object):
default:
t.dropped.Add(1)
}
}
// noteDemoteCandidate is called by the scanner for an object it found on a
// fast pool that has gone quiet. The leader picks these up on the next merge.
func (t *accessTracker) noteDemoteCandidate(bucket, object string, pool int) {
if t == nil || !t.enabled.Load() {
return
}
t.demoteMu.Lock()
defer t.demoteMu.Unlock()
if len(t.demote) >= accessMaxDemoteCandidates {
return
}
t.demote[accessKey(bucket, object)] = demoteCandidate{Bucket: bucket, Object: object, Pool: pool}
}
// hits reports cluster-wide accesses to an object over window.
func (t *accessTracker) hits(bucket, object string, window time.Duration) uint64 {
if t == nil {
return 0
}
return t.merged.Load().hits(accessKey(bucket, object), window)
}
// lastAccess reports the cluster-wide time an object was last read. A zero
// time means "no read on record", which for demotion purposes is idle.
func (t *accessTracker) lastAccess(bucket, object string) time.Time {
if t == nil {
return time.Time{}
}
sec := t.merged.Load().lastAccess(accessKey(bucket, object))
if sec == 0 {
return time.Time{}
}
return time.Unix(sec, 0)
}
// snapshot returns the current merged view, or nil if none has been published.
func (t *accessTracker) snapshot() *mergedAccess {
if t == nil {
return nil
}
return t.merged.Load()
}
// takeDemoteCandidates drains and returns the pending candidates.
func (t *accessTracker) takeDemoteCandidates() []demoteCandidate {
t.demoteMu.Lock()
defer t.demoteMu.Unlock()
if len(t.demote) == 0 {
return nil
}
out := make([]demoteCandidate, 0, len(t.demote))
for _, c := range t.demote {
out = append(out, c)
}
t.demote = make(map[string]demoteCandidate)
return out
}
// restoreDemoteCandidates puts candidates back when the publisher is still
// busy. Dropping access samples is acceptable; dropping the only scanner
// discovery of an idle promoted object would delay demotion by a full scan.
func (t *accessTracker) restoreDemoteCandidates(candidates []demoteCandidate) {
if len(candidates) == 0 {
return
}
t.demoteMu.Lock()
defer t.demoteMu.Unlock()
for _, c := range candidates {
if len(t.demote) >= accessMaxDemoteCandidates {
return
}
t.demote[accessKey(c.Bucket, c.Object)] = c
}
}
// shardName is this node's counter object. The node name is hashed so that
// host:port never has to be escaped into an object key.
func (t *accessTracker) shardName() string {
return fmt.Sprintf("%s/%016x.bin", accessTrackerPrefix, xxh3.HashString(globalLocalNodeName))
}
// run owns the live counter map. It drains reads, and on every flush interval
// hands a marshaled shard to a background publisher.
//
// Everything here is best effort: this is accounting for a background data
// movement decision, not a durability path.
func (t *accessTracker) run(ctx context.Context, objAPI ObjectLayer) {
cfg := globalILMConfig.accessCfg()
t.enabled.Store(cfg.AccessTiering)
live := make(map[string]accessEntry)
if cfg.AccessTiering {
if buf, err := readConfig(ctx, objAPI, t.shardName()); err == nil {
var previous accessShard
if _, err = previous.UnmarshalMsg(buf); err == nil && previous.BinWidth == int64(cfg.AccessBinWidth/time.Second) {
now := time.Now().Unix()
for key, entry := range previous.Entries {
entry.rollTo(now, previous.BinWidth, cfg.AccessBins)
if entry.total() != 0 {
live[key] = entry
}
}
}
}
}
ticker := time.NewTicker(cfg.AccessFlush)
defer ticker.Stop()
// One publisher goroutine keeps object-layer I/O off the drain loop.
type pub struct {
shard accessShard
cfg ilm.Config
}
pubCh := make(chan pub, 1)
go func() {
for p := range pubCh {
t.publish(ctx, objAPI, p.shard, p.cfg)
}
}()
defer close(pubCh)
for {
select {
case <-ctx.Done():
return
case key := <-t.ch:
now := time.Now().Unix()
e := live[key]
e.rollTo(now, int64(cfg.AccessBinWidth/time.Second), cfg.AccessBins)
if len(e.Bins) > 0 {
if e.Bins[0] != ^uint32(0) {
e.Bins[0]++
}
}
e.LastAt = now
live[key] = e
case <-ticker.C:
newCfg := globalILMConfig.accessCfg()
t.enabled.Store(newCfg.AccessTiering)
if newCfg.AccessFlush != cfg.AccessFlush && newCfg.AccessFlush > 0 {
ticker.Reset(newCfg.AccessFlush)
}
cfg = newCfg
if !cfg.AccessTiering {
// Feature turned off: release the counters rather than
// holding a stale working set for the process lifetime.
clear(live)
t.merged.Store(nil)
continue
}
now := time.Now().Unix()
binWidth := int64(cfg.AccessBinWidth / time.Second)
t.evict(live, now, binWidth, cfg)
shard := accessShard{
UpdatedAt: now,
BinWidth: binWidth,
Entries: make(map[string]accessEntry, len(live)),
Demote: t.takeDemoteCandidates(),
}
for k, e := range live {
e.Bins = slices.Clone(e.Bins)
shard.Entries[k] = e
}
select {
case pubCh <- pub{shard: shard, cfg: cfg}:
default:
// Previous publish still running; skip this round
// rather than queueing work we cannot keep up with.
t.restoreDemoteCandidates(shard.Demote)
}
}
}
}
// evict rolls every counter forward and drops the ones with no hits left in
// the retained history, then enforces the tracked-object cap.
//
// ponytail: amortized O(n) sweep on the flush tick; a heap would only pay off
// past ~10M tracked keys.
func (t *accessTracker) evict(live map[string]accessEntry, now, binWidth int64, cfg ilm.Config) {
for k, e := range live {
e.rollTo(now, binWidth, cfg.AccessBins)
if e.total() == 0 {
delete(live, k)
continue
}
live[k] = e
}
if len(live) <= cfg.AccessMaxTracked {
return
}
// Over the cap: keep the hottest and drop the rest. Objects that fall
// out are cold by construction, which is exactly what the demotion path
// already assumes about anything missing from the map.
type kt struct {
key string
total uint64
}
all := make([]kt, 0, len(live))
for k, e := range live {
all = append(all, kt{k, e.total()})
}
slices.SortFunc(all, func(a, b kt) int { return cmp.Compare(b.total, a.total) })
for _, x := range all[cfg.AccessMaxTracked:] {
delete(live, x.key)
}
}
// publish writes this node's shard and rebuilds the merged snapshot from every
// node's shard. Both halves are best effort.
func (t *accessTracker) publish(ctx context.Context, objAPI ObjectLayer, shard accessShard, cfg ilm.Config) {
buf, err := shard.MarshalMsg(nil)
if err != nil {
ilmLogIf(ctx, err)
return
}
if err := saveConfig(ctx, objAPI, t.shardName(), buf); err != nil {
ilmLogIf(ctx, err)
// Still rebuild the snapshot below: our own counters are already
// in hand and a stale peer view beats no view.
}
t.merged.Store(t.mergeShards(ctx, objAPI, shard, cfg))
}
// mergeShards sums every live node's counters, including our own in-memory
// shard so this node's most recent reads are never a flush behind.
func (t *accessTracker) mergeShards(ctx context.Context, objAPI ObjectLayer, own accessShard, cfg ilm.Config) *mergedAccess {
now := time.Now().Unix()
binWidth := int64(cfg.AccessBinWidth / time.Second)
out := &mergedAccess{
entries: make(map[string]accessEntry, len(own.Entries)),
binWidth: binWidth,
at: now,
}
add := func(s accessShard) {
if s.BinWidth != binWidth {
// A peer has not yet picked up a bin-width change; merging
// its bins would silently mis-scale the counts.
return
}
for k, e := range s.Entries {
e.rollTo(now, binWidth, cfg.AccessBins)
cur, ok := out.entries[k]
if !ok {
cur = accessEntry{Bins: make([]uint32, cfg.AccessBins)}
}
cur.mergeFrom(e)
out.entries[k] = cur
}
}
add(own)
ownName := t.shardName()
stale := int64(cfg.AccessFlush/time.Second) * accessShardStaleFactor
for _, name := range t.listShards(ctx, objAPI) {
if name == ownName {
continue
}
buf, err := readConfig(ctx, objAPI, name)
if err != nil {
continue
}
var s accessShard
if _, err := s.UnmarshalMsg(buf); err != nil {
continue
}
if !accessShardFresh(now, s, stale) {
continue // node is gone or wedged
}
add(s)
}
return out
}
func (t *accessTracker) listShards(ctx context.Context, objAPI ObjectLayer) []string {
res, err := objAPI.ListObjects(ctx, minioMetaBucket, accessTrackerPrefix+"/", "", "", maxObjectList)
if err != nil {
return nil
}
names := make([]string, 0, len(res.Objects))
for _, o := range res.Objects {
if strings.HasSuffix(o.Name, ".bin") {
names = append(names, o.Name)
}
}
return names
}
// collectDemoteCandidates returns every node's pending demote candidates. Only
// the leader calls this, right before running a demotion pass.
//
// Our own shard is read back rather than skipped: run() drains the local map
// into the published shard, so anything found since the last leader pass lives
// there, not in memory. The local map is still drained here to pick up
// candidates recorded since that publish.
func (t *accessTracker) collectDemoteCandidates(ctx context.Context, objAPI ObjectLayer, cfg ilm.Config) []demoteCandidate {
seen := make(map[string]struct{})
var out []demoteCandidate
for _, c := range t.takeDemoteCandidates() {
key := accessKey(c.Bucket, c.Object)
seen[key] = struct{}{}
out = append(out, c)
}
now := time.Now().Unix()
stale := int64(cfg.AccessFlush/time.Second) * accessShardStaleFactor
for _, name := range t.listShards(ctx, objAPI) {
buf, err := readConfig(ctx, objAPI, name)
if err != nil {
continue
}
var s accessShard
if _, err := s.UnmarshalMsg(buf); err != nil {
continue
}
if !accessShardFresh(now, s, stale) {
continue
}
for _, c := range s.Demote {
key := accessKey(c.Bucket, c.Object)
if _, dup := seen[key]; dup {
continue
}
seen[key] = struct{}{}
out = append(out, c)
}
}
return out
}
+742
View File
@@ -0,0 +1,742 @@
// Code generated by github.com/tinylib/msgp DO NOT EDIT.
package cmd
import (
"github.com/tinylib/msgp/msgp"
)
// DecodeMsg implements msgp.Decodable
func (z *accessEntry) DecodeMsg(dc *msgp.Reader) (err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "b":
var zb0002 uint32
zb0002, err = dc.ReadArrayHeader()
if err != nil {
err = msgp.WrapError(err, "Bins")
return
}
if cap(z.Bins) >= int(zb0002) {
z.Bins = (z.Bins)[:zb0002]
} else {
z.Bins = make([]uint32, zb0002)
}
for za0001 := range z.Bins {
z.Bins[za0001], err = dc.ReadUint32()
if err != nil {
err = msgp.WrapError(err, "Bins", za0001)
return
}
}
case "h":
z.HeadAt, err = dc.ReadInt64()
if err != nil {
err = msgp.WrapError(err, "HeadAt")
return
}
case "l":
z.LastAt, err = dc.ReadInt64()
if err != nil {
err = msgp.WrapError(err, "LastAt")
return
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
return
}
// EncodeMsg implements msgp.Encodable
func (z *accessEntry) EncodeMsg(en *msgp.Writer) (err error) {
// map header, size 3
// write "b"
err = en.Append(0x83, 0xa1, 0x62)
if err != nil {
return
}
err = en.WriteArrayHeader(uint32(len(z.Bins)))
if err != nil {
err = msgp.WrapError(err, "Bins")
return
}
for za0001 := range z.Bins {
err = en.WriteUint32(z.Bins[za0001])
if err != nil {
err = msgp.WrapError(err, "Bins", za0001)
return
}
}
// write "h"
err = en.Append(0xa1, 0x68)
if err != nil {
return
}
err = en.WriteInt64(z.HeadAt)
if err != nil {
err = msgp.WrapError(err, "HeadAt")
return
}
// write "l"
err = en.Append(0xa1, 0x6c)
if err != nil {
return
}
err = en.WriteInt64(z.LastAt)
if err != nil {
err = msgp.WrapError(err, "LastAt")
return
}
return
}
// MarshalMsg implements msgp.Marshaler
func (z *accessEntry) MarshalMsg(b []byte) (o []byte, err error) {
o = msgp.Require(b, z.Msgsize())
// map header, size 3
// string "b"
o = append(o, 0x83, 0xa1, 0x62)
o = msgp.AppendArrayHeader(o, uint32(len(z.Bins)))
for za0001 := range z.Bins {
o = msgp.AppendUint32(o, z.Bins[za0001])
}
// string "h"
o = append(o, 0xa1, 0x68)
o = msgp.AppendInt64(o, z.HeadAt)
// string "l"
o = append(o, 0xa1, 0x6c)
o = msgp.AppendInt64(o, z.LastAt)
return
}
// UnmarshalMsg implements msgp.Unmarshaler
func (z *accessEntry) UnmarshalMsg(bts []byte) (o []byte, err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "b":
var zb0002 uint32
zb0002, bts, err = msgp.ReadArrayHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Bins")
return
}
if cap(z.Bins) >= int(zb0002) {
z.Bins = (z.Bins)[:zb0002]
} else {
z.Bins = make([]uint32, zb0002)
}
for za0001 := range z.Bins {
z.Bins[za0001], bts, err = msgp.ReadUint32Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "Bins", za0001)
return
}
}
case "h":
z.HeadAt, bts, err = msgp.ReadInt64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "HeadAt")
return
}
case "l":
z.LastAt, bts, err = msgp.ReadInt64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "LastAt")
return
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
o = bts
return
}
// Msgsize returns an upper bound estimate of the number of bytes occupied by the serialized message
func (z *accessEntry) Msgsize() (s int) {
s = 1 + 2 + msgp.ArrayHeaderSize + (len(z.Bins) * (msgp.Uint32Size)) + 2 + msgp.Int64Size + 2 + msgp.Int64Size
return
}
// DecodeMsg implements msgp.Decodable
func (z *accessShard) DecodeMsg(dc *msgp.Reader) (err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "u":
z.UpdatedAt, err = dc.ReadInt64()
if err != nil {
err = msgp.WrapError(err, "UpdatedAt")
return
}
case "bw":
z.BinWidth, err = dc.ReadInt64()
if err != nil {
err = msgp.WrapError(err, "BinWidth")
return
}
case "e":
var zb0002 uint32
zb0002, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err, "Entries")
return
}
if z.Entries == nil {
z.Entries = make(map[string]accessEntry, zb0002)
} else if len(z.Entries) > 0 {
clear(z.Entries)
}
for zb0002 > 0 {
zb0002--
var za0001 string
za0001, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "Entries")
return
}
var za0002 accessEntry
err = za0002.DecodeMsg(dc)
if err != nil {
err = msgp.WrapError(err, "Entries", za0001)
return
}
z.Entries[za0001] = za0002
}
case "d":
var zb0003 uint32
zb0003, err = dc.ReadArrayHeader()
if err != nil {
err = msgp.WrapError(err, "Demote")
return
}
if cap(z.Demote) >= int(zb0003) {
z.Demote = (z.Demote)[:zb0003]
} else {
z.Demote = make([]demoteCandidate, zb0003)
}
for za0003 := range z.Demote {
var zb0004 uint32
zb0004, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err, "Demote", za0003)
return
}
for zb0004 > 0 {
zb0004--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err, "Demote", za0003)
return
}
switch msgp.UnsafeString(field) {
case "b":
z.Demote[za0003].Bucket, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Bucket")
return
}
case "o":
z.Demote[za0003].Object, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Object")
return
}
case "p":
z.Demote[za0003].Pool, err = dc.ReadInt()
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Pool")
return
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err, "Demote", za0003)
return
}
}
}
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
return
}
// EncodeMsg implements msgp.Encodable
func (z *accessShard) EncodeMsg(en *msgp.Writer) (err error) {
// map header, size 4
// write "u"
err = en.Append(0x84, 0xa1, 0x75)
if err != nil {
return
}
err = en.WriteInt64(z.UpdatedAt)
if err != nil {
err = msgp.WrapError(err, "UpdatedAt")
return
}
// write "bw"
err = en.Append(0xa2, 0x62, 0x77)
if err != nil {
return
}
err = en.WriteInt64(z.BinWidth)
if err != nil {
err = msgp.WrapError(err, "BinWidth")
return
}
// write "e"
err = en.Append(0xa1, 0x65)
if err != nil {
return
}
err = en.WriteMapHeader(uint32(len(z.Entries)))
if err != nil {
err = msgp.WrapError(err, "Entries")
return
}
for za0001, za0002 := range z.Entries {
err = en.WriteString(za0001)
if err != nil {
err = msgp.WrapError(err, "Entries")
return
}
err = za0002.EncodeMsg(en)
if err != nil {
err = msgp.WrapError(err, "Entries", za0001)
return
}
}
// write "d"
err = en.Append(0xa1, 0x64)
if err != nil {
return
}
err = en.WriteArrayHeader(uint32(len(z.Demote)))
if err != nil {
err = msgp.WrapError(err, "Demote")
return
}
for za0003 := range z.Demote {
// map header, size 3
// write "b"
err = en.Append(0x83, 0xa1, 0x62)
if err != nil {
return
}
err = en.WriteString(z.Demote[za0003].Bucket)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Bucket")
return
}
// write "o"
err = en.Append(0xa1, 0x6f)
if err != nil {
return
}
err = en.WriteString(z.Demote[za0003].Object)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Object")
return
}
// write "p"
err = en.Append(0xa1, 0x70)
if err != nil {
return
}
err = en.WriteInt(z.Demote[za0003].Pool)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Pool")
return
}
}
return
}
// MarshalMsg implements msgp.Marshaler
func (z *accessShard) MarshalMsg(b []byte) (o []byte, err error) {
o = msgp.Require(b, z.Msgsize())
// map header, size 4
// string "u"
o = append(o, 0x84, 0xa1, 0x75)
o = msgp.AppendInt64(o, z.UpdatedAt)
// string "bw"
o = append(o, 0xa2, 0x62, 0x77)
o = msgp.AppendInt64(o, z.BinWidth)
// string "e"
o = append(o, 0xa1, 0x65)
o = msgp.AppendMapHeader(o, uint32(len(z.Entries)))
for za0001, za0002 := range z.Entries {
o = msgp.AppendString(o, za0001)
o, err = za0002.MarshalMsg(o)
if err != nil {
err = msgp.WrapError(err, "Entries", za0001)
return
}
}
// string "d"
o = append(o, 0xa1, 0x64)
o = msgp.AppendArrayHeader(o, uint32(len(z.Demote)))
for za0003 := range z.Demote {
// map header, size 3
// string "b"
o = append(o, 0x83, 0xa1, 0x62)
o = msgp.AppendString(o, z.Demote[za0003].Bucket)
// string "o"
o = append(o, 0xa1, 0x6f)
o = msgp.AppendString(o, z.Demote[za0003].Object)
// string "p"
o = append(o, 0xa1, 0x70)
o = msgp.AppendInt(o, z.Demote[za0003].Pool)
}
return
}
// UnmarshalMsg implements msgp.Unmarshaler
func (z *accessShard) UnmarshalMsg(bts []byte) (o []byte, err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "u":
z.UpdatedAt, bts, err = msgp.ReadInt64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "UpdatedAt")
return
}
case "bw":
z.BinWidth, bts, err = msgp.ReadInt64Bytes(bts)
if err != nil {
err = msgp.WrapError(err, "BinWidth")
return
}
case "e":
var zb0002 uint32
zb0002, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Entries")
return
}
if z.Entries == nil {
z.Entries = make(map[string]accessEntry, zb0002)
} else if len(z.Entries) > 0 {
clear(z.Entries)
}
for zb0002 > 0 {
var za0002 accessEntry
zb0002--
var za0001 string
za0001, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Entries")
return
}
bts, err = za0002.UnmarshalMsg(bts)
if err != nil {
err = msgp.WrapError(err, "Entries", za0001)
return
}
z.Entries[za0001] = za0002
}
case "d":
var zb0003 uint32
zb0003, bts, err = msgp.ReadArrayHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Demote")
return
}
if cap(z.Demote) >= int(zb0003) {
z.Demote = (z.Demote)[:zb0003]
} else {
z.Demote = make([]demoteCandidate, zb0003)
}
for za0003 := range z.Demote {
var zb0004 uint32
zb0004, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003)
return
}
for zb0004 > 0 {
zb0004--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003)
return
}
switch msgp.UnsafeString(field) {
case "b":
z.Demote[za0003].Bucket, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Bucket")
return
}
case "o":
z.Demote[za0003].Object, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Object")
return
}
case "p":
z.Demote[za0003].Pool, bts, err = msgp.ReadIntBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003, "Pool")
return
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err, "Demote", za0003)
return
}
}
}
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
o = bts
return
}
// Msgsize returns an upper bound estimate of the number of bytes occupied by the serialized message
func (z *accessShard) Msgsize() (s int) {
s = 1 + 2 + msgp.Int64Size + 3 + msgp.Int64Size + 2 + msgp.MapHeaderSize
if z.Entries != nil {
for za0001, za0002 := range z.Entries {
_ = za0002
s += msgp.StringPrefixSize + len(za0001) + za0002.Msgsize()
}
}
s += 2 + msgp.ArrayHeaderSize
for za0003 := range z.Demote {
s += 1 + 2 + msgp.StringPrefixSize + len(z.Demote[za0003].Bucket) + 2 + msgp.StringPrefixSize + len(z.Demote[za0003].Object) + 2 + msgp.IntSize
}
return
}
// DecodeMsg implements msgp.Decodable
func (z *demoteCandidate) DecodeMsg(dc *msgp.Reader) (err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, err = dc.ReadMapHeader()
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, err = dc.ReadMapKeyPtr()
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "b":
z.Bucket, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "Bucket")
return
}
case "o":
z.Object, err = dc.ReadString()
if err != nil {
err = msgp.WrapError(err, "Object")
return
}
case "p":
z.Pool, err = dc.ReadInt()
if err != nil {
err = msgp.WrapError(err, "Pool")
return
}
default:
err = dc.Skip()
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
return
}
// EncodeMsg implements msgp.Encodable
func (z demoteCandidate) EncodeMsg(en *msgp.Writer) (err error) {
// map header, size 3
// write "b"
err = en.Append(0x83, 0xa1, 0x62)
if err != nil {
return
}
err = en.WriteString(z.Bucket)
if err != nil {
err = msgp.WrapError(err, "Bucket")
return
}
// write "o"
err = en.Append(0xa1, 0x6f)
if err != nil {
return
}
err = en.WriteString(z.Object)
if err != nil {
err = msgp.WrapError(err, "Object")
return
}
// write "p"
err = en.Append(0xa1, 0x70)
if err != nil {
return
}
err = en.WriteInt(z.Pool)
if err != nil {
err = msgp.WrapError(err, "Pool")
return
}
return
}
// MarshalMsg implements msgp.Marshaler
func (z demoteCandidate) MarshalMsg(b []byte) (o []byte, err error) {
o = msgp.Require(b, z.Msgsize())
// map header, size 3
// string "b"
o = append(o, 0x83, 0xa1, 0x62)
o = msgp.AppendString(o, z.Bucket)
// string "o"
o = append(o, 0xa1, 0x6f)
o = msgp.AppendString(o, z.Object)
// string "p"
o = append(o, 0xa1, 0x70)
o = msgp.AppendInt(o, z.Pool)
return
}
// UnmarshalMsg implements msgp.Unmarshaler
func (z *demoteCandidate) UnmarshalMsg(bts []byte) (o []byte, err error) {
var field []byte
_ = field
var zb0001 uint32
zb0001, bts, err = msgp.ReadMapHeaderBytes(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
for zb0001 > 0 {
zb0001--
field, bts, err = msgp.ReadMapKeyZC(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
switch msgp.UnsafeString(field) {
case "b":
z.Bucket, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Bucket")
return
}
case "o":
z.Object, bts, err = msgp.ReadStringBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Object")
return
}
case "p":
z.Pool, bts, err = msgp.ReadIntBytes(bts)
if err != nil {
err = msgp.WrapError(err, "Pool")
return
}
default:
bts, err = msgp.Skip(bts)
if err != nil {
err = msgp.WrapError(err)
return
}
}
}
o = bts
return
}
// Msgsize returns an upper bound estimate of the number of bytes occupied by the serialized message
func (z demoteCandidate) Msgsize() (s int) {
s = 1 + 2 + msgp.StringPrefixSize + len(z.Bucket) + 2 + msgp.StringPrefixSize + len(z.Object) + 2 + msgp.IntSize
return
}
+349
View File
@@ -0,0 +1,349 @@
// Code generated by github.com/tinylib/msgp DO NOT EDIT.
package cmd
import (
"bytes"
"testing"
"github.com/tinylib/msgp/msgp"
)
func TestMarshalUnmarshalaccessEntry(t *testing.T) {
v := accessEntry{}
bts, err := v.MarshalMsg(nil)
if err != nil {
t.Fatal(err)
}
left, err := v.UnmarshalMsg(bts)
if err != nil {
t.Fatal(err)
}
if len(left) > 0 {
t.Errorf("%d bytes left over after UnmarshalMsg(): %q", len(left), left)
}
left, err = msgp.Skip(bts)
if err != nil {
t.Fatal(err)
}
if len(left) > 0 {
t.Errorf("%d bytes left over after Skip(): %q", len(left), left)
}
}
func BenchmarkMarshalMsgaccessEntry(b *testing.B) {
v := accessEntry{}
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
v.MarshalMsg(nil)
}
}
func BenchmarkAppendMsgaccessEntry(b *testing.B) {
v := accessEntry{}
bts := make([]byte, 0, v.Msgsize())
bts, _ = v.MarshalMsg(bts[0:0])
b.SetBytes(int64(len(bts)))
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
bts, _ = v.MarshalMsg(bts[0:0])
}
}
func BenchmarkUnmarshalaccessEntry(b *testing.B) {
v := accessEntry{}
bts, _ := v.MarshalMsg(nil)
b.ReportAllocs()
b.SetBytes(int64(len(bts)))
b.ResetTimer()
for i := 0; i < b.N; i++ {
_, err := v.UnmarshalMsg(bts)
if err != nil {
b.Fatal(err)
}
}
}
func TestEncodeDecodeaccessEntry(t *testing.T) {
v := accessEntry{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
m := v.Msgsize()
if buf.Len() > m {
t.Log("WARNING: TestEncodeDecodeaccessEntry Msgsize() is inaccurate")
}
vn := accessEntry{}
err := msgp.Decode(&buf, &vn)
if err != nil {
t.Error(err)
}
buf.Reset()
msgp.Encode(&buf, &v)
err = msgp.NewReader(&buf).Skip()
if err != nil {
t.Error(err)
}
}
func BenchmarkEncodeaccessEntry(b *testing.B) {
v := accessEntry{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
b.SetBytes(int64(buf.Len()))
en := msgp.NewWriter(msgp.Nowhere)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
v.EncodeMsg(en)
}
en.Flush()
}
func BenchmarkDecodeaccessEntry(b *testing.B) {
v := accessEntry{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
b.SetBytes(int64(buf.Len()))
rd := msgp.NewEndlessReader(buf.Bytes(), b)
dc := msgp.NewReader(rd)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
err := v.DecodeMsg(dc)
if err != nil {
b.Fatal(err)
}
}
}
func TestMarshalUnmarshalaccessShard(t *testing.T) {
v := accessShard{}
bts, err := v.MarshalMsg(nil)
if err != nil {
t.Fatal(err)
}
left, err := v.UnmarshalMsg(bts)
if err != nil {
t.Fatal(err)
}
if len(left) > 0 {
t.Errorf("%d bytes left over after UnmarshalMsg(): %q", len(left), left)
}
left, err = msgp.Skip(bts)
if err != nil {
t.Fatal(err)
}
if len(left) > 0 {
t.Errorf("%d bytes left over after Skip(): %q", len(left), left)
}
}
func BenchmarkMarshalMsgaccessShard(b *testing.B) {
v := accessShard{}
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
v.MarshalMsg(nil)
}
}
func BenchmarkAppendMsgaccessShard(b *testing.B) {
v := accessShard{}
bts := make([]byte, 0, v.Msgsize())
bts, _ = v.MarshalMsg(bts[0:0])
b.SetBytes(int64(len(bts)))
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
bts, _ = v.MarshalMsg(bts[0:0])
}
}
func BenchmarkUnmarshalaccessShard(b *testing.B) {
v := accessShard{}
bts, _ := v.MarshalMsg(nil)
b.ReportAllocs()
b.SetBytes(int64(len(bts)))
b.ResetTimer()
for i := 0; i < b.N; i++ {
_, err := v.UnmarshalMsg(bts)
if err != nil {
b.Fatal(err)
}
}
}
func TestEncodeDecodeaccessShard(t *testing.T) {
v := accessShard{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
m := v.Msgsize()
if buf.Len() > m {
t.Log("WARNING: TestEncodeDecodeaccessShard Msgsize() is inaccurate")
}
vn := accessShard{}
err := msgp.Decode(&buf, &vn)
if err != nil {
t.Error(err)
}
buf.Reset()
msgp.Encode(&buf, &v)
err = msgp.NewReader(&buf).Skip()
if err != nil {
t.Error(err)
}
}
func BenchmarkEncodeaccessShard(b *testing.B) {
v := accessShard{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
b.SetBytes(int64(buf.Len()))
en := msgp.NewWriter(msgp.Nowhere)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
v.EncodeMsg(en)
}
en.Flush()
}
func BenchmarkDecodeaccessShard(b *testing.B) {
v := accessShard{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
b.SetBytes(int64(buf.Len()))
rd := msgp.NewEndlessReader(buf.Bytes(), b)
dc := msgp.NewReader(rd)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
err := v.DecodeMsg(dc)
if err != nil {
b.Fatal(err)
}
}
}
func TestMarshalUnmarshaldemoteCandidate(t *testing.T) {
v := demoteCandidate{}
bts, err := v.MarshalMsg(nil)
if err != nil {
t.Fatal(err)
}
left, err := v.UnmarshalMsg(bts)
if err != nil {
t.Fatal(err)
}
if len(left) > 0 {
t.Errorf("%d bytes left over after UnmarshalMsg(): %q", len(left), left)
}
left, err = msgp.Skip(bts)
if err != nil {
t.Fatal(err)
}
if len(left) > 0 {
t.Errorf("%d bytes left over after Skip(): %q", len(left), left)
}
}
func BenchmarkMarshalMsgdemoteCandidate(b *testing.B) {
v := demoteCandidate{}
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
v.MarshalMsg(nil)
}
}
func BenchmarkAppendMsgdemoteCandidate(b *testing.B) {
v := demoteCandidate{}
bts := make([]byte, 0, v.Msgsize())
bts, _ = v.MarshalMsg(bts[0:0])
b.SetBytes(int64(len(bts)))
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
bts, _ = v.MarshalMsg(bts[0:0])
}
}
func BenchmarkUnmarshaldemoteCandidate(b *testing.B) {
v := demoteCandidate{}
bts, _ := v.MarshalMsg(nil)
b.ReportAllocs()
b.SetBytes(int64(len(bts)))
b.ResetTimer()
for i := 0; i < b.N; i++ {
_, err := v.UnmarshalMsg(bts)
if err != nil {
b.Fatal(err)
}
}
}
func TestEncodeDecodedemoteCandidate(t *testing.T) {
v := demoteCandidate{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
m := v.Msgsize()
if buf.Len() > m {
t.Log("WARNING: TestEncodeDecodedemoteCandidate Msgsize() is inaccurate")
}
vn := demoteCandidate{}
err := msgp.Decode(&buf, &vn)
if err != nil {
t.Error(err)
}
buf.Reset()
msgp.Encode(&buf, &v)
err = msgp.NewReader(&buf).Skip()
if err != nil {
t.Error(err)
}
}
func BenchmarkEncodedemoteCandidate(b *testing.B) {
v := demoteCandidate{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
b.SetBytes(int64(buf.Len()))
en := msgp.NewWriter(msgp.Nowhere)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
v.EncodeMsg(en)
}
en.Flush()
}
func BenchmarkDecodedemoteCandidate(b *testing.B) {
v := demoteCandidate{}
var buf bytes.Buffer
msgp.Encode(&buf, &v)
b.SetBytes(int64(buf.Len()))
rd := msgp.NewEndlessReader(buf.Bytes(), b)
dc := msgp.NewReader(rd)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
err := v.DecodeMsg(dc)
if err != nil {
b.Fatal(err)
}
}
}
+95
View File
@@ -0,0 +1,95 @@
// Copyright (c) 2015-2026 MinIO, Inc.
package cmd
import (
"math"
"testing"
"time"
"github.com/minio/minio/internal/config/ilm"
)
func TestAccessEntryRollAndHits(t *testing.T) {
entry := accessEntry{Bins: []uint32{3, 2, 1}, HeadAt: 100, LastAt: 107}
entry.rollTo(120, 10, 3)
want := []uint32{0, 0, 3}
for i := range want {
if entry.Bins[i] != want[i] {
t.Fatalf("bins = %v, want %v", entry.Bins, want)
}
}
if entry.HeadAt != 120 || entry.LastAt != 107 {
t.Fatalf("head/last = %d/%d", entry.HeadAt, entry.LastAt)
}
entry = accessEntry{Bins: []uint32{10, 20, 30}, HeadAt: 120}
if got := entry.hits(20*time.Second, 10); got != 30 {
t.Fatalf("20s hits = %d, want 30", got)
}
if got := entry.hits(21*time.Second, 10); got != 60 {
t.Fatalf("21s hits = %d, want 60", got)
}
entry.rollTo(200, 10, 3)
if got := entry.total(); got != 0 {
t.Fatalf("expired total = %d, want 0", got)
}
}
func TestAccessEntryMergeSaturates(t *testing.T) {
entry := accessEntry{Bins: []uint32{math.MaxUint32 - 1}}
entry.mergeFrom(accessEntry{Bins: []uint32{10}, LastAt: 50})
if entry.Bins[0] != math.MaxUint32 || entry.LastAt != 50 {
t.Fatalf("merged entry = %+v", entry)
}
}
func TestAccessTrackerEvictsColdest(t *testing.T) {
tracker := newAccessTracker()
live := map[string]accessEntry{
"a": {Bins: []uint32{1}, HeadAt: 100},
"b": {Bins: []uint32{5}, HeadAt: 100},
"c": {Bins: []uint32{3}, HeadAt: 100},
}
tracker.evict(live, 100, 10, ilm.Config{AccessBins: 1, AccessMaxTracked: 2})
if len(live) != 2 {
t.Fatalf("len = %d, want 2", len(live))
}
if _, ok := live["a"]; ok {
t.Fatal("coldest entry was retained")
}
}
func TestAccessKeyRoundTrip(t *testing.T) {
key := accessKey("bucket", "a/b/c")
bucket, object, ok := splitAccessKey(key)
if !ok || bucket != "bucket" || object != "a/b/c" {
t.Fatalf("split %q = %q/%q/%v", key, bucket, object, ok)
}
}
func TestAccessShardFresh(t *testing.T) {
const now = int64(1000)
if !accessShardFresh(now, accessShard{UpdatedAt: 950}, 100) {
t.Fatal("fresh shard rejected")
}
if accessShardFresh(now, accessShard{UpdatedAt: 899}, 100) {
t.Fatal("stale shard accepted")
}
if accessShardFresh(now, accessShard{UpdatedAt: 1101}, 100) {
t.Fatal("far-future shard accepted")
}
if accessShardFresh(now, accessShard{}, 100) {
t.Fatal("zero timestamp accepted")
}
}
func TestAccessTrackerRestoresDemoteCandidates(t *testing.T) {
tracker := newAccessTracker()
candidate := demoteCandidate{Bucket: "bucket", Object: "object", Pool: 1}
tracker.restoreDemoteCandidates([]demoteCandidate{candidate})
got := tracker.takeDemoteCandidates()
if len(got) != 1 || got[0] != candidate {
t.Fatalf("restored candidates = %+v, want %+v", got, candidate)
}
}
+30
View File
@@ -19,6 +19,7 @@ package cmd
import (
"sync"
"time"
"github.com/minio/minio/internal/config/ilm"
)
@@ -27,6 +28,16 @@ var globalILMConfig = ilmConfig{
cfg: ilm.Config{
ExpirationWorkers: 100,
TransitionWorkers: 100,
// Access tiering stays off until configured, but the counter
// geometry must be sane from the start: the tracker divides by
// AccessBinWidth before any config is loaded.
AccessPromoteWatermark: 85,
AccessBinWidth: time.Minute,
AccessBins: 12,
AccessFlush: time.Minute,
AccessMinResidency: 24 * time.Hour,
AccessWorkers: 10,
AccessMaxTracked: 1000000,
},
}
@@ -49,6 +60,25 @@ func (c *ilmConfig) getTransitionWorkers() int {
return c.cfg.TransitionWorkers
}
// accessCfg returns a copy of the access tiering settings. Callers take the
// whole struct rather than one getter per field because the promotion and
// demotion paths need a consistent view of several knobs at once.
func (c *ilmConfig) accessCfg() ilm.Config {
c.mu.RLock()
defer c.mu.RUnlock()
return c.cfg
}
// accessTieringEnabled is the cheap gate used on the scanner path.
func (c *ilmConfig) accessTieringEnabled() bool {
c.mu.RLock()
defer c.mu.RUnlock()
_, ok := c.cfg.HotPool()
return ok
}
func (c *ilmConfig) update(cfg ilm.Config) {
c.mu.Lock()
defer c.mu.Unlock()
+3 -2
View File
@@ -19,11 +19,12 @@ func _() {
_ = x[lcEventSrc_s3PutObject-8]
_ = x[lcEventSrc_s3CopyObject-9]
_ = x[lcEventSrc_s3CompleteMultipartUpload-10]
_ = x[lcEventSrc_AccessTier-11]
}
const _lcEventSrc_name = "NoneHealScannerDecomRebals3HeadObjects3GetObjects3ListObjectss3PutObjects3CopyObjects3CompleteMultipartUpload"
const _lcEventSrc_name = "NoneHealScannerDecomRebals3HeadObjects3GetObjects3ListObjectss3PutObjects3CopyObjects3CompleteMultipartUploadAccessTier"
var _lcEventSrc_index = [...]uint8{0, 4, 8, 15, 20, 25, 37, 48, 61, 72, 84, 109}
var _lcEventSrc_index = [...]uint8{0, 4, 8, 15, 20, 25, 37, 48, 61, 72, 84, 109, 119}
func (i lcEventSrc) String() string {
idx := int(i) - 0
+24 -2
View File
@@ -992,6 +992,26 @@ func getClusterReplMRFFailedOperationsMD() MetricDescription {
}
}
func getClusterReplMRFDroppedOperationsMD() MetricDescription {
return MetricDescription{
Namespace: nodeMetricNamespace,
Subsystem: replicationSubsystem,
Name: "mrf_dropped_operations_total",
Help: "Total number of replication MRF entries dropped since server start; entries may refer to the same object",
Type: counterMetric,
}
}
func getClusterReplMRFDroppedBytesMD() MetricDescription {
return MetricDescription{
Namespace: nodeMetricNamespace,
Subsystem: replicationSubsystem,
Name: "mrf_dropped_bytes_total",
Help: "Total known bytes of replication MRF entries dropped since server start; delete entries count as zero bytes",
Type: counterMetric,
}
}
func getClusterRepCredentialErrorsMD(namespace MetricNamespace) MetricDescription {
return MetricDescription{
Namespace: namespace,
@@ -2423,6 +2443,8 @@ func getReplicationNodeMetrics(opts MetricsGroupOpts) *MetricsGroupV2 {
avgTransferRate,
maxTransferRate,
mrfCount,
{Description: getClusterReplMRFDroppedOperationsMD(), Value: float64(qs.MRFStats.TotalDroppedCount)},
{Description: getClusterReplMRFDroppedBytesMD(), Value: float64(qs.MRFStats.TotalDroppedBytes)},
}
}
for ep, health := range globalBucketTargetSys.healthStats() {
@@ -3314,10 +3336,10 @@ func getBucketUsageMetrics(opts MetricsGroupOpts) *MetricsGroupV2 {
VariableLabels: map[string]string{"bucket": bucket},
})
if quota != nil && quota.Quota > 0 {
if quotaSize := getBucketQuotaSize(quota); quotaSize > 0 {
metrics = append(metrics, MetricV2{
Description: getBucketUsageQuotaTotalBytesMD(),
Value: float64(quota.Quota),
Value: float64(quotaSize),
VariableLabels: map[string]string{"bucket": bucket},
})
}
+2 -2
View File
@@ -167,8 +167,8 @@ func loadClusterUsageBucketMetrics(ctx context.Context, m MetricValues, c *metri
m.Set(usageBucketVersionsCount, float64(usage.VersionsCount), "bucket", bucket)
m.Set(usageBucketDeleteMarkersCount, float64(usage.DeleteMarkersCount), "bucket", bucket)
if quota != nil && quota.Quota > 0 {
m.Set(usageBucketQuotaTotalBytes, float64(quota.Quota), "bucket", bucket)
if quotaSize := getBucketQuotaSize(quota); quotaSize > 0 {
m.Set(usageBucketQuotaTotalBytes, float64(quotaSize), "bucket", bucket)
}
for k, v := range usage.ObjectSizesHistogram {
+37
View File
@@ -26,6 +26,17 @@ const (
transitionActiveTasks = "transition_active_tasks"
transitionPendingTasks = "transition_pending_tasks"
transitionMissedImmediateTasks = "transition_missed_immediate_tasks"
accessTierActiveTasks = "access_tier_active_tasks"
accessTierPendingTasks = "access_tier_pending_tasks"
accessTierPromotionsTotal = "access_tier_promotions_total"
accessTierDemotionsTotal = "access_tier_demotions_total"
accessTierBytesMovedTotal = "access_tier_bytes_moved_total"
accessTierFailuresTotal = "access_tier_failures_total"
accessTierSkippedWatermark = "access_tier_skipped_watermark_total"
accessTierSkippedMaxSize = "access_tier_skipped_max_size_total"
accessTierSkippedQuota = "access_tier_skipped_bucket_quota_total"
accessTierHotBytes = "access_tier_hot_bytes"
accessTierSamplesDropped = "access_tier_samples_dropped_total"
versionsScanned = "versions_scanned"
)
@@ -34,6 +45,17 @@ var (
ilmTransitionActiveTasksMD = NewGaugeMD(transitionActiveTasks, "Number of active ILM transition tasks")
ilmTransitionPendingTasksMD = NewGaugeMD(transitionPendingTasks, "Number of pending ILM transition tasks in the queue")
ilmTransitionMissedImmediateTasksMD = NewCounterMD(transitionMissedImmediateTasks, "Number of missed immediate ILM transition tasks")
ilmAccessTierActiveTasksMD = NewGaugeMD(accessTierActiveTasks, "Number of active access-tier pool moves")
ilmAccessTierPendingTasksMD = NewGaugeMD(accessTierPendingTasks, "Number of pending access-tier pool moves")
ilmAccessTierPromotionsTotalMD = NewCounterMD(accessTierPromotionsTotal, "Total objects promoted by access-tier ILM")
ilmAccessTierDemotionsTotalMD = NewCounterMD(accessTierDemotionsTotal, "Total objects demoted by access-tier ILM")
ilmAccessTierBytesMovedTotalMD = NewCounterMD(accessTierBytesMovedTotal, "Total logical bytes moved by access-tier ILM")
ilmAccessTierFailuresTotalMD = NewCounterMD(accessTierFailuresTotal, "Total failed access-tier ILM moves")
ilmAccessTierSkippedWatermarkMD = NewCounterMD(accessTierSkippedWatermark, "Promotions skipped because the hot pool reached its watermark")
ilmAccessTierSkippedMaxSizeMD = NewCounterMD(accessTierSkippedMaxSize, "Promotions skipped because the cluster hot-tier size cap was reached")
ilmAccessTierSkippedQuotaMD = NewCounterMD(accessTierSkippedQuota, "Promotions skipped because the bucket hot-tier quota was reached")
ilmAccessTierHotBytesMD = NewGaugeMD(accessTierHotBytes, "Logical bytes currently accounted to the hot tier", "bucket")
ilmAccessTierSamplesDroppedMD = NewCounterMD(accessTierSamplesDropped, "GET samples dropped because the access tracker queue was full")
ilmVersionsScannedMD = NewCounterMD(versionsScanned, "Total number of object versions checked for ILM actions since server start")
)
@@ -47,6 +69,21 @@ func loadILMMetrics(_ context.Context, m MetricValues, _ *metricsCache) error {
m.Set(transitionPendingTasks, float64(globalTransitionState.PendingTasks()))
m.Set(transitionMissedImmediateTasks, float64(globalTransitionState.MissedImmediateTasks()))
}
if globalAccessTierState != nil {
m.Set(accessTierActiveTasks, float64(globalAccessTierState.ActiveTasks()))
m.Set(accessTierPendingTasks, float64(globalAccessTierState.PendingTasks()))
m.Set(accessTierPromotionsTotal, float64(globalAccessTierState.promotions.Load()))
m.Set(accessTierDemotionsTotal, float64(globalAccessTierState.demotions.Load()))
m.Set(accessTierBytesMovedTotal, float64(globalAccessTierState.bytesMoved.Load()))
m.Set(accessTierFailuresTotal, float64(globalAccessTierState.failures.Load()))
m.Set(accessTierSkippedWatermark, float64(globalAccessTierState.skippedWatermark.Load()))
m.Set(accessTierSkippedMaxSize, float64(globalAccessTierState.skippedMaxSize.Load()))
m.Set(accessTierSkippedQuota, float64(globalAccessTierState.skippedQuota.Load()))
for bucket, bytes := range globalAccessTierState.hotUsageSnapshot() {
m.Set(accessTierHotBytes, float64(bytes), "bucket", bucket)
}
}
m.Set(accessTierSamplesDropped, float64(globalAccessTracker.dropped.Load()))
m.Set(versionsScanned, float64(globalScannerMetrics.lifetime(scannerMetricILM)))
return nil
+8
View File
@@ -35,6 +35,8 @@ const (
replicationMaxQueuedCount = "max_queued_count"
replicationMaxDataTransferRate = "max_data_transfer_rate"
replicationRecentBacklogCount = "recent_backlog_count"
replicationMRFDroppedOperations = "mrf_dropped_operations_total"
replicationMRFDroppedBytes = "mrf_dropped_bytes_total"
)
var (
@@ -64,6 +66,10 @@ var (
"Maximum replication data transfer rate in bytes/sec seen since server start")
replicationRecentBacklogCountMD = NewGaugeMD(replicationRecentBacklogCount,
"Total number of objects seen in replication backlog in the last 5 minutes")
replicationMRFDroppedOperationsMD = NewCounterMD(replicationMRFDroppedOperations,
"Total number of replication MRF entries dropped since server start; entries may refer to the same object")
replicationMRFDroppedBytesMD = NewCounterMD(replicationMRFDroppedBytes,
"Total known bytes of replication MRF entries dropped since server start; delete entries count as zero bytes")
)
// loadClusterReplicationMetrics - `MetricsLoaderFn` for cluster replication metrics
@@ -96,6 +102,8 @@ func loadClusterReplicationMetrics(ctx context.Context, m MetricValues, c *metri
m.Set(replicationMaxDataTransferRate, tots.Peak)
}
m.Set(replicationRecentBacklogCount, float64(qs.MRFStats.LastFailedCount))
m.Set(replicationMRFDroppedOperations, float64(qs.MRFStats.TotalDroppedCount))
m.Set(replicationMRFDroppedBytes, float64(qs.MRFStats.TotalDroppedBytes))
return nil
}
+13
View File
@@ -342,6 +342,8 @@ func newMetricGroups(r *prometheus.Registry) *metricsV3Collection {
replicationMaxQueuedCountMD,
replicationMaxDataTransferRateMD,
replicationRecentBacklogCountMD,
replicationMRFDroppedOperationsMD,
replicationMRFDroppedBytesMD,
},
loadClusterReplicationMetrics,
)
@@ -390,6 +392,17 @@ func newMetricGroups(r *prometheus.Registry) *metricsV3Collection {
ilmTransitionActiveTasksMD,
ilmTransitionPendingTasksMD,
ilmTransitionMissedImmediateTasksMD,
ilmAccessTierActiveTasksMD,
ilmAccessTierPendingTasksMD,
ilmAccessTierPromotionsTotalMD,
ilmAccessTierDemotionsTotalMD,
ilmAccessTierBytesMovedTotalMD,
ilmAccessTierFailuresTotalMD,
ilmAccessTierSkippedWatermarkMD,
ilmAccessTierSkippedMaxSizeMD,
ilmAccessTierSkippedQuotaMD,
ilmAccessTierHotBytesMD,
ilmAccessTierSamplesDroppedMD,
ilmVersionsScannedMD,
},
loadILMMetrics,
+10 -2
View File
@@ -99,9 +99,13 @@ type ObjectOptions struct {
ReplicationSourceTaggingTimestamp time.Time // set if MinIOSourceTaggingTimestamp received
ReplicationSourceLegalholdTimestamp time.Time // set if MinIOSourceObjectLegalholdTimestamp received
ReplicationSourceRetentionTimestamp time.Time // set if MinIOSourceObjectRetentionTimestamp received
ReplicaLockReconcile bool // re-order a trusted replica write against the destination version under the object write lock
DeletePrefix bool // set true to enforce a prefix deletion, only application for DeleteObject API,
DeletePrefixObject bool // set true when object's erasure set is resolvable by object name (using getHashedSetIndex)
// Pools-layer version resolver; the caller holds the shared object lock.
replicaObjectInfo GetObjectInfoFn
Speedtest bool // object call specifically meant for SpeedTest code, set to 'true' when invoked by SpeedtestHandler.
// Use the maximum parity (N/2), used when saving server configuration files
@@ -118,6 +122,10 @@ type ObjectOptions struct {
SkipRebalancing bool
SrcPoolIdx int // set by PutObject/CompleteMultipart operations due to rebalance; used to prevent rebalance src, dst pools to be the same
// DstPoolIdx forces a data-movement write onto a specific server pool.
// It is ignored unless DataMovement is true; a pointer keeps pool zero
// distinguishable from the unset value.
DstPoolIdx *int
DataMovement bool // indicates an going decommisionning or rebalacing
@@ -131,8 +139,8 @@ type ObjectOptions struct {
// when looking up a version by fi.VersionID
InclFreeVersions bool
// SkipFreeVersion skips adding a free version when a tiered version is
// being 'replaced'
// Note: Used only when a tiered object is being expired.
// being replaced. Used when expiring tiered content or retiring a copy
// whose tier reference is still owned by another copy.
SkipFreeVersion bool
MetadataChg bool // is true if it is a metadata update operation.
+33 -1
View File
@@ -184,6 +184,16 @@ func getAndValidateAttributesOpts(ctx context.Context, w http.ResponseWriter, r
return opts, valid
}
// Reject out-of-range page sizes as ListObjectParts does, instead of
// answering an invalid request with an empty parts listing.
if opts.MaxParts < 0 {
apiErr = errorCodes.ToAPIErr(ErrInvalidMaxParts)
argumentName = strings.ToLower(xhttp.AmzMaxParts)
argumentValue = r.Header.Get(xhttp.AmzMaxParts)
valid = false
return opts, valid
}
if opts.MaxParts == 0 {
opts.MaxParts = maxPartsList
}
@@ -196,6 +206,14 @@ func getAndValidateAttributesOpts(ctx context.Context, w http.ResponseWriter, r
return opts, valid
}
if opts.PartNumberMarker < 0 {
apiErr = errorCodes.ToAPIErr(ErrInvalidPartNumberMarker)
argumentName = strings.ToLower(xhttp.AmzPartNumberMarker)
argumentValue = r.Header.Get(xhttp.AmzPartNumberMarker)
valid = false
return opts, valid
}
opts.ObjectAttributes = parseObjectAttributes(r.Header)
if len(opts.ObjectAttributes) < 1 {
apiErr = errorCodes.ToAPIErr(ErrInvalidAttributeName)
@@ -415,7 +433,16 @@ func putOptsFromHeaders(ctx context.Context, hdr http.Header, metadata map[strin
if err != nil {
return ObjectOptions{}, err
}
sseKms, err := encrypt.NewSSEKMS(keyID, context)
// kms.Context implements encoding.TextMarshaler, so handing it to the
// SDK's interface{} parameter would serialize the context as a JSON
// string, which the receiving ParseHTTP rejects; a nil Context is a
// typed nil there and would be sent as "{}". Pass a plain map, or
// nothing when no context was requested.
var sdkContext any
if context != nil {
sdkContext = map[string]string(context)
}
sseKms, err := encrypt.NewSSEKMS(keyID, sdkContext)
if err != nil {
return ObjectOptions{}, err
}
@@ -425,6 +452,11 @@ func putOptsFromHeaders(ctx context.Context, hdr http.Header, metadata map[strin
MTime: mtime,
PreserveETag: etag,
ReplicationRequest: trustedReplication,
// The Object Lock timestamps order replicated retention and legal
// hold updates. Dropping them here would leave every update on an
// SSE-KMS destination unordered.
ReplicationSourceLegalholdTimestamp: lholdtimestmp,
ReplicationSourceRetentionTimestamp: retaintimestmp,
}
return op, nil
}
+108
View File
@@ -18,9 +18,11 @@
package cmd
import (
"encoding/xml"
"net/http"
"net/http/httptest"
"reflect"
"strings"
"testing"
xhttp "github.com/minio/minio/internal/http"
@@ -76,3 +78,109 @@ func TestGetAndValidateAttributesOpts(t *testing.T) {
})
}
}
// TestGetAndValidateAttributesOptsPartsRange asserts that GetObjectAttributes
// range checks the ObjectParts pagination headers the way ListObjectParts
// does: a negative value is rejected with the same API error, an absent or
// zero x-amz-max-parts means the default page size, and in-range values are
// passed through unchanged.
func TestGetAndValidateAttributesOptsPartsRange(t *testing.T) {
globalBucketVersioningSys = &BucketVersioningSys{}
bucket := minioMetaBucket
ctx := t.Context()
testCases := []struct {
name string
maxParts string
marker string
wantValid bool
wantErr APIErrorCode
wantArgument string
wantValue string
wantMaxParts int
wantMarker int
}{
{
name: "defaults",
wantValid: true,
wantMaxParts: maxPartsList,
},
{
name: "zero max-parts means the default page size",
maxParts: "0",
wantValid: true,
wantMaxParts: maxPartsList,
},
{
name: "in range values pass through",
maxParts: "10",
marker: "3",
wantValid: true,
wantMaxParts: 10,
wantMarker: 3,
},
{
name: "negative max-parts is rejected",
maxParts: "-1",
wantValid: false,
wantErr: ErrInvalidMaxParts,
wantArgument: strings.ToLower(xhttp.AmzMaxParts),
wantValue: "-1",
},
{
name: "negative part-number-marker is rejected",
marker: "-1",
wantValid: false,
wantErr: ErrInvalidPartNumberMarker,
wantArgument: strings.ToLower(xhttp.AmzPartNumberMarker),
wantValue: "-1",
},
}
for _, testCase := range testCases {
t.Run(testCase.name, func(t *testing.T) {
rec := httptest.NewRecorder()
req := httptest.NewRequest(http.MethodGet, "/testbucket/testobject?attributes", nil)
req.Header.Set(xhttp.AmzObjectAttributes, "ObjectParts")
if testCase.maxParts != "" {
req.Header.Set(xhttp.AmzMaxParts, testCase.maxParts)
}
if testCase.marker != "" {
req.Header.Set(xhttp.AmzPartNumberMarker, testCase.marker)
}
opts, valid := getAndValidateAttributesOpts(ctx, rec, req, bucket, "testobject")
if valid != testCase.wantValid {
t.Fatalf("want valid %v, got %v (%s)", testCase.wantValid, valid, rec.Body.String())
}
if testCase.wantValid {
if opts.MaxParts != testCase.wantMaxParts {
t.Errorf("want MaxParts %d, got %d", testCase.wantMaxParts, opts.MaxParts)
}
if opts.PartNumberMarker != testCase.wantMarker {
t.Errorf("want PartNumberMarker %d, got %d", testCase.wantMarker, opts.PartNumberMarker)
}
return
}
wantErr := errorCodes.ToAPIErr(testCase.wantErr)
if rec.Code != wantErr.HTTPStatusCode {
t.Errorf("want HTTP status %d, got %d", wantErr.HTTPStatusCode, rec.Code)
}
var errResp objectAttributesErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &errResp); err != nil {
t.Fatalf("decode error response: %v (%s)", err, rec.Body.String())
}
if errResp.Code != wantErr.Code || errResp.Message != wantErr.Description {
t.Errorf("want error %s/%q, got %s/%q", wantErr.Code, wantErr.Description, errResp.Code, errResp.Message)
}
if errResp.ArgumentName == nil || *errResp.ArgumentName != testCase.wantArgument {
t.Errorf("want ArgumentName %q, got %v", testCase.wantArgument, errResp.ArgumentName)
}
if errResp.ArgumentValue == nil || *errResp.ArgumentValue != testCase.wantValue {
t.Errorf("want ArgumentValue %q, got %v", testCase.wantValue, errResp.ArgumentValue)
}
})
}
}
+29 -6
View File
@@ -49,6 +49,7 @@ import (
xhttp "github.com/minio/minio/internal/http"
xioutil "github.com/minio/minio/internal/ioutil"
"github.com/minio/minio/internal/logger"
"github.com/minio/sio"
"github.com/pgsty/silo-pkg/v3/trie"
"github.com/pgsty/silo-pkg/v3/wildcard"
"github.com/valyala/bytebufferpool"
@@ -610,7 +611,10 @@ func excludeForCompression(header http.Header, object string, cfg compress.Confi
return true
}
if crypto.Requested(header) && !cfg.AllowEncrypted {
// SSE-C replication sends raw ciphertext without compression metadata.
// Exclude new SSE-C data from compression; other modes follow allow_encryption.
if crypto.SSEC.IsRequested(header) ||
(crypto.Requested(header) && !cfg.AllowEncrypted) {
return true
}
@@ -675,19 +679,35 @@ func getPartFile(entriesTrie *trie.Trie, partNumber int, etag string) (partFile
return partFile
}
func partNumberToRangeSpec(oi ObjectInfo, partNumber int) *HTTPRangeSpec {
func partNumberToRangeSpec(oi ObjectInfo, partNumber int) (*HTTPRangeSpec, error) {
if oi.Size == 0 || len(oi.Parts) == 0 {
return nil
return nil, nil
}
// For an encrypted, uncompressed object derive each part's plaintext length
// from the stored ciphertext length instead of trusting ActualSize: parts
// written before this was normalised record the ciphertext length there.
// The range returned here is consumed in the plaintext domain, where
// GetDecryptedRange and DecryptedSize both use exactly this arithmetic.
_, isEncrypted := crypto.IsEncrypted(oi.UserDefined)
deriveFromSize := isEncrypted && !oi.IsCompressed()
var start int64
end := int64(-1)
for i := 0; i < len(oi.Parts) && i < partNumber; i++ {
partSize := oi.Parts[i].ActualSize
if deriveFromSize {
decrypted, err := sio.DecryptedSize(uint64(oi.Parts[i].Size))
if err != nil {
return nil, errObjectTampered
}
partSize = int64(decrypted)
}
start = end + 1
end = start + oi.Parts[i].ActualSize - 1
end = start + partSize - 1
}
return &HTTPRangeSpec{Start: start, End: end}
return &HTTPRangeSpec{Start: start, End: end}, nil
}
// Returns the compressed offset which should be skipped.
@@ -806,7 +826,10 @@ func NewGetObjectReader(rs *HTTPRangeSpec, oi ObjectInfo, opts ObjectOptions, h
}
if rs == nil && opts.PartNumber > 0 {
rs = partNumberToRangeSpec(oi, opts.PartNumber)
rs, err = partNumberToRangeSpec(oi, opts.PartNumber)
if err != nil {
return nil, 0, 0, err
}
}
_, isEncrypted := crypto.IsEncrypted(oi.UserDefined)
@@ -0,0 +1,185 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful,
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"encoding/xml"
"fmt"
"net/http"
"net/http/httptest"
"strconv"
"testing"
"github.com/minio/minio/internal/auth"
xhttp "github.com/minio/minio/internal/http"
)
type attributesPartsPage struct {
ObjectParts struct {
IsTruncated bool
MaxParts int
NextPartNumberMarker int
PartNumberMarker int
PartsCount int
Parts []struct {
PartNumber int
Size int64
} `xml:"Part"`
}
}
// TestAPIGetObjectAttributesPartsPagination asserts that ObjectParts pagination
// terminates for sparse part numbers: truncation is decided by whether parts
// remain, not by comparing the last part number with the part count.
func TestAPIGetObjectAttributesPartsPagination(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIGetObjectAttributesPartsPagination,
})
}
func testAPIGetObjectAttributesPartsPagination(_ ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
signedRequest := func(method, target string, body []byte, headers map[string]string) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, target, int64(len(body)), bytes.NewReader(body),
credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return rec
}
const partSize = 5 * 1024 * 1024
// upload completes a multipart object with the given part numbers. Part
// numbers must be strictly increasing, but gaps are allowed, so the last
// part number need not equal the number of parts.
upload := func(object string, partNumbers []int) {
t.Helper()
initRec := signedRequest(http.MethodPost, getNewMultipartURL("", bucketName, object), nil, nil)
if initRec.Code != http.StatusOK {
t.Fatalf("%s INIT: %d %s", instanceType, initRec.Code, initRec.Body.String())
}
var initiated struct {
UploadID string `xml:"UploadId"`
}
if err := xml.Unmarshal(initRec.Body.Bytes(), &initiated); err != nil {
t.Fatal(err)
}
var complete bytes.Buffer
complete.WriteString("<CompleteMultipartUpload>")
for i, number := range partNumbers {
body := bytes.Repeat([]byte("abcd"), partSize/4)
if i == len(partNumbers)-1 {
body = bytes.Repeat([]byte("12345"), 103)
}
put := signedRequest(http.MethodPut,
getPutObjectPartURL("", bucketName, object, initiated.UploadID, strconv.Itoa(number)), body, nil)
if put.Code != http.StatusOK {
t.Fatalf("%s PART %d: %d %s", instanceType, number, put.Code, put.Body.String())
}
fmt.Fprintf(&complete, "<Part><PartNumber>%d</PartNumber><ETag>%s</ETag></Part>",
number, put.Header()[xhttp.ETag][0])
}
complete.WriteString("</CompleteMultipartUpload>")
finish := signedRequest(http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, initiated.UploadID), complete.Bytes(), nil)
if finish.Code != http.StatusOK {
t.Fatalf("%s COMPLETE: %d %s", instanceType, finish.Code, finish.Body.String())
}
}
attributes := func(object string, marker, maxParts int) attributesPartsPage {
t.Helper()
headers := map[string]string{xhttp.AmzObjectAttributes: "ObjectParts"}
if marker > 0 {
headers[xhttp.AmzPartNumberMarker] = strconv.Itoa(marker)
}
if maxParts > 0 {
headers[xhttp.AmzMaxParts] = strconv.Itoa(maxParts)
}
rec := signedRequest(http.MethodGet, getGetObjectURL("", bucketName, object)+"?attributes", nil, headers)
if rec.Code != http.StatusOK {
t.Fatalf("%s ATTRIBUTES marker=%d max=%d: %d %s", instanceType, marker, maxParts, rec.Code, rec.Body.String())
}
var page attributesPartsPage
if err := xml.Unmarshal(rec.Body.Bytes(), &page); err != nil {
t.Fatal(err)
}
return page
}
check := func(name string, page attributesPartsPage, wantParts []int, wantTruncated bool, wantNext, wantCount int) {
t.Helper()
got := make([]int, 0, len(page.ObjectParts.Parts))
for _, p := range page.ObjectParts.Parts {
got = append(got, p.PartNumber)
}
if fmt.Sprint(got) != fmt.Sprint(wantParts) {
t.Errorf("%s %s: parts=%v want=%v", instanceType, name, got, wantParts)
}
if page.ObjectParts.IsTruncated != wantTruncated {
t.Errorf("%s %s: IsTruncated=%v want=%v", instanceType, name, page.ObjectParts.IsTruncated, wantTruncated)
}
if page.ObjectParts.NextPartNumberMarker != wantNext {
t.Errorf("%s %s: NextPartNumberMarker=%d want=%d", instanceType, name,
page.ObjectParts.NextPartNumberMarker, wantNext)
}
if page.ObjectParts.PartsCount != wantCount {
t.Errorf("%s %s: PartsCount=%d want=%d", instanceType, name, page.ObjectParts.PartsCount, wantCount)
}
}
sparse := "review/attributes-pagination-sparse"
upload(sparse, []int{1, 3, 5})
// A single page holding every sparse part is complete.
check("sparse full page", attributes(sparse, 0, 0), []int{1, 3, 5}, false, 0, 3)
// One part per page walks 1, 3, 5 and stops. Page 2 returns part 3,
// whose number equals the part count, and must still be truncated:
// the old comparison declared that page final and silently dropped
// part 5.
check("sparse max-parts=1 page 1", attributes(sparse, 0, 1), []int{1}, true, 1, 3)
check("sparse max-parts=1 page 2", attributes(sparse, 1, 1), []int{3}, true, 3, 3)
check("sparse max-parts=1 page 3", attributes(sparse, 3, 1), []int{5}, false, 0, 3)
// A page that exactly holds the remainder is not truncated, so no
// empty trailing page is requested.
check("sparse max-parts=2 page 2", attributes(sparse, 1, 2), []int{3, 5}, false, 0, 3)
// A marker at or past the last part ends the listing instead of
// looping back through NextPartNumberMarker=0.
check("sparse marker at last part", attributes(sparse, 5, 0), nil, false, 0, 3)
check("sparse marker past last part", attributes(sparse, 9, 0), nil, false, 0, 3)
// Contiguous numbering is affected too: the empty page past the end
// used to report IsTruncated=true with NextPartNumberMarker=0.
contiguous := "review/attributes-pagination-contiguous"
upload(contiguous, []int{1, 2})
check("contiguous full page", attributes(contiguous, 0, 0), []int{1, 2}, false, 0, 2)
check("contiguous max-parts=1 page 1", attributes(contiguous, 0, 1), []int{1}, true, 1, 2)
check("contiguous max-parts=1 page 2", attributes(contiguous, 1, 1), []int{2}, false, 0, 2)
check("contiguous marker at last part", attributes(contiguous, 2, 0), nil, false, 0, 2)
}
+410
View File
@@ -0,0 +1,410 @@
// Copyright (c) 2015-2026 MinIO, Inc.
// Copyright (c) 2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful,
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"crypto/md5"
"encoding/base64"
"encoding/xml"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"strconv"
"testing"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
xhttp "github.com/minio/minio/internal/http"
)
// attributesPartsResponse is the subset of the GetObjectAttributes response
// the ObjectParts tests assert on.
type attributesPartsResponse struct {
ObjectSize int64
ObjectParts struct {
IsTruncated bool
NextPartNumberMarker int
PartsCount int
Parts []struct {
PartNumber int
Size int64
} `xml:"Part"`
}
}
func attributesPartsSSECHeaders(key []byte) map[string]string {
digest := md5.Sum(key)
return map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(digest[:]),
}
}
func attributesPartsSignedRequest(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
method, target string, body []byte, headers map[string]string,
) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, target, int64(len(body)), bytes.NewReader(body),
credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return rec
}
// attributesPartsUpload completes a multipart upload of bodies under
// partNumbers and returns the concatenated plaintext.
func attributesPartsUpload(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
bucketName, object string, headers map[string]string, bodies [][]byte, partNumbers []int,
) []byte {
t.Helper()
initRec := attributesPartsSignedRequest(t, apiRouter, credentials, http.MethodPost,
getNewMultipartURL("", bucketName, object), nil, headers)
if initRec.Code != http.StatusOK {
t.Fatalf("NewMultipart %s: %d %s", object, initRec.Code, initRec.Body.String())
}
var initiated struct {
UploadID string `xml:"UploadId"`
}
if err := xml.Unmarshal(initRec.Body.Bytes(), &initiated); err != nil {
t.Fatal(err)
}
var complete bytes.Buffer
complete.WriteString("<CompleteMultipartUpload>")
var data []byte
for i, body := range bodies {
data = append(data, body...)
put := attributesPartsSignedRequest(t, apiRouter, credentials, http.MethodPut,
getPutObjectPartURL("", bucketName, object, initiated.UploadID, strconv.Itoa(partNumbers[i])), body, headers)
if put.Code != http.StatusOK {
t.Fatalf("PutObjectPart %s part %d: %d %s", object, partNumbers[i], put.Code, put.Body.String())
}
fmt.Fprintf(&complete, "<Part><PartNumber>%d</PartNumber><ETag>%s</ETag></Part>",
partNumbers[i], put.Header()[xhttp.ETag][0])
}
complete.WriteString("</CompleteMultipartUpload>")
finish := attributesPartsSignedRequest(t, apiRouter, credentials, http.MethodPost,
getCompleteMultipartUploadURL("", bucketName, object, initiated.UploadID), complete.Bytes(), headers)
if finish.Code != http.StatusOK {
t.Fatalf("CompleteMultipartUpload %s: %d %s", object, finish.Code, finish.Body.String())
}
return data
}
// attributesPartsFetch issues GetObjectAttributes for ObjectSize and
// ObjectParts and decodes the response.
func attributesPartsFetch(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
bucketName, object string, headers map[string]string,
) attributesPartsResponse {
t.Helper()
attributeHeaders := maps.Clone(headers)
if attributeHeaders == nil {
attributeHeaders = make(map[string]string)
}
attributeHeaders[xhttp.AmzObjectAttributes] = "ObjectSize,ObjectParts"
rec := attributesPartsSignedRequest(t, apiRouter, credentials, http.MethodGet,
getGetObjectURL("", bucketName, object)+"?attributes", nil, attributeHeaders)
if rec.Code != http.StatusOK {
t.Fatalf("GetObjectAttributes %s: %d %s", object, rec.Code, rec.Body.String())
}
var response attributesPartsResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatalf("decode GetObjectAttributes response: %v (%s)", err, rec.Body.String())
}
return response
}
// enableAttributesPartsCompression turns on compression for ".txt" objects,
// including encrypted ones, for the duration of the test.
func enableAttributesPartsCompression(t *testing.T) {
t.Helper()
globalCompressConfigMu.Lock()
previous := globalCompressConfig
globalCompressConfig.Enabled = true
globalCompressConfig.Extensions = []string{".txt"}
globalCompressConfig.MimeTypes = nil
globalCompressConfig.AllowEncrypted = true
globalCompressConfigMu.Unlock()
t.Cleanup(func() {
globalCompressConfigMu.Lock()
globalCompressConfig = previous
globalCompressConfigMu.Unlock()
})
}
// TestAPIGetObjectAttributesMultipartLogicalPartSize asserts that ObjectPart.Size
// reports the uploaded plaintext length of every part, not the transformed
// length stored on disk. Compressed parts must report the pre-compression
// length and encrypted parts the pre-encryption length, for consecutive as
// well as sparse part numbering.
func TestAPIGetObjectAttributesMultipartLogicalPartSize(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIGetObjectAttributesMultipartLogicalPartSize,
})
}
func testAPIGetObjectAttributesMultipartLogicalPartSize(_ ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
enableAttributesPartsCompression(t)
for _, variant := range []struct {
name string
extension string
encrypted bool
}{
{name: "plain", extension: ".bin"},
{name: "compressed", extension: ".txt"},
{name: "ssec", extension: ".bin", encrypted: true},
{name: "compressed-ssec", extension: ".txt", encrypted: true},
} {
for _, numbering := range []struct {
name string
partNumbers []int
}{
{name: "consecutive", partNumbers: []int{1, 2}},
{name: "sparse", partNumbers: []int{1, 3}},
} {
t.Run(variant.name+"/"+numbering.name, func(t *testing.T) {
var headers map[string]string
if variant.encrypted {
headers = attributesPartsSSECHeaders(bytes.Repeat([]byte{0x19}, 32))
}
object := "attributes/parts-" + variant.name + "-" + numbering.name + variant.extension
bodies := [][]byte{
bytes.Repeat([]byte("abcd"), 5*1024*1024/4),
bytes.Repeat([]byte("12345"), 103),
}
data := attributesPartsUpload(t, apiRouter, credentials, bucketName, object, headers, bodies, numbering.partNumbers)
get := attributesPartsSignedRequest(t, apiRouter, credentials, http.MethodGet,
getGetObjectURL("", bucketName, object), nil, headers)
if get.Code != http.StatusOK || !bytes.Equal(get.Body.Bytes(), data) {
t.Fatalf("%s GET: %d bytes=%d want=%d", instanceType, get.Code, get.Body.Len(), len(data))
}
response := attributesPartsFetch(t, apiRouter, credentials, bucketName, object, headers)
if response.ObjectSize != int64(len(data)) {
t.Errorf("%s ObjectSize=%d want=%d", instanceType, response.ObjectSize, len(data))
}
if len(response.ObjectParts.Parts) != len(bodies) {
t.Fatalf("%s part count=%d want=%d", instanceType, len(response.ObjectParts.Parts), len(bodies))
}
if response.ObjectParts.IsTruncated {
t.Errorf("%s lists all %d parts %v, but IsTruncated=true NextPartNumberMarker=%d",
instanceType, len(bodies), numbering.partNumbers, response.ObjectParts.NextPartNumberMarker)
}
var total int64
for i, part := range response.ObjectParts.Parts {
if part.PartNumber != numbering.partNumbers[i] {
t.Errorf("%s part %d number=%d want=%d", instanceType, i, part.PartNumber, numbering.partNumbers[i])
}
if part.Size != int64(len(bodies[i])) {
t.Errorf("%s part %d size=%d want logical size=%d",
instanceType, part.PartNumber, part.Size, len(bodies[i]))
}
total += part.Size
}
if total != response.ObjectSize {
t.Errorf("%s part sizes sum to %d, ObjectSize=%d", instanceType, total, response.ObjectSize)
}
})
}
}
}
// TestAPIGetObjectAttributesCompressedEmptyTrailingPart pins the reported
// size of a compressed part carrying no payload. Such a part stores
// Size 0 and ActualSize 0, so it exercises the lower bound of the
// ActualSize guard and must keep reporting 0.
func TestAPIGetObjectAttributesCompressedEmptyTrailingPart(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIGetObjectAttributesCompressedEmptyTrailingPart,
})
}
func testAPIGetObjectAttributesCompressedEmptyTrailingPart(_ ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
enableAttributesPartsCompression(t)
object := "attributes/parts-compressed-empty-tail.txt"
bodies := [][]byte{bytes.Repeat([]byte("abcd"), 5*1024*1024/4), {}}
data := attributesPartsUpload(t, apiRouter, credentials, bucketName, object, nil, bodies, []int{1, 2})
response := attributesPartsFetch(t, apiRouter, credentials, bucketName, object, nil)
if response.ObjectSize != int64(len(data)) {
t.Errorf("%s ObjectSize=%d want=%d", instanceType, response.ObjectSize, len(data))
}
if len(response.ObjectParts.Parts) != len(bodies) {
t.Fatalf("%s part count=%d want=%d", instanceType, len(response.ObjectParts.Parts), len(bodies))
}
for i, part := range response.ObjectParts.Parts {
if part.Size != int64(len(bodies[i])) {
t.Errorf("%s part %d size=%d want logical size=%d",
instanceType, part.PartNumber, part.Size, len(bodies[i]))
}
}
}
// TestAPIGetObjectAttributesEncryptedPartLengths pins how a part length that
// cannot be a valid encrypted stream is reported, for both encrypted layouts.
// Parts of an encrypted multipart object are separate streams, so an
// unconvertible one is corrupt and must fail the request. DecryptObjectInfo
// does not catch that: ObjectInfo.isMultipart gives up on the first part that
// fails sio.DecryptedSize, after which ObjectInfo.DecryptedSize validates only
// the object total. A legacy encrypted object carries no multipart marker and
// is one continuous stream that the erasure writer split into storage
// fragments; those fragments are not independently decryptable, so they must
// keep their stored size rather than fail an intact object. Both fixtures use
// part lengths 5245473 and 1, whose sum 5245474 is a valid stream length while
// the second part alone is not.
func TestAPIGetObjectAttributesEncryptedPartLengths(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIGetObjectAttributesEncryptedPartLengths,
})
}
func testAPIGetObjectAttributesEncryptedPartLengths(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
// Part lengths 5245473 and 1: the sum 5245474 is a valid encrypted stream
// length while the second part alone is not. Since pgsty/silo#119,
// PutObjectPart derives an encrypted part's plaintext length from the bytes
// written and rejects one that cannot be a valid stream, so an object with
// these per-part sizes can no longer be created through a normal write; it
// only exists as pre-#119 on-disk state or from an old peer. Inject that
// stored shape directly through the object layer and exercise the handler,
// which is what this test pins.
partLengths := []int64{5245473, 1}
for _, variant := range []struct {
name string
metadata map[string]string
tampered bool
}{
{
// Parts of an encrypted multipart object are separate streams, so an
// unconvertible one is corrupt and must fail the request.
name: "separately-encrypted-parts",
metadata: map[string]string{crypto.MetaMultipart: ""},
tampered: true,
},
{
// A legacy encrypted object carries no multipart marker and is one
// continuous stream split into storage fragments that are not
// independently decryptable, so they keep their stored size.
name: "legacy-single-stream",
metadata: map[string]string{crypto.MetaIV: "legacy"},
},
} {
t.Run(variant.name, func(t *testing.T) {
object := "attributes/parts-encrypted-" + variant.name
var total int64
infoParts := make([]ObjectPartInfo, 0, len(partLengths))
for i, length := range partLengths {
total += length
infoParts = append(infoParts, ObjectPartInfo{Number: i + 1, Size: length, ActualSize: length})
}
info := ObjectInfo{
Bucket: bucketName,
Name: object,
Size: total,
ModTime: UTCNow(),
IsLatest: true,
UserDefined: maps.Clone(variant.metadata),
Parts: infoParts,
}
previous := newObjectLayerFn()
setObjectLayer(&attributesPartsObjectLayer{ObjectLayer: obj, bucket: bucketName, object: object, info: info})
defer setObjectLayer(previous)
rec := attributesPartsSignedRequest(t, apiRouter, credentials, http.MethodGet,
getGetObjectURL("", bucketName, object)+"?attributes", nil,
map[string]string{xhttp.AmzObjectAttributes: "ObjectParts"})
if variant.tampered {
wantErr := errorCodes.ToAPIErr(ErrObjectTampered)
if rec.Code != wantErr.HTTPStatusCode {
t.Fatalf("%s status %d, want %d: %s", instanceType, rec.Code, wantErr.HTTPStatusCode, rec.Body.String())
}
var errResp APIErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &errResp); err != nil {
t.Fatalf("decode error response: %v (%s)", err, rec.Body.String())
}
if errResp.Code != wantErr.Code {
t.Errorf("%s error code %q, want %q", instanceType, errResp.Code, wantErr.Code)
}
return
}
if rec.Code != http.StatusOK {
t.Fatalf("%s status %d, want %d: %s", instanceType, rec.Code, http.StatusOK, rec.Body.String())
}
var response attributesPartsResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatalf("decode GetObjectAttributes response: %v (%s)", err, rec.Body.String())
}
if len(response.ObjectParts.Parts) != len(partLengths) {
t.Fatalf("%s part count=%d want=%d", instanceType, len(response.ObjectParts.Parts), len(partLengths))
}
for i, part := range response.ObjectParts.Parts {
if part.Size != partLengths[i] {
t.Errorf("%s part %d size=%d, want the stored fragment size %d",
instanceType, part.PartNumber, part.Size, partLengths[i])
}
}
})
}
}
// attributesPartsObjectLayer returns a crafted ObjectInfo for one object so a
// stored shape that pgsty/silo#119 no longer lets PutObjectPart create can be
// handed to the GetObjectAttributes handler under test; every other call falls
// through to the real layer.
type attributesPartsObjectLayer struct {
ObjectLayer
bucket, object string
info ObjectInfo
}
func (o *attributesPartsObjectLayer) GetObjectInfo(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
if bucket == o.bucket && object == o.object {
return o.info.Clone(), nil
}
return o.ObjectLayer.GetObjectInfo(ctx, bucket, object, opts)
}
+270 -1
View File
@@ -30,6 +30,7 @@ import (
"testing"
"github.com/minio/minio/internal/auth"
objectreplication "github.com/minio/minio/internal/bucket/replication"
"github.com/minio/minio/internal/hash"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
@@ -61,6 +62,13 @@ func copyChecksumRequest(t *testing.T, apiRouter http.Handler, credentials auth.
t.Fatalf("failed to build CopyObject request: %v", err)
}
req.Header.Set(xhttp.AmzCopySource, SlashSeparator+pathJoin(bucket, source))
// Re-sign so x-amz-copy-source is covered by the signature, as real S3
// clients send it; the verifier rejects unsigned x-amz-* headers.
if credentials.AccessKey != "" && credentials.SecretKey != "" {
if err := signRequestV4(req, credentials.AccessKey, credentials.SecretKey); err != nil {
t.Fatalf("failed to re-sign CopyObject request: %v", err)
}
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return rec
@@ -363,7 +371,10 @@ func testAPICopyObjectServerSideChecksumEncryption(obj ObjectLayer, instanceType
compressed bool
}{
{name: "encrypted-only", extension: ".bin"},
{name: "compressed-encrypted", extension: ".txt", compressed: true},
// SSE-C is excluded from compression whatever allow_encryption says,
// so a compressible destination extension changes nothing here. The
// SSE-S3 sibling above keeps the compressed-encrypted coverage.
{name: "compressible-extension", extension: ".txt"},
} {
t.Run(variant.name, func(t *testing.T) {
destination := "copy-checksum/sse-c-" + variant.name + variant.extension
@@ -592,3 +603,261 @@ func testPutObjectRejectsMissingServerSideChecksum(obj ObjectLayer, instanceType
})
}
}
// copyChecksumSSECHeaders returns the SSE-C headers naming key for a request
// that reads or writes the object itself.
func copyChecksumSSECHeaders(key []byte) map[string]string {
digest := md5.Sum(key)
return map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(digest[:]),
}
}
// copyChecksumSSECCopySource returns the SSE-C headers naming key as the
// CopyObject source key.
func copyChecksumSSECCopySource(key []byte) map[string]string {
digest := md5.Sum(key)
return map[string]string{
xhttp.AmzServerSideEncryptionCopyCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCopyCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCopyCustomerKeyMD5: base64.StdEncoding.EncodeToString(digest[:]),
}
}
// TestAPICopyObjectSSECKeyRotationChecksumAlgorithm covers an in-place SSE-C key
// rotation that also requests a different checksum algorithm. The rotation fast
// path only rewraps the object key and never reads the object data, so it cannot
// honor the request; the copy has to fall through to the re-encrypting path that
// recomputes, stores and reports the requested algorithm.
func TestAPICopyObjectSSECKeyRotationChecksumAlgorithm(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPICopyObjectSSECKeyRotationChecksumAlgorithm,
endpoints: []string{"CopyObject", "PutObject", "GetObject"},
})
}
func testAPICopyObjectSSECKeyRotationChecksumAlgorithm(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
data := bytes.Repeat([]byte("rotate-and-upgrade-the-checksum-"), 32*1024)
object := "copy-checksum/rotate-checksum-algorithm.bin"
oldKey := bytes.Repeat([]byte{0x31}, 32)
newKey := bytes.Repeat([]byte{0x42}, 32)
putHeaders := copyChecksumSSECHeaders(oldKey)
putHeaders[xhttp.AmzChecksumCRC32] = mustChecksum(t, hash.ChecksumCRC32, data)
putCopyChecksumSource(t, apiRouter, credentials, bucketName, object, data, putHeaders)
rotate := copyChecksumSSECHeaders(newKey)
rotate[xhttp.AmzChecksumAlgo] = hash.ChecksumSHA256.String()
for name, value := range copyChecksumSSECCopySource(oldKey) {
rotate[name] = value
}
rec := copyChecksumRequest(t, apiRouter, credentials, bucketName, object, object, rotate)
if rec.Code != http.StatusOK {
t.Fatalf("%s: rotation with a requested algorithm failed: %d %s",
instanceType, rec.Code, rec.Body.String())
}
assertCopyChecksumResponse(t, rec, hash.ChecksumSHA256, data)
decryptHeaders := http.Header{}
for name, value := range copyChecksumSSECHeaders(newKey) {
decryptHeaders.Set(name, value)
}
oi := assertCopyChecksum(t, obj, bucketName, object, hash.ChecksumSHA256, data, false, decryptHeaders)
if stored, _ := oi.decryptChecksums(0, decryptHeaders); stored[hash.ChecksumCRC32.String()] != "" {
t.Fatalf("%s: rotation kept the superseded CRC32 checksum: %v", instanceType, stored)
}
getHeaders := copyChecksumSSECHeaders(newKey)
getHeaders[xhttp.AmzChecksumMode] = "ENABLED"
req, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, getHeaders)
if err != nil {
t.Fatalf("failed to build GetObject request: %v", err)
}
response := httptest.NewRecorder()
apiRouter.ServeHTTP(response, req)
if response.Code != http.StatusOK || !bytes.Equal(response.Body.Bytes(), data) {
t.Fatalf("%s: post-rotation GetObject returned %d with %d bytes, want 200 with %d bytes: %s",
instanceType, response.Code, response.Body.Len(), len(data), response.Body.String())
}
if got, want := response.Header().Get(xhttp.AmzChecksumSHA256),
mustChecksum(t, hash.ChecksumSHA256, data); got != want {
t.Fatalf("%s: post-rotation GetObject SHA256 %q, want %q", instanceType, got, want)
}
if got := response.Header().Get(xhttp.AmzChecksumCRC32); got != "" {
t.Fatalf("%s: post-rotation GetObject still returns the superseded CRC32 %q", instanceType, got)
}
stale, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, copyChecksumSSECHeaders(oldKey))
if err != nil {
t.Fatalf("failed to build stale-key GetObject request: %v", err)
}
staleResponse := httptest.NewRecorder()
apiRouter.ServeHTTP(staleResponse, stale)
if staleResponse.Code != http.StatusForbidden {
t.Fatalf("%s: GetObject with the rotated-out key returned %d, want 403",
instanceType, staleResponse.Code)
}
}
// TestAPICopyObjectSSECKeyRotationKeepsChecksumAbsence pins the compatibility
// limitation accepted with pgsty/silo#113: without a requested algorithm an
// in-place SSE-C rotation preserves the stored checksum state, including its
// absence, so a checksum-less object does not gain the CRC64NVME that every
// re-encrypting CopyObject adds. Deliberate, and the counterpart of the
// preserved CRC32 that TestAPICopyObjectSSECKeyRotationKeepsCompressionState
// pins for a checksum-bearing source.
func TestAPICopyObjectSSECKeyRotationKeepsChecksumAbsence(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPICopyObjectSSECKeyRotationKeepsChecksumAbsence,
endpoints: []string{"CopyObject", "PutObject", "GetObject"},
})
}
func testAPICopyObjectSSECKeyRotationKeepsChecksumAbsence(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
data := bytes.Repeat([]byte("rotate-without-a-checksum-"), 32*1024)
object := "copy-checksum/rotate-keeps-checksum-absence.bin"
oldKey := bytes.Repeat([]byte{0x53}, 32)
newKey := bytes.Repeat([]byte{0x64}, 32)
// Assert the raw stored bytes rather than decryptChecksums, which also
// returns an empty map when it cannot unseal a checksum that is there.
storedChecksum := func() []byte {
t.Helper()
oi, err := obj.GetObjectInfo(t.Context(), bucketName, object, ObjectOptions{})
if err != nil {
t.Fatalf("GetObjectInfo(%s) failed: %v", object, err)
}
return oi.Checksum
}
putCopyChecksumSource(t, apiRouter, credentials, bucketName, object, data, copyChecksumSSECHeaders(oldKey))
if before := storedChecksum(); len(before) != 0 {
t.Fatalf("%s: invalid precondition, the source already carries a checksum: %x", instanceType, before)
}
rotate := copyChecksumSSECHeaders(newKey)
for name, value := range copyChecksumSSECCopySource(oldKey) {
rotate[name] = value
}
rec := copyChecksumRequest(t, apiRouter, credentials, bucketName, object, object, rotate)
if rec.Code != http.StatusOK {
t.Fatalf("%s: headerless rotation failed: %d %s", instanceType, rec.Code, rec.Body.String())
}
var copyResponse CopyObjectResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &copyResponse); err != nil {
t.Fatalf("unable to decode CopyObjectResult: %v", err)
}
if copyResponse.ChecksumCRC32 != "" || copyResponse.ChecksumCRC32C != "" ||
copyResponse.ChecksumSHA1 != "" || copyResponse.ChecksumSHA256 != "" ||
copyResponse.ChecksumCRC64NVME != "" || copyResponse.ChecksumType != "" {
t.Fatalf("%s: headerless rotation reported a checksum it did not compute: %s",
instanceType, rec.Body.String())
}
if after := storedChecksum(); len(after) != 0 {
t.Fatalf("%s: headerless rotation attached a checksum: %x", instanceType, after)
}
req, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, copyChecksumSSECHeaders(newKey))
if err != nil {
t.Fatalf("failed to build GetObject request: %v", err)
}
response := httptest.NewRecorder()
apiRouter.ServeHTTP(response, req)
if response.Code != http.StatusOK || !bytes.Equal(response.Body.Bytes(), data) {
t.Fatalf("%s: post-rotation GetObject returned %d with %d bytes, want 200 with %d bytes: %s",
instanceType, response.Code, response.Body.Len(), len(data), response.Body.String())
}
}
// TestAPICopyObjectSSECKeyRotationReplicaKeepsFastPath pins that a replica-trusted
// rotation keeps the in-place fast path even when it carries a checksum algorithm
// header. Such a request reads its source without decrypting it, so a rewrite
// would hash ciphertext and would skip the source-key check that a zero byte read
// performs, and a replica has to keep the checksum its source assigned anyway.
func TestAPICopyObjectSSECKeyRotationReplicaKeepsFastPath(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPICopyObjectSSECKeyRotationReplicaKeepsFastPath,
endpoints: []string{"CopyObject", "PutObject", "GetObject"},
})
}
func testAPICopyObjectSSECKeyRotationReplicaKeepsFastPath(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
previousTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = previousTLS }()
oldKey := bytes.Repeat([]byte{0x71}, 32)
newKey := bytes.Repeat([]byte{0x82}, 32)
wrongKey := bytes.Repeat([]byte{0x93}, 32)
rotateAsReplica := func(sourceKey []byte) map[string]string {
headers := copyChecksumSSECHeaders(newKey)
headers[xhttp.AmzChecksumAlgo] = hash.ChecksumSHA256.String()
for name, value := range copyChecksumSSECCopySource(sourceKey) {
headers[name] = value
}
headers[xhttp.MinIOSourceReplicationRequest] = "true"
headers[xhttp.AmzBucketReplicationStatus] = objectreplication.Replica.String()
return headers
}
data := bytes.Repeat([]byte("replica-rotation-keeps-its-checksum-"), 1024)
object := "copy-checksum/replica-rotate.bin"
putHeaders := copyChecksumSSECHeaders(oldKey)
putHeaders[xhttp.AmzChecksumCRC32] = mustChecksum(t, hash.ChecksumCRC32, data)
putCopyChecksumSource(t, apiRouter, credentials, bucketName, object, data, putHeaders)
rec := copyChecksumRequest(t, apiRouter, credentials, bucketName, object, object, rotateAsReplica(oldKey))
if rec.Code != http.StatusOK {
t.Fatalf("%s: replica rotation with an algorithm header failed: %d %s",
instanceType, rec.Code, rec.Body.String())
}
assertCopyChecksumResponse(t, rec, hash.ChecksumCRC32, data)
req, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, copyChecksumSSECHeaders(newKey))
if err != nil {
t.Fatalf("failed to build GetObject request: %v", err)
}
response := httptest.NewRecorder()
apiRouter.ServeHTTP(response, req)
if response.Code != http.StatusOK || !bytes.Equal(response.Body.Bytes(), data) {
t.Fatalf("%s: post-rotation GetObject returned %d with %d bytes, want 200 with %d bytes: %s",
instanceType, response.Code, response.Body.Len(), len(data), response.Body.String())
}
// A zero byte source is the case where the fast path is the only thing that
// still authenticates the rotated-out key.
empty := "copy-checksum/replica-rotate-empty.bin"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, empty, nil, copyChecksumSSECHeaders(oldKey))
wrong := copyChecksumRequest(t, apiRouter, credentials, bucketName, empty, empty, rotateAsReplica(wrongKey))
if wrong.Code != http.StatusForbidden {
t.Fatalf("%s: replica rotation of an empty object with the wrong source key returned %d, want 403: %s",
instanceType, wrong.Code, wrong.Body.String())
}
}
@@ -0,0 +1,241 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"net/http"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/auth"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// enableBucketObjectLock puts a lock-enabled configuration on an existing
// bucket so checkPutObjectLockAllowed accepts a legal-hold header instead of
// rejecting the request with ErrInvalidBucketObjectLockConfiguration.
func enableBucketObjectLock(t *testing.T, bucket string) {
t.Helper()
meta, err := globalBucketMetadataSys.Get(bucket)
if err != nil {
t.Fatalf("unable to read bucket metadata for %s: %v", bucket, err)
}
updated := meta
updated.ObjectLockConfigXML = enabledBucketObjectLockConfig
updated.VersioningConfigXML = enabledBucketVersioningConfig
// The XML alone is not enough: BucketMetadata keeps a parsed copy that the
// lookups actually read, and it is only populated by parseAllConfigs.
if err := updated.parseAllConfigs(t.Context(), newObjectLayerFn()); err != nil {
t.Fatalf("unable to parse bucket metadata for %s: %v", bucket, err)
}
globalBucketMetadataSys.Set(bucket, updated)
}
// Exercise #165 and #166 together, including hold OFF, default retention and
// explicit subsecond retention. Ordinary copies must not inherit source locks.
func TestAPIFederatedCopyObjectLockParity(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
testKMS, err := kms.NewBuiltin(federationTestKMSKeyID, bytes.Repeat([]byte{0x58}, 32))
if err != nil {
t.Fatal(err)
}
previousKMS := GlobalKMS
GlobalKMS = testKMS
defer func() { GlobalKMS = previousKMS }()
remoteBucket, capture, cleanup := setupCopyObjectFederation(t, obj, router, instanceType, bucket)
defer cleanup()
enableBucketObjectLock(t, bucket)
until := UTCNow().Add(7 * 24 * time.Hour).Truncate(time.Second).Add(789 * time.Millisecond).Format(time.RFC3339Nano)
for _, sourceType := range []string{"plain", "s3"} {
source := sourceType + "-held-source"
headers := federationSSEHeaders(sourceType, 0, false)
headers[xhttp.AmzObjectLockLegalHold] = "ON"
putCopyChecksumSource(t, router, cred, bucket, source, []byte("held source"), headers)
for _, tc := range []struct {
name, hold, defaultMode, explicitMode string
replace bool
}{
{name: "on", hold: "ON"},
{name: "off", hold: "OFF"},
{name: "no source inheritance"},
{name: "on with default", hold: "ON", defaultMode: "GOVERNANCE"},
{name: "off with default", hold: "OFF", defaultMode: "COMPLIANCE"},
{name: "explicit retention", explicitMode: "GOVERNANCE"},
{name: "hold and explicit retention", hold: "ON", explicitMode: "COMPLIANCE"},
{name: "replace metadata", hold: "ON", explicitMode: "GOVERNANCE", replace: true},
} {
t.Run(instanceType+"/"+sourceType+"/"+tc.name, func(t *testing.T) {
setTestBucketDefaultRetention(t, bucket, tc.defaultMode)
setTestBucketDefaultRetention(t, remoteBucket, tc.defaultMode)
headers := map[string]string{}
if tc.hold != "" {
headers[xhttp.AmzObjectLockLegalHold] = tc.hold
}
if tc.explicitMode != "" {
headers[xhttp.AmzObjectLockMode] = tc.explicitMode
headers[xhttp.AmzObjectLockRetainUntilDate] = until
}
if tc.replace {
headers[xhttp.AmzMetadataDirective] = "REPLACE"
headers["X-Amz-Meta-Origin"] = "replacement"
}
capture.mu.Lock()
capture.headers = nil
capture.mu.Unlock()
for _, destinationBucket := range []string{bucket, remoteBucket} {
destination := sourceType + "-copy-" + strings.ReplaceAll(tc.name, " ", "-")
rec := federatedCopyRequest(t, router, cred, bucket, source, destinationBucket, destination, headers)
if rec.Code != http.StatusOK {
t.Fatalf("copy to %s: %d %s", destinationBucket, rec.Code, rec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), destinationBucket, destination, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if got := objectlock.GetObjectLegalHoldMeta(info.UserDefined).Status; string(got) != tc.hold {
t.Errorf("stored hold = %q, want %q", got, tc.hold)
}
retention := objectlock.GetObjectRetentionMeta(info.UserDefined)
mode := tc.explicitMode
if mode == "" {
mode = tc.defaultMode
}
if string(retention.Mode) != mode {
t.Errorf("stored retention = %q, want %q", retention.Mode, mode)
}
if tc.explicitMode != "" && retention.RetainUntilDate.Format(time.RFC3339Nano) != until {
t.Errorf("retention date lost precision: %s, want %s", retention.RetainUntilDate, until)
}
for key := range info.UserDefined {
if stringsHasPrefixFold(key, "X-Amz-Meta-X-Amz-Object-Lock-") {
t.Errorf("lock state became user metadata: %s", key)
}
}
if tc.replace && info.UserDefined["X-Amz-Meta-Origin"] != "replacement" {
t.Errorf("replacement metadata was lost: %v", info.UserDefined)
}
}
typed, asMetadata := capture.legalHoldHeaders()
if len(asMetadata) != 0 || (tc.hold != "" && strings.Join(typed, "") != tc.hold) || (tc.hold == "" && len(typed) != 0) {
t.Errorf("forwarded hold headers = %v, metadata = %v; want %q", typed, asMetadata, tc.hold)
}
})
}
}
},
})
}
// legalHoldHeaders returns the forwarded legal-hold headers the remote saw,
// separated into the real Object Lock header and the user-metadata spelling
// minio-go produces for an unrecognized UserMetadata key.
func (c *federationRemoteCapture) legalHoldHeaders() (typed, asMetadata []string) {
c.mu.Lock()
defer c.mu.Unlock()
metaKey := "X-Amz-Meta-" + xhttp.AmzObjectLockLegalHold
for _, h := range c.headers {
for k, v := range h {
switch {
case strings.EqualFold(k, xhttp.AmzObjectLockLegalHold):
typed = append(typed, strings.Join(v, ","))
case strings.EqualFold(k, metaKey):
asMetadata = append(asMetadata, strings.Join(v, ","))
}
}
}
return typed, asMetadata
}
// TestAPIFederatedCopyObjectLegalHold drives the legacy etcd federation branch
// of CopyObjectHandler with an explicit legal hold on the copy.
//
// Before the fix the resolved hold was forwarded inside
// PutObjectOptions.UserMetadata. minio-go's Header() prefixes every
// UserMetadata key it does not recognize with "x-amz-meta-", and
// x-amz-object-lock-legal-hold is in neither supportedHeaders nor isAmzHeader,
// so the hold reached the remote as X-Amz-Meta-X-Amz-Object-Lock-Legal-Hold.
// The destination stored no hold and the copy still answered 200 -- a silent
// loss of a WORM control (#166). Retention requested on the same copy survived,
// which is what made it easy to miss.
func TestAPIFederatedCopyObjectLegalHold(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectLegalHold,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
func testAPIFederatedCopyObjectLegalHold(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
data := []byte("federated copy with a legal hold")
srcObject := "federation/legal-hold-source"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, data, nil)
remoteBucket, capture, cleanup := setupCopyObjectFederation(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
// Both roles read bucket metadata from the shared backend in this fixture,
// so one configuration covers the proxy's own check and the remote's.
enableBucketObjectLock(t, bucketName)
enableBucketObjectLock(t, remoteBucket)
dstObject := "federation/legal-hold-destination"
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject,
map[string]string{xhttp.AmzObjectLockLegalHold: string(objectlock.LegalHoldOn)})
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject with a legal hold failed: %d %s",
instanceType, rec.Code, rec.Body.String())
}
// The wire is the point: the hold must arrive as the Object Lock header,
// never as user metadata. A 200 with the metadata spelling is exactly the
// silent loss this test exists for.
typed, asMetadata := capture.legalHoldHeaders()
if len(asMetadata) != 0 {
t.Fatalf("%s: legal hold forwarded as user metadata %v; the destination stores no hold",
instanceType, asMetadata)
}
if len(typed) == 0 {
t.Fatalf("%s: no %s header reached the remote deployment", instanceType, xhttp.AmzObjectLockLegalHold)
}
for _, got := range typed {
if !strings.EqualFold(got, string(objectlock.LegalHoldOn)) {
t.Fatalf("%s: forwarded legal hold = %q, want %q", instanceType, got, objectlock.LegalHoldOn)
}
}
// And it must actually be stored on the destination version.
oi, err := objectAPI.GetObjectInfo(t.Context(), remoteBucket, dstObject, ObjectOptions{})
if err != nil {
t.Fatalf("%s: unable to stat the federated copy destination: %v", instanceType, err)
}
if hold := objectlock.GetObjectLegalHoldMeta(oi.UserDefined); hold.Status != objectlock.LegalHoldOn {
t.Fatalf("%s: destination legal hold = %q, want %q (metadata: %v)",
instanceType, hold.Status, objectlock.LegalHoldOn, oi.UserDefined)
}
}
+167
View File
@@ -0,0 +1,167 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-only
package cmd
import (
"bytes"
"encoding/xml"
"net/http"
"net/url"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/amztime"
"github.com/minio/minio/internal/auth"
sse "github.com/minio/minio/internal/bucket/encryption"
"github.com/minio/minio/internal/event"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
"github.com/minio/minio/internal/pubsub"
)
func TestAPIFederatedCopyObjectVersionAndEvent(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
endpoints: []string{"CopyObject", "PutObject", "GetObject", "HeadObject"},
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
testKMS, err := kms.NewBuiltin(federationTestKMSKeyID, bytes.Repeat([]byte{0x58}, 32))
if err != nil {
t.Fatal(err)
}
previousKMS := GlobalKMS
GlobalKMS = testKMS
defer func() { GlobalKMS = previousKMS }()
restore := setCopyChecksumCompression(true)
defer restore()
remoteBucket, _, cleanup := setupCopyObjectFederation(t, obj, router, instanceType, bucket)
defer cleanup()
enableBucketObjectLock(t, remoteBucket)
events := make(chan event.Event, 8)
done := make(chan struct{})
defer close(done)
if err := globalHTTPListen.Subscribe(pubsub.MaskFromMaskable(event.ObjectCreatedCopy), events, done, nil); err != nil {
t.Fatal(err)
}
for _, kind := range []string{"plain", "compressed", "encrypted"} {
t.Run(instanceType+"/"+kind, func(t *testing.T) {
data := []byte("logical object size")
source := kind + "-source.bin"
var headers map[string]string
if kind == "compressed" {
data = bytes.Repeat(data, 8192)
source = kind + "-source.txt"
}
if kind == "encrypted" {
headers = federationSSEHeaders("s3", 0, false)
}
putCopyChecksumSource(t, router, cred, bucket, source, data, headers)
destination := "result/" + kind + " with space.bin"
rec := federatedCopyRequest(t, router, cred, bucket, source, remoteBucket, destination, nil)
if rec.Code != http.StatusOK {
t.Fatalf("copy: %d %s", rec.Code, rec.Body.String())
}
versionID := strings.Join(rec.Header()[xhttp.AmzVersionID], "")
if versionID == "" {
t.Error("copy response omitted destination version ID")
} else {
info, err := obj.GetObjectInfo(t.Context(), remoteBucket, destination, ObjectOptions{VersionID: versionID})
if err != nil || info.VersionID != versionID {
t.Fatalf("response does not name the written version: %v, %q", err, info.VersionID)
}
var response CopyObjectResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatal(err)
}
if want := amztime.ISO8601Format(info.ModTime.UTC()); response.LastModified != want {
t.Errorf("copy time = %q, want written version time %q", response.LastModified, want)
}
}
select {
case evt := <-events:
key, err := url.QueryUnescape(evt.S3.Object.Key)
if err != nil || key != destination || evt.S3.Bucket.Name != remoteBucket || evt.S3.Object.Size != int64(len(data)) || evt.S3.Object.VersionID == "" || evt.S3.Object.VersionID != versionID {
t.Errorf("copy event does not describe the written object: %+v", evt.S3)
}
case <-time.After(5 * time.Second):
t.Fatal("copy event was not emitted")
}
})
}
},
})
}
func TestAPIFederatedCopyObjectDestinationSSEDefaults(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
endpoints: []string{"CopyObject", "PutObject", "GetObject", "HeadObject"},
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
data := []byte("destination chooses default encryption")
putCopyChecksumSource(t, router, cred, bucket, "source", data, nil)
testKMS, err := kms.NewBuiltin(federationTestKMSKeyID, bytes.Repeat([]byte{0x58}, 32))
if err != nil {
t.Fatal(err)
}
previousKMS, previousAuto := GlobalKMS, globalAutoEncryption
GlobalKMS, globalAutoEncryption = testKMS, true
defer func() { GlobalKMS, globalAutoEncryption = previousKMS, previousAuto }()
remoteBucket, capture, cleanup := setupCopyObjectFederationRemote(t, obj, router, instanceType, bucket, true)
defer cleanup()
destinationConfig, err := sse.ParseBucketSSEConfig(strings.NewReader(`<ServerSideEncryptionConfiguration><Rule><ApplyServerSideEncryptionByDefault><SSEAlgorithm>aws:kms</SSEAlgorithm><KMSMasterKeyID>destination-bucket-key</KMSMasterKeyID></ApplyServerSideEncryptionByDefault></Rule></ServerSideEncryptionConfiguration>`))
if err != nil {
t.Fatal(err)
}
for _, kind := range []string{"plain", "s3", "kms", "kms-default", "c"} {
t.Run(instanceType+"/"+kind, func(t *testing.T) {
var headers map[string]string
if kind == "kms-default" {
headers = map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS}
} else {
headers = federationSSEHeaders(kind, 0x22, false)
}
rec := federatedCopyRequest(t, router, cred, bucket, "source", remoteBucket, kind, headers)
if rec.Code != http.StatusOK {
t.Fatalf("copy: %d %s", rec.Code, rec.Body.String())
}
capture.mu.Lock()
forwarded := capture.headers[len(capture.headers)-1].Clone()
capture.mu.Unlock()
if kind == "plain" {
if forwarded.Get(xhttp.AmzServerSideEncryption) != "" {
t.Errorf("proxy injected SSE %q into a request with no client SSE", forwarded.Get(xhttp.AmzServerSideEncryption))
}
// With nothing injected, the KMS encryption the destination
// stores can only be the remote applying its own defaults.
info, err := obj.GetObjectInfo(t.Context(), remoteBucket, kind, ObjectOptions{})
if err != nil || federationStoredSSE(info.UserDefined) != "kms" {
t.Errorf("remote destination did not apply its own auto-encryption: %v, %v", err, info.UserDefined)
}
}
// Replay the actual inbound headers against an independent bucket
// configuration. No handler flips shared globals while serving.
destinationConfig.Apply(forwarded, sse.ApplyOptions{})
wantKey := headers[xhttp.AmzServerSideEncryptionKmsID]
if kind == "plain" {
wantKey = "destination-bucket-key"
}
if got := forwarded.Get(xhttp.AmzServerSideEncryptionKmsID); got != wantKey {
t.Errorf("destination selected key %q, want %q", got, wantKey)
}
})
}
// A local destination still inherits the server's auto-encryption.
rec := federatedCopyRequest(t, router, cred, bucket, "source", bucket, "local-copy", nil)
if rec.Code != http.StatusOK {
t.Fatalf("local copy: %d %s", rec.Code, rec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), bucket, "local-copy", ObjectOptions{})
if err != nil || federationStoredSSE(info.UserDefined) != "kms" {
t.Errorf("local copy lost auto-encryption: %v, %v", err, info.UserDefined)
}
},
})
}
+956
View File
@@ -0,0 +1,956 @@
// Copyright (c) 2015-2025 MinIO, Inc.
// Copyright (c) 2025-2026 PGSTY
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"crypto/md5"
"encoding/base64"
"encoding/json"
"encoding/xml"
"io"
"maps"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"sync"
"testing"
miniogo "github.com/minio/minio-go/v7"
miniocredentials "github.com/minio/minio-go/v7/pkg/credentials"
"github.com/minio/minio-go/v7/pkg/set"
"github.com/minio/minio/internal/amztime"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/config/dns"
"github.com/minio/minio/internal/crypto"
"github.com/minio/minio/internal/hash"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// federationRemoteCapture records the exact request headers the remote
// deployment receives on each forwarded request, so a test can assert what
// actually crossed the wire rather than what the destination ends up storing.
type federationRemoteCapture struct {
mu sync.Mutex
headers []http.Header
}
type federationResponseFilter struct {
http.ResponseWriter
filter func(http.Header)
}
func (w federationResponseFilter) WriteHeader(status int) {
w.filter(w.Header())
w.ResponseWriter.WriteHeader(status)
}
func (c *federationRemoteCapture) record(h http.Header) {
c.mu.Lock()
c.headers = append(c.headers, h.Clone())
c.mu.Unlock()
}
// reservedKeys returns every reserved-prefix header key seen across all
// forwarded requests.
func (c *federationRemoteCapture) reservedKeys() []string {
c.mu.Lock()
defer c.mu.Unlock()
var keys []string
for _, h := range c.headers {
for k := range h {
if stringsHasPrefixFold(k, ReservedMetadataPrefix) {
keys = append(keys, k)
}
}
}
return keys
}
// setupCopyObjectFederation makes the current process play both federation
// roles for a whole-object CopyObject. The destination bucket really exists in
// the shared backend, but the proxy's own bucket lookup is told it lives on a
// remote deployment, so CopyObjectHandler takes the legacy etcd federation
// branch and forwards the write through getRemoteInstanceClient and minio-go
// into a second HTTP endpoint that serves the real PutObjectHandler. That
// endpoint records the inbound headers first so a caller can inspect the wire.
//
// GetBucketLocation is deliberately left unregistered on that endpoint: the
// remoteBucketObjectLayer reports remoteBucket as missing (so the proxy takes
// the federation branch), which would also fail a real GetBucketLocation the
// minio-go client probes for. Leaving it unregistered makes the probe fall back
// to the default region, exactly as TestAPIFederatedCopyObjectPartChecksum does.
func setupCopyObjectFederation(t *testing.T, objectAPI ObjectLayer, apiRouter http.Handler,
instanceType, srcBucket string, responseFilters ...func(http.Header),
) (remoteBucket string, capture *federationRemoteCapture, cleanup func()) {
t.Helper()
return setupCopyObjectFederationRemote(t, objectAPI, apiRouter, instanceType, srcBucket, false, responseFilters...)
}
// setupCopyObjectFederationTLS is setupCopyObjectFederation with a TLS remote
// endpoint and globalIsTLS set, so an SSE-C copy is accepted on both hops: the
// proxy's own SSE-C transport gate and the minio-go client's SSE-C policy. The
// proxy client is built exactly as in production, trusting the test
// certificate for the duration.
func setupCopyObjectFederationTLS(t *testing.T, objectAPI ObjectLayer, apiRouter http.Handler,
instanceType, srcBucket string,
) (remoteBucket string, cleanup func()) {
t.Helper()
remoteBucket, _, cleanup = setupCopyObjectFederationRemote(t, objectAPI, apiRouter, instanceType, srcBucket, true)
return remoteBucket, cleanup
}
func setupCopyObjectFederationRemote(t *testing.T, objectAPI ObjectLayer, apiRouter http.Handler,
instanceType, srcBucket string, secure bool, responseFilters ...func(http.Header),
) (remoteBucket string, capture *federationRemoteCapture, cleanup func()) {
t.Helper()
remoteBucket = getRandomBucketName()
if err := objectAPI.MakeBucket(t.Context(), remoteBucket, MakeBucketOptions{}); err != nil {
t.Fatalf("%s: unable to create the remote bucket: %v", instanceType, err)
}
capture = &federationRemoteCapture{}
handler := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
capture.record(r.Header)
if len(responseFilters) != 0 && r.Method == http.MethodPut {
w = federationResponseFilter{ResponseWriter: w, filter: responseFilters[0]}
}
apiRouter.ServeHTTP(w, r)
})
var remote *httptest.Server
if secure {
remote = httptest.NewTLSServer(handler)
} else {
remote = httptest.NewServer(handler)
}
host, port, _ := strings.Cut(remote.Listener.Addr().String(), ":")
globalObjLayerMutex.Lock()
previousLayer := globalObjectAPI
globalObjectAPI = remoteBucketObjectLayer{ObjectLayer: previousLayer, remoteBucket: remoteBucket}
globalObjLayerMutex.Unlock()
previousDNS, previousFederation, previousIPs := globalDNSConfig, globalBucketFederation, globalDomainIPs
previousTLS, previousClient := globalIsTLS, getRemoteInstanceClient
globalDNSConfig = federationTestDNS{records: map[string][]dns.SrvRecord{
srcBucket: {{Host: host, Port: json.Number(port)}},
remoteBucket: {{Host: host, Port: json.Number(port)}},
}}
// Every DNS record resolves to this process, so the bucket forwarding
// middleware always serves locally and only the handler proxies.
globalDomainIPs = set.CreateStringSet(remote.Listener.Addr().String())
globalBucketFederation = true
if secure {
globalIsTLS = true
transport := remote.Client().Transport
getRemoteInstanceClient = func(r *http.Request, host string) (*miniogo.Core, error) {
cred := getReqAccessCred(r, globalSite.Region())
core, err := miniogo.NewCore(host, &miniogo.Options{
Creds: miniocredentials.NewStaticV4(cred.AccessKey, cred.SecretKey, ""),
Secure: true,
Transport: federatedWriteTransport{transport},
})
if err != nil {
return nil, err
}
core.SetAppInfo(federatedInternalAppName, ReleaseTag)
return core, nil
}
}
cleanup = func() {
remote.Close()
globalObjLayerMutex.Lock()
globalObjectAPI = previousLayer
globalObjLayerMutex.Unlock()
globalDNSConfig, globalBucketFederation, globalDomainIPs = previousDNS, previousFederation, previousIPs
globalIsTLS, getRemoteInstanceClient = previousTLS, previousClient
}
return remoteBucket, capture, cleanup
}
// federatedCopyRequest drives a whole-object CopyObject at the proxy that
// forwards across deployments, and returns the recorded response.
func federatedCopyRequest(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
srcBucket, srcObject, dstBucket, dstObject string, headers map[string]string,
) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(http.MethodPut, getCopyObjectURL("", dstBucket, dstObject),
0, nil, credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatalf("failed to build federated CopyObject request: %v", err)
}
req.Header.Set(xhttp.AmzCopySource, SlashSeparator+pathJoin(srcBucket, srcObject))
// Re-sign so x-amz-copy-source is covered by the signature, as real S3
// clients send it; the verifier rejects unsigned x-amz-* headers.
if credentials.AccessKey != "" && credentials.SecretKey != "" {
if err := signRequestV4(req, credentials.AccessKey, credentials.SecretKey); err != nil {
t.Fatalf("failed to re-sign federated CopyObject request: %v", err)
}
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return rec
}
func TestAPIFederatedCopyObjectRejectsRawSSECReplica(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
endpoints: []string{"CopyObject", "PutObject"},
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
remoteBucket, capture, cleanup := setupCopyObjectFederationRemote(t, obj, router, instanceType, bucket, true)
defer cleanup()
for _, body := range []string{"", "raw SSE-C replica"} {
for _, withKey := range []bool{false, true} {
t.Run("size="+strconv.Itoa(len(body))+"/key="+strconv.FormatBool(withKey), func(t *testing.T) {
putCopyChecksumSource(t, router, cred, bucket, "source", []byte(body), federationSSEHeaders("c", 0x11, false))
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
}
if withKey {
maps.Copy(headers, federationSSEHeaders("c", 0x11, true))
}
destination := "destination-" + strconv.Itoa(len(body)) + "-" + strconv.FormatBool(withKey)
rec := federatedCopyRequest(t, router, cred, bucket, "source", remoteBucket, destination, headers)
var response APIErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatal(err)
}
if rec.Code != http.StatusNotImplemented || response.Code != "NotImplemented" ||
!strings.Contains(response.Message, "federated raw SSE-C replica CopyObject") {
t.Errorf("copy = %d %s, want explicit 501 rejection", rec.Code, rec.Body.String())
}
if _, err := obj.GetObjectInfo(t.Context(), remoteBucket, destination, ObjectOptions{}); !isErrObjectNotFound(err) {
t.Errorf("rejected copy created a destination: %v", err)
}
})
}
}
capture.mu.Lock()
defer capture.mu.Unlock()
if len(capture.headers) != 0 {
t.Errorf("rejected copies made %d remote requests", len(capture.headers))
}
},
})
}
// TestAPIFederatedCopyObjectInlineSource drives the legacy etcd federation
// branch of CopyObjectHandler end to end for a source object stored inline.
// Such a source carries x-minio-internal-inline-data in its stored metadata,
// which the remote deployment rejects as a reserved-prefix header. Before the
// fix the forwarded write failed with 400 InvalidArgument; the copy must now
// succeed and forward no reserved-prefix metadata at all.
func TestAPIFederatedCopyObjectInlineSource(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectInlineSource,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
func testAPIFederatedCopyObjectInlineSource(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
// A small object is stored inline, so its stored metadata carries
// x-minio-internal-inline-data. A user metadata key rides along to prove
// the fix strips only the reserved class, never ordinary metadata.
data := []byte("federated inline copy 26b!")
srcObject := "federation/inline-source"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, data,
map[string]string{"X-Amz-Meta-Origin": "inline-source"})
remoteBucket, capture, cleanup := setupCopyObjectFederation(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
dstObject := "federation/inline-destination"
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject, nil)
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject of an inline source failed: %d %s",
instanceType, rec.Code, rec.Body.String())
}
// No reserved-prefix header may reach the remote deployment on any of the
// forwarded requests.
if leaked := capture.reservedKeys(); len(leaked) != 0 {
t.Fatalf("%s: forwarded reserved metadata to the remote: %v", instanceType, leaked)
}
// The destination object must be readable and byte-identical, and it must
// still carry the copied user metadata.
gr, err := objectAPI.GetObjectNInfo(t.Context(), remoteBucket, dstObject, nil, nil, ObjectOptions{})
if err != nil {
t.Fatalf("%s: unable to read the federated copy destination: %v", instanceType, err)
}
got, err := io.ReadAll(gr)
closeErr := gr.Close()
if err != nil {
t.Fatalf("%s: reading the federated copy destination failed: %v", instanceType, err)
}
if closeErr != nil {
t.Fatalf("%s: closing the federated copy destination failed: %v", instanceType, closeErr)
}
if !bytes.Equal(got, data) {
t.Fatalf("%s: federated copy destination body = %q, want %q", instanceType, got, data)
}
if origin, ok := gr.ObjInfo.UserDefined["X-Amz-Meta-Origin"]; !ok || origin != "inline-source" {
t.Fatalf("%s: destination lost copied user metadata: %v", instanceType, gr.ObjInfo.UserDefined)
}
}
// TestAPIFederatedCopyObjectRequestedChecksum drives the legacy etcd federation
// branch of CopyObjectHandler and verifies that a server-side checksum is both
// returned and persisted, matching the local CopyObject path. Before the fix
// the federated copy forwarded the write without asking for a checksum and
// discarded whatever the remote returned, so the response carried an empty
// checksum even when the client requested one (#99).
//
// The no-algorithm case is included deliberately: a checksum-less source gains
// the S3 default CRC-64NVME full-object checksum on the local path, so the
// federated path must return the same. "No requested algorithm" does not mean
// "no checksum".
func TestAPIFederatedCopyObjectRequestedChecksum(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectRequestedChecksum,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
func testAPIFederatedCopyObjectRequestedChecksum(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
data := []byte("federated copy checksum body")
srcObject := "federation/checksum-source"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, data, nil)
remoteBucket, _, cleanup := setupCopyObjectFederation(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
cases := []struct {
name string
typ hash.ChecksumType
explicit bool
}{
{name: "CRC32", typ: hash.ChecksumCRC32, explicit: true},
{name: "CRC32C", typ: hash.ChecksumCRC32C, explicit: true},
{name: "SHA1", typ: hash.ChecksumSHA1, explicit: true},
{name: "SHA256", typ: hash.ChecksumSHA256, explicit: true},
{name: "CRC64NVME", typ: hash.ChecksumCRC64NVME, explicit: true},
// Default: no requested algorithm still yields the S3 CRC-64NVME.
{name: "default", typ: hash.ChecksumCRC64NVME},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
var headers map[string]string
if tc.explicit {
headers = map[string]string{xhttp.AmzChecksumAlgo: tc.typ.String()}
}
dstObject := "federation/checksum-destination-" + tc.name
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject, headers)
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject failed: %d %s", instanceType, rec.Code, rec.Body.String())
}
// The CopyObjectResult must carry the checksum of the copied bytes.
assertCopyChecksumResponse(t, rec, tc.typ, data)
// The remote must have persisted that same checksum.
assertCopyChecksum(t, objectAPI, remoteBucket, dstObject, tc.typ, data, false, nil)
})
}
}
func TestAPIFederatedCopyObjectRejectsInvalidRemoteChecksum(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, credentials auth.Credentials, t *testing.T) {
data := []byte("remote checksum response fixture")
putCopyChecksumSource(t, router, credentials, bucket, "source", data, nil)
// "" missing, "invalid-base64" unparseable, "YQ==" wrong digest
// length, and "<valid>-0" a multipart-marked value that a single
// forwarded PutObject must never yield (it would mislabel the
// destination as composite).
for _, value := range []string{"", "invalid-base64", "YQ==", mustChecksum(t, hash.ChecksumCRC32, data) + "-0"} {
t.Run("checksum="+value, func(t *testing.T) {
remoteBucket, _, cleanup := setupCopyObjectFederation(t, obj, router, instanceType, bucket, func(header http.Header) {
header.Set(xhttp.AmzChecksumCRC32, value)
})
defer cleanup()
rec := federatedCopyRequest(t, router, credentials, bucket, "source", remoteBucket, "destination",
map[string]string{xhttp.AmzChecksumAlgo: "CRC32"})
if rec.Code < 500 {
t.Fatalf("invalid remote checksum must fail the copy, got %d: %s", rec.Code, rec.Body.String())
}
})
}
},
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
// TestAPIFederatedCopyObjectChecksumIsBoundToWrite guards the checksum
// representation: a federated copy that requests one algorithm must return only
// that algorithm, and a copy of a checksum-less source without a requested
// algorithm must never fabricate one other than the S3 default.
func TestAPIFederatedCopyObjectChecksumIsBoundToWrite(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectChecksumIsBoundToWrite,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
func testAPIFederatedCopyObjectChecksumIsBoundToWrite(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
data := []byte("federated copy single checksum body")
srcObject := "federation/single-checksum-source"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, data, nil)
remoteBucket, _, cleanup := setupCopyObjectFederation(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
dstObject := "federation/single-checksum-destination"
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject,
map[string]string{xhttp.AmzChecksumAlgo: hash.ChecksumCRC32.String()})
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject failed: %d %s", instanceType, rec.Code, rec.Body.String())
}
var response CopyObjectResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatalf("%s: unable to decode CopyObjectResult: %v", instanceType, err)
}
if response.ChecksumCRC32 == "" {
t.Fatalf("%s: requested CRC32 checksum missing from response: %s", instanceType, rec.Body.String())
}
// Only the requested algorithm may be present.
if response.ChecksumCRC32C != "" || response.ChecksumSHA1 != "" ||
response.ChecksumSHA256 != "" || response.ChecksumCRC64NVME != "" {
t.Fatalf("%s: response carried checksums beyond the requested CRC32: %s", instanceType, rec.Body.String())
}
}
// TestAPIFederatedCopyObjectEmptySource guards the empty-body regression: a
// checksum-less object gains the S3 default CRC-64NVME, but minio-go streams no
// trailing checksum for a 0-byte body, so the remote returned none, the bind
// found nothing, and every empty-object federated copy 500'd. An empty source
// must now copy with 200 and carry the empty-content checksum, both when a
// checksum is requested explicitly and via the default.
func TestAPIFederatedCopyObjectEmptySource(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectEmptySource,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
func testAPIFederatedCopyObjectEmptySource(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
srcObject := "federation/empty-source"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, nil, nil)
remoteBucket, _, cleanup := setupCopyObjectFederation(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
cases := []struct {
name string
typ hash.ChecksumType
explicit bool
}{
{name: "explicit-CRC32", typ: hash.ChecksumCRC32, explicit: true},
{name: "explicit-SHA256", typ: hash.ChecksumSHA256, explicit: true},
// No requested algorithm: the S3 default CRC-64NVME still applies.
{name: "default-CRC64NVME", typ: hash.ChecksumCRC64NVME},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
var headers map[string]string
if tc.explicit {
headers = map[string]string{xhttp.AmzChecksumAlgo: tc.typ.String()}
}
dstObject := "federation/empty-destination-" + tc.name
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject, headers)
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject of an empty source failed: %d %s",
instanceType, rec.Code, rec.Body.String())
}
// The empty-content digest must be returned and persisted.
assertCopyChecksumResponse(t, rec, tc.typ, nil)
assertCopyChecksum(t, objectAPI, remoteBucket, dstObject, tc.typ, nil, false, nil)
})
}
}
// TestAPIFederatedCopyObjectInheritedChecksum guards that a full-object checksum
// already stored on the source is preserved across a federated copy that
// requests no algorithm. That checksum sets dstOpts.WantChecksum (not
// WantServerSideChecksumType), which the federated branch previously ignored,
// silently dropping the checksum the local path keeps.
func TestAPIFederatedCopyObjectInheritedChecksum(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectInheritedChecksum,
endpoints: []string{"CopyObject", "PutObject", "HeadObject", "GetObject"},
})
}
func testAPIFederatedCopyObjectInheritedChecksum(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
data := []byte("abc")
want := mustChecksum(t, hash.ChecksumCRC32, data) // "NSRBwg=="
srcObject := "federation/inherited-checksum-source"
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, data,
map[string]string{xhttp.AmzChecksumCRC32: want})
remoteBucket, _, cleanup := setupCopyObjectFederation(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
// No algorithm header: the source's stored CRC32 must survive the copy.
dstObject := "federation/inherited-checksum-destination"
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject, nil)
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject failed: %d %s", instanceType, rec.Code, rec.Body.String())
}
assertCopyChecksumResponse(t, rec, hash.ChecksumCRC32, data)
assertCopyChecksum(t, objectAPI, remoteBucket, dstObject, hash.ChecksumCRC32, data, false, nil)
}
// federationTestKMSKeyID names the single key of the builtin KMS the SSE test
// installs; SSE-S3 seals with it implicitly, SSE-KMS names it by ID.
const federationTestKMSKeyID = "federation-test-key"
// federationSSEHeaders returns the request headers that select one server-side
// encryption kind: "plain", "s3", "kms", "kms-context" (SSE-KMS with an
// explicit encryption context), "kms-empty-context" or "c". SSE-C keys derive from keyByte so a
// case can name two distinct customer keys. With copySource the SSE-C key is
// returned in its x-amz-copy-source-* form, the only kind a copy has to name
// for its source; the other kinds then return nothing.
func federationSSEHeaders(kind string, keyByte byte, copySource bool) map[string]string {
h := map[string]string{}
if copySource && kind != "c" {
return h
}
switch kind {
case "plain":
case "s3":
h[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionAES
case "kms", "kms-context", "kms-empty-context":
h[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionKMS
h[xhttp.AmzServerSideEncryptionKmsID] = federationTestKMSKeyID
switch kind {
case "kms-context":
h[xhttp.AmzServerSideEncryptionKmsContext] = base64.StdEncoding.EncodeToString([]byte(`{"tenant":"federation"}`))
case "kms-empty-context":
h[xhttp.AmzServerSideEncryptionKmsContext] = base64.StdEncoding.EncodeToString([]byte(`{}`))
}
case "c":
key := bytes.Repeat([]byte{keyByte}, 32)
sum := md5.Sum(key)
algorithm, customerKey, keyMD5 := xhttp.AmzServerSideEncryptionCustomerAlgorithm,
xhttp.AmzServerSideEncryptionCustomerKey, xhttp.AmzServerSideEncryptionCustomerKeyMD5
if copySource {
algorithm, customerKey, keyMD5 = xhttp.AmzServerSideEncryptionCopyCustomerAlgorithm,
xhttp.AmzServerSideEncryptionCopyCustomerKey, xhttp.AmzServerSideEncryptionCopyCustomerKeyMD5
}
h[algorithm] = xhttp.AmzEncryptionAES
h[customerKey] = base64.StdEncoding.EncodeToString(key)
h[keyMD5] = base64.StdEncoding.EncodeToString(sum[:])
default:
panic("unknown SSE kind " + kind)
}
return h
}
// federationStoredSSE reports which SSE kind stored object metadata declares.
func federationStoredSSE(metadata map[string]string) string {
switch {
case crypto.SSEC.IsEncrypted(metadata):
return "c"
case crypto.S3KMS.IsEncrypted(metadata):
return "kms"
case crypto.S3.IsEncrypted(metadata):
return "s3"
}
return "plain"
}
// federationGetObject reads an object through GetObjectHandler.
func federationGetObject(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
bucket, object string, headers map[string]string,
) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucket, object),
0, nil, credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatalf("failed to build GetObject request: %v", err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return rec
}
// assertFederationCopyReadback checks storage and the public GET/HEAD checksum
// surface with the destination key, including the committed write's time.
func assertFederationCopyReadback(t *testing.T, obj ObjectLayer, router http.Handler, cred auth.Credentials,
bucket, object string, rec *httptest.ResponseRecorder, typ hash.ChecksumType, data []byte, compressed bool, keyHeaders map[string]string,
) ObjectInfo {
t.Helper()
assertCopyChecksumResponse(t, rec, typ, data)
headers := make(http.Header)
for k, v := range keyHeaders {
headers.Set(k, v)
}
info := assertCopyChecksum(t, obj, bucket, object, typ, data, compressed, headers)
var response CopyObjectResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatal(err)
}
if want := amztime.ISO8601Format(info.ModTime.UTC()); info.ModTime.IsZero() || response.LastModified != want {
t.Errorf("copy time = %q, want stored time %q", response.LastModified, want)
}
readHeaders := map[string]string{xhttp.AmzChecksumMode: "ENABLED"}
maps.Copy(readHeaders, keyHeaders)
for _, method := range []string{http.MethodGet, http.MethodHead} {
req, err := newTestSignedRequestV4(method, getGetObjectURL("", bucket, object), 0, nil, cred.AccessKey, cred.SecretKey, readHeaders)
if err != nil {
t.Fatal(err)
}
got := httptest.NewRecorder()
router.ServeHTTP(got, req)
if got.Code != http.StatusOK {
t.Fatalf("destination %s = %d %s", method, got.Code, got.Body.String())
}
if method == http.MethodGet && !bytes.Equal(got.Body.Bytes(), data) {
t.Fatalf("destination GET differs from the %d-byte plaintext", len(data))
}
if got.Header().Get(xhttp.ContentLength) != strconv.Itoa(len(data)) ||
got.Header().Get(typ.Key()) != mustChecksum(t, typ, data) ||
got.Header().Get(xhttp.AmzChecksumType) != xhttp.AmzChecksumTypeFullObject {
t.Fatalf("destination %s length/checksum headers = %v", method, got.Header())
}
}
return info
}
func assertFederationKMSContext(t *testing.T, info ObjectInfo, requested string) {
t.Helper()
decode := func(encoded string) map[string]string {
t.Helper()
value, err := base64.StdEncoding.DecodeString(encoded)
if err != nil {
t.Fatal(err)
}
var context map[string]string
if err := json.Unmarshal(value, &context); err != nil {
t.Fatal(err)
}
return context
}
want := map[string]string{}
if requested != "" {
want = decode(requested)
}
stored, present := info.UserDefined[crypto.MetaContext]
if len(want) == 0 {
if present {
t.Errorf("absent/empty destination context stored client metadata %q", stored)
}
} else if !present || !maps.Equal(decode(stored), want) {
t.Errorf("stored KMS context = %q, want %v", stored, want)
}
}
// TestAPIFederatedCopyObjectSSE guards the legacy etcd federation branch of
// CopyObjectHandler for encrypted sources and destinations (#158). The proxy
// reads its source through getObjectNInfo, which yields the decrypted and
// decompressed bytes, but it used to run the destination encryption locally
// as well and then forward that stream with the source's stored size under
// the destination SSE option. SSE to plain and plain to SSE failed on the
// length mismatch; SSE to SSE matched by coincidence, so the remote encrypted
// the ciphertext a second time and a destination GET returned the inner
// ciphertext with HTTP 200. The proxy must forward the logical bytes with
// their logical size and let the remote encrypt exactly once.
func TestAPIFederatedCopyObjectSSE(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIFederatedCopyObjectSSE,
// Register PutObjectPart and CopyObject before the query-less
// PutObject PUT route; otherwise PutObject shadows them.
endpoints: []string{
"NewMultipart", "PutObjectPart", "CompleteMultipart",
"CopyObject", "PutObject", "HeadObject", "GetObject",
},
})
}
// putFederationMultipartSource uploads parts as one multipart object through
// the API router. Headers select the upload's encryption/checksum; SSE-C keys
// and checksums are also supplied on the parts and completion as required.
func putFederationMultipartSource(t *testing.T, apiRouter http.Handler, credentials auth.Credentials,
bucket, object string, parts [][]byte, headers map[string]string,
) {
t.Helper()
do := func(method, url string, body []byte, headers map[string]string) *httptest.ResponseRecorder {
req, err := newTestSignedRequestV4(method, url, int64(len(body)), bytes.NewReader(body),
credentials.AccessKey, credentials.SecretKey, headers)
if err != nil {
t.Fatalf("failed to build %s %s request: %v", method, url, err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("%s %s failed: %d %s", method, url, rec.Code, rec.Body.String())
}
return rec
}
var upload InitiateMultipartUploadResponse
rec := do(http.MethodPost, getNewMultipartURL("", bucket, object), nil, headers)
if err := xml.Unmarshal(rec.Body.Bytes(), &upload); err != nil {
t.Fatalf("failed to decode NewMultipartUpload response: %v", err)
}
completion := CompleteMultipartUpload{}
partHeaders := map[string]string{}
for _, key := range []string{xhttp.AmzServerSideEncryptionCustomerAlgorithm, xhttp.AmzServerSideEncryptionCustomerKey, xhttp.AmzServerSideEncryptionCustomerKeyMD5} {
if value := headers[key]; value != "" {
partHeaders[key] = value
}
}
checksumType := hash.NewChecksumType(headers[xhttp.AmzChecksumAlgo], headers[xhttp.AmzChecksumType])
for i, part := range parts {
if checksumType.IsSet() {
partHeaders[checksumType.Key()] = mustChecksum(t, checksumType, part)
}
rec := do(http.MethodPut, getPutObjectPartURL("", bucket, object, upload.UploadID, strconv.Itoa(i+1)), part, partHeaders)
completion.Parts = append(completion.Parts, completePartWithChecksum(checksumType, i+1,
canonicalizeETag(rec.Header()[xhttp.ETag][0]), partHeaders[checksumType.Key()]))
}
body, err := xml.Marshal(completion)
if err != nil {
t.Fatalf("failed to encode CompleteMultipartUpload: %v", err)
}
delete(partHeaders, checksumType.Key())
do(http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, upload.UploadID), body, partHeaders)
}
func testAPIFederatedCopyObjectSSE(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
testKMS, err := kms.NewBuiltin(federationTestKMSKeyID, bytes.Repeat([]byte{0x58}, 32))
if err != nil {
t.Fatalf("unable to create the test KMS: %v", err)
}
previousKMS := GlobalKMS
GlobalKMS = testKMS
defer func() { GlobalKMS = previousKMS }()
// .txt sources are stored compressed, SSE-S3 and SSE-KMS ones included;
// SSE-C data is never compressed.
restoreCompression := setCopyChecksumCompression(true)
defer restoreCompression()
remoteBucket, cleanup := setupCopyObjectFederationTLS(t, objectAPI, apiRouter, instanceType, bucketName)
defer cleanup()
// A DARE package holds 64 KiB of plaintext, so the larger bodies span two
// packages and catch a size that accounts for only one.
large := bytes.Repeat([]byte("federated sse copy body "), 64*1024/24+1)[:64*1024+1]
bodies := []struct {
name, ext string
data []byte
requested, inherited hash.ChecksumType
}{
{name: "empty", ext: ".bin"},
{name: "empty-explicit", ext: ".bin", requested: hash.ChecksumSHA256},
{name: "inherited", ext: ".bin", data: []byte("stored encrypted checksum"), inherited: hash.ChecksumCRC32},
{name: "small", ext: ".bin", data: []byte("abc")},
{name: "two-packages", ext: ".bin", data: large},
{name: "compressed", ext: ".txt", data: large},
}
pairs := []struct{ src, dst string }{
{"plain", "s3"},
{"s3", "plain"},
{"s3", "s3"},
{"plain", "c"},
{"c", "plain"},
{"c", "c"},
{"s3", "c"},
{"plain", "kms"},
{"kms", "plain"},
{"kms", "kms"},
{"plain", "kms-context"},
{"kms-context", "kms-context"},
{"kms-context", "kms"},
{"kms-context", "kms-empty-context"},
}
const srcKeyByte, dstKeyByte = 0x11, 0x22
for _, body := range bodies {
for _, pair := range pairs {
// Focus empty and inherited-checksum cases on each encrypted kind.
if body.name == "empty" || body.name == "empty-explicit" || body.name == "inherited" {
if pair.src != pair.dst || pair.src == "kms-context" {
continue
}
}
t.Run(body.name+"/"+pair.src+"-to-"+pair.dst, func(t *testing.T) {
prefix := "federation/sse-" + body.name + "-" + pair.src + "-to-" + pair.dst
srcObject, dstObject := prefix+"-source"+body.ext, prefix+"-destination"+body.ext
sourceHeaders := federationSSEHeaders(pair.src, srcKeyByte, false)
if pair.src == "kms-context" {
sourceHeaders[xhttp.AmzServerSideEncryptionKmsContext] = base64.StdEncoding.EncodeToString([]byte(`{"tenant":"source"}`))
}
if body.inherited.IsSet() {
sourceHeaders[body.inherited.Key()] = mustChecksum(t, body.inherited, body.data)
}
putCopyChecksumSource(t, apiRouter, credentials, bucketName, srcObject, body.data, sourceHeaders)
before, err := objectAPI.GetObjectInfo(t.Context(), bucketName, srcObject, ObjectOptions{})
if err != nil {
t.Fatalf("%s: GetObjectInfo(source) failed: %v", instanceType, err)
}
if got, want := federationStoredSSE(before.UserDefined), strings.Split(pair.src, "-")[0]; got != want {
t.Fatalf("%s: source stored as %s, want %s", instanceType, got, want)
}
if compressed := body.ext == ".txt" && pair.src != "c"; before.IsCompressed() != compressed {
t.Fatalf("%s: source compressed=%v, want %v", instanceType, before.IsCompressed(), compressed)
}
assertFederationKMSContext(t, before, sourceHeaders[xhttp.AmzServerSideEncryptionKmsContext])
if body.inherited.IsSet() {
sourceKeys := make(http.Header)
for k, v := range federationSSEHeaders(pair.src, srcKeyByte, true) {
sourceKeys.Set(k, v)
}
assertCopyChecksum(t, objectAPI, bucketName, srcObject, body.inherited, body.data, false, sourceKeys)
}
headers := federationSSEHeaders(pair.dst, dstKeyByte, false)
if body.requested.IsSet() {
headers[xhttp.AmzChecksumAlgo] = body.requested.String()
}
maps.Copy(headers, federationSSEHeaders(pair.src, srcKeyByte, true))
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject, headers)
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject failed: %d %s", instanceType, rec.Code, rec.Body.String())
}
wantChecksum := hash.ChecksumCRC64NVME
if body.requested.IsSet() {
wantChecksum = body.requested
} else if body.inherited.IsSet() {
wantChecksum = body.inherited
}
var getHeaders map[string]string
if pair.dst == "c" {
getHeaders = federationSSEHeaders("c", dstKeyByte, false)
}
after := assertFederationCopyReadback(t, objectAPI, apiRouter, credentials, remoteBucket, dstObject,
rec, wantChecksum, body.data, body.ext == ".txt" && pair.dst != "c", getHeaders)
assertFederationKMSContext(t, after, headers[xhttp.AmzServerSideEncryptionKmsContext])
// A second encryption layer would add its own DARE overhead.
if got, want := federationStoredSSE(after.UserDefined), strings.Split(pair.dst, "-")[0]; got != want {
t.Fatalf("%s: destination stored as %s, want %s", instanceType, got, want)
}
if pair.dst != "plain" && !after.IsCompressed() {
once := ObjectInfo{Size: int64(len(body.data))}
if want := once.EncryptedSize(); after.Size != want {
t.Fatalf("%s: destination stored %d bytes, want %d for %d bytes encrypted once",
instanceType, after.Size, want, len(body.data))
}
}
// The source must be untouched.
source, err := objectAPI.GetObjectInfo(t.Context(), bucketName, srcObject, ObjectOptions{})
if err != nil {
t.Fatalf("%s: GetObjectInfo(source) after the copy failed: %v", instanceType, err)
}
if source.Size != before.Size || source.ETag != before.ETag || !source.ModTime.Equal(before.ModTime) {
t.Fatalf("%s: federated copy modified its source: size %d->%d etag %s->%s",
instanceType, before.Size, source.Size, before.ETag, source.ETag)
}
})
}
}
// Multipart sources cross an encryption boundary per part. SHA256 also
// exercises promotion from a stored composite to a full-object checksum.
parts := [][]byte{bytes.Repeat([]byte("p"), globalMinPartSize+1), []byte("q!")}
data := bytes.Join(parts, nil)
for _, pair := range []struct {
src, dst string
composite bool
}{
{src: "s3", dst: "plain"},
{src: "s3", dst: "s3"},
{src: "c", dst: "c"},
{src: "kms", dst: "kms", composite: true},
} {
t.Run("multipart/"+pair.src+"-to-"+pair.dst, func(t *testing.T) {
srcObject, dstObject := "federation/multipart-source-"+pair.src+"-"+pair.dst+".bin", "federation/multipart-destination-"+pair.src+"-"+pair.dst+".bin"
sourceHeaders := federationSSEHeaders(pair.src, srcKeyByte, false)
wantChecksum := hash.ChecksumCRC64NVME
if pair.composite {
wantChecksum = hash.ChecksumSHA256
sourceHeaders[xhttp.AmzChecksumAlgo] = wantChecksum.String()
sourceHeaders[xhttp.AmzChecksumType] = xhttp.AmzChecksumTypeComposite
}
putFederationMultipartSource(t, apiRouter, credentials, bucketName, srcObject, parts, sourceHeaders)
before, err := objectAPI.GetObjectInfo(t.Context(), bucketName, srcObject, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if len(before.Parts) != len(parts) || federationStoredSSE(before.UserDefined) != pair.src {
t.Fatalf("source has %d parts stored as %s, want %d parts as %s", len(before.Parts), federationStoredSSE(before.UserDefined), len(parts), pair.src)
}
if pair.composite {
checksums, _ := before.decryptChecksums(0, nil)
if checksums[xhttp.AmzChecksumType] != xhttp.AmzChecksumTypeComposite || !strings.HasSuffix(checksums[wantChecksum.String()], "-2") {
t.Fatalf("source did not persist a two-part composite checksum: %v", checksums)
}
}
headers := federationSSEHeaders(pair.dst, dstKeyByte, false)
maps.Copy(headers, federationSSEHeaders(pair.src, srcKeyByte, true))
rec := federatedCopyRequest(t, apiRouter, credentials, bucketName, srcObject, remoteBucket, dstObject, headers)
if rec.Code != http.StatusOK {
t.Fatalf("%s: federated CopyObject failed: %d %s", instanceType, rec.Code, rec.Body.String())
}
var getHeaders map[string]string
if pair.dst == "c" {
getHeaders = federationSSEHeaders("c", dstKeyByte, false)
}
after := assertFederationCopyReadback(t, objectAPI, apiRouter, credentials, remoteBucket, dstObject, rec, wantChecksum, data, false, getHeaders)
if got := federationStoredSSE(after.UserDefined); got != pair.dst {
t.Fatalf("destination stored as %s, want %s", got, pair.dst)
}
if once := (ObjectInfo{Size: int64(len(data))}); pair.dst != "plain" && after.Size != once.EncryptedSize() {
t.Fatalf("destination stored %d bytes, want %d for %d bytes encrypted once", after.Size, once.EncryptedSize(), len(data))
}
})
}
}
+8 -4
View File
@@ -404,16 +404,16 @@ func testAPICopyObjectSSECKeyRotationNullVersion(obj ObjectLayer, instanceType,
apiRouter, credentials, false, t)
}
func TestAPICopyObjectSSECKeyRotationNullVersionCompressesRewrite(t *testing.T) {
func TestAPICopyObjectSSECKeyRotationNullVersionSkipsCompression(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPICopyObjectSSECKeyRotationNullVersionCompressesRewrite,
objAPITest: testAPICopyObjectSSECKeyRotationNullVersionSkipsCompression,
endpoints: []string{"CopyObject", "PutObject", "GetObject"},
})
}
func testAPICopyObjectSSECKeyRotationNullVersionCompressesRewrite(obj ObjectLayer, instanceType, bucketName string,
func testAPICopyObjectSSECKeyRotationNullVersionSkipsCompression(obj ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
testAPICopyObjectSSECKeyRotationNullVersionWithCompression(obj, instanceType, bucketName,
@@ -496,7 +496,11 @@ func testAPICopyObjectSSECKeyRotationNullVersionWithCompression(obj ObjectLayer,
for key, value := range getHeaders {
decryptHeaders.Set(key, value)
}
after := assertCopyChecksum(t, obj, bucketName, object, hash.ChecksumCRC32, data, compressAtCopy, decryptHeaders)
// This SSE-C rewrite stays uncompressed even with compression enabled at copy
// time. The plaintext sibling
// TestAPICopyObjectMetadataOnlyNullVersionCompressesRewrite keeps the
// coverage that a compressed rewrite records matching metadata.
after := assertCopyChecksum(t, obj, bucketName, object, hash.ChecksumCRC32, data, false, decryptHeaders)
if after.VersionID == "" {
t.Fatalf("%s: rotation into a versioned bucket did not create a new version", instanceType)
}
+243
View File
@@ -0,0 +1,243 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-only
package cmd
import (
"bytes"
"encoding/xml"
"fmt"
"io"
"maps"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
miniogo "github.com/minio/minio-go/v7"
"github.com/minio/minio-go/v7/pkg/credentials"
"github.com/minio/minio/internal/amztime"
"github.com/minio/minio/internal/auth"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// Exercise the selected SDK's actual retry loop and concurrent calls sharing a
// transport. Timestamps must follow the final PUT response, never a location
// probe, another operation, a failed attempt, or the HTTP Date header.
func TestFederatedWriteTimeTransport(t *testing.T) {
var mu sync.Mutex
attempts := make(map[string]int)
written := time.Date(2001, 2, 3, 4, 5, 6, 123456789, time.UTC)
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = io.Copy(io.Discard, r.Body)
w.Header().Set(federatedLastModified, written.Add(-time.Hour).Format(time.RFC3339Nano))
if r.Method == http.MethodGet && r.URL.Query().Has("location") {
_, _ = io.WriteString(w, `<LocationConstraint>us-east-1</LocationConstraint>`)
return
}
if r.Method != http.MethodPut {
t.Errorf("unexpected follow-up read: %s %s", r.Method, r.URL)
w.WriteHeader(http.StatusForbidden)
return
}
mu.Lock()
attempts[r.URL.Path]++
attempt := attempts[r.URL.Path]
mu.Unlock()
if strings.Contains(r.URL.Path, "retry-") && attempt == 1 {
w.WriteHeader(http.StatusServiceUnavailable)
_, _ = io.WriteString(w, `<Error><Code>SlowDown</Code></Error>`)
return
}
stamp := written.Add(time.Duration(len(r.URL.Path)) * time.Nanosecond)
if r.URL.Query().Has("partNumber") {
stamp = stamp.Add(time.Second)
}
w.Header().Set(federatedLastModified, stamp.Format(time.RFC3339Nano))
switch {
case strings.HasSuffix(r.URL.Path, "missing"):
w.Header().Del(federatedLastModified)
case strings.HasSuffix(r.URL.Path, "malformed"):
w.Header().Set(federatedLastModified, "invalid-time")
}
w.Header().Set("ETag", `"`+r.URL.Path+`"`)
}))
t.Cleanup(server.Close)
core, err := miniogo.NewCore(server.Listener.Addr().String(), &miniogo.Options{
Creds: credentials.NewStaticV4("federation", "federation-secret", ""),
Transport: federatedWriteTransport{server.Client().Transport},
BucketLookup: miniogo.BucketLookupPath,
MaxRetries: 2,
})
if err != nil {
t.Fatal(err)
}
for _, part := range []bool{false, true} {
for _, scenario := range []string{"present", "missing", "malformed", "retry-present", "retry-missing", "retry-malformed"} {
t.Run(fmt.Sprintf("part=%v/%s", part, scenario), func(t *testing.T) {
t.Parallel()
ctx, modified := withFederatedWriteTime(t.Context())
object := fmt.Sprintf("part-%v/%s", part, scenario)
var etag string
if part {
info, err := core.PutObjectPart(ctx, "bucket", object, "upload-id", 1, bytes.NewReader([]byte("abc")), 3, miniogo.PutObjectPartOptions{})
if err != nil {
t.Fatal(err)
}
etag = info.ETag
} else {
info, err := core.PutObject(ctx, "bucket", object, bytes.NewReader([]byte("abc")), 3, "", "", miniogo.PutObjectOptions{})
if err != nil {
t.Fatal(err)
}
etag = info.ETag
}
want := time.Time{}
if strings.HasSuffix(scenario, "present") {
want = written.Add(time.Duration(len("/bucket/"+object)) * time.Nanosecond)
if part {
want = want.Add(time.Second)
}
}
if !modified.Equal(want) || etag != "/bucket/"+object {
t.Errorf("write result: time=%s ETag=%s, want time=%s ETag=/bucket/%s", modified, etag, want, object)
}
mu.Lock()
count := attempts["/bucket/"+object]
mu.Unlock()
wantAttempts := 1
if strings.HasPrefix(scenario, "retry-") {
wantAttempts = 2
}
if count != wantAttempts {
t.Errorf("PUT attempts = %d, want %d", count, wantAttempts)
}
})
}
}
}
func TestAPIFederatedCopyWriteTimeSSEAndCompatibility(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
endpoints: []string{"CopyObjectPart", "NewMultipart", "PutObjectPart", "ListObjectParts", "CompleteMultipart", "CopyObject", "PutObject", "GetObject"},
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
testKMS, err := kms.NewBuiltin(federationTestKMSKeyID, bytes.Repeat([]byte{0x58}, 32))
if err != nil {
t.Fatal(err)
}
previous := GlobalKMS
GlobalKMS = testKMS
defer func() { GlobalKMS = previous }()
for _, mode := range []string{"current", "missing", "malformed"} {
t.Run(mode, func(t *testing.T) {
remote, _, cleanup := setupCopyObjectFederationRemote(t, obj, router, instanceType, bucket, true, func(h http.Header) {
switch mode {
case "missing":
h.Del(federatedLastModified)
case "malformed":
h.Set(federatedLastModified, "invalid-time")
}
})
defer cleanup()
for _, kind := range []string{"plain", "s3", "c", "kms"} {
if mode != "current" && kind != "plain" && kind != "c" {
continue
}
t.Run(kind, func(t *testing.T) {
data := []byte("federated write timestamp")
putCopyChecksumSource(t, router, cred, bucket, "source", data, federationSSEHeaders(kind, 0x11, false))
headers := federationSSEHeaders(kind, 0x22, false)
maps.Copy(headers, federationSSEHeaders(kind, 0x11, true))
rec := federatedCopyRequest(t, router, cred, bucket, "source", remote, kind+"-copy", headers)
if rec.Code != http.StatusOK {
t.Fatalf("CopyObject: %d %s", rec.Code, rec.Body.String())
}
info, err := obj.GetObjectInfo(t.Context(), remote, kind+"-copy", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
var copied CopyObjectResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &copied); err != nil {
t.Fatal(err)
}
assertTime := func(got string, stored time.Time) {
t.Helper()
if stored.IsZero() {
t.Fatal("storage returned a zero timestamp")
}
if mode != "current" {
stored = time.Time{}
}
if got != amztime.ISO8601Format(stored.UTC()) {
t.Errorf("copy time = %q, want %q", got, amztime.ISO8601Format(stored.UTC()))
}
}
assertTime(copied.LastModified, info.ModTime)
object := kind + "-multipart"
req, err := newTestSignedRequestV4(http.MethodPost, getNewMultipartURL("", remote, object), 0, nil,
cred.AccessKey, cred.SecretKey, federationSSEHeaders(kind, 0x22, false))
if err != nil {
t.Fatal(err)
}
rec = httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("NewMultipartUpload: %d %s", rec.Code, rec.Body.String())
}
var upload InitiateMultipartUploadResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &upload); err != nil {
t.Fatal(err)
}
// UploadPartCopy selects KMS/S3 encryption from the upload;
// only SSE-C keys accompany the part request. The replica
// markers must still take the ordinary decrypting copy path.
headers = federationSSEHeaders(kind, 0x11, true)
var keys map[string]string
if kind == "c" {
keys = federationSSEHeaders("c", 0x22, false)
maps.Copy(headers, keys)
}
headers[xhttp.AmzCopySource] = SlashSeparator + pathJoin(bucket, "source")
headers[xhttp.MinIOSourceReplicationRequest] = "true"
headers[xhttp.AmzBucketReplicationStatus] = "REPLICA"
req, err = newTestSignedRequestV4(http.MethodPut, getCopyObjectPartURL("", remote, object, upload.UploadID, "1"), 0, nil, cred.AccessKey, cred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
rec = httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("UploadPartCopy: %d %s", rec.Code, rec.Body.String())
}
var part CopyObjectPartResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &part); err != nil {
t.Fatal(err)
}
parts, err := obj.ListObjectParts(t.Context(), remote, object, upload.UploadID, 0, 100, ObjectOptions{})
if err != nil || len(parts.Parts) != 1 {
t.Fatalf("stored part: %v, %v", parts, err)
}
assertTime(part.LastModified, parts.Parts[0].LastModified)
rec = completePartsHTTP(t, router, cred, remote, object, upload.UploadID, []CompletePart{{PartNumber: 1, ETag: canonicalizeETag(part.ETag)}}, keys)
if rec.Code != http.StatusOK {
t.Fatalf("CompleteMultipartUpload: %d %s", rec.Code, rec.Body.String())
}
for _, name := range []string{kind + "-copy", object} {
got := federationGetObject(t, router, cred, remote, name, keys)
if got.Code != http.StatusOK || !bytes.Equal(got.Body.Bytes(), data) {
t.Fatalf("readback %s: %d %s", name, got.Code, got.Body.String())
}
}
})
}
})
}
},
})
}
@@ -0,0 +1,138 @@
// Copyright (c) 2015-2026 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"encoding/base64"
"encoding/binary"
"encoding/xml"
"hash/crc32"
"io"
"net/http"
"net/http/httptest"
"testing"
"github.com/minio/minio/internal/auth"
)
func crc32Checksum(data []byte) string {
var c [4]byte
binary.BigEndian.PutUint32(c[:], crc32.ChecksumIEEE(data))
return base64.StdEncoding.EncodeToString(c[:])
}
// TestAPIPutObjectChunkedChecksum exercises PutObject with aws-chunked streaming
// transfer encoding combined with a CRC32 checksum. It reproduces issue #107:
// the AWS Java SDK v2, with chunked encoding enabled, sends a non-trailer signed
// chunked body (x-amz-content-sha256: STREAMING-AWS4-HMAC-SHA256-PAYLOAD), places
// the precomputed checksum in the x-amz-checksum-crc32 header, yet still advertises
// it in x-amz-trailer even though no trailer is ever sent. Before the fix the
// server treated the checksum as trailing (empty value) and returned HTTP 400
// XAmzContentChecksumMismatch.
func TestAPIPutObjectChunkedChecksum(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: testAPIPutObjectChunkedChecksum,
endpoints: []string{"PutObject", "GetObject"},
})
}
func testAPIPutObjectChunkedChecksum(obj ObjectLayer, instanceType, bucketName string, apiRouter http.Handler,
credentials auth.Credentials, t *testing.T,
) {
data := bytes.Repeat([]byte("a"), 4096)
goodCRC := crc32Checksum(data)
// A well-formed 4-byte value that does not match the content.
wrongCRC := base64.StdEncoding.EncodeToString([]byte{0x01, 0x02, 0x03, 0x04})
apiCode := func(rec *httptest.ResponseRecorder) string {
var apiErr APIErrorResponse
b, _ := io.ReadAll(rec.Body)
_ = xml.Unmarshal(b, &apiErr)
return apiErr.Code
}
// newChunkedJavaForm builds a non-trailer signed chunked PutObject request that
// mirrors the AWS Java SDK v2 wire form: the checksum value sits in the header
// while x-amz-trailer still advertises it (no trailer is actually sent). The
// checksum headers are set BEFORE signing so they are part of the signed
// headers, exactly as the captured Java SDK request sends them.
newChunkedJavaForm := func(object, crc string) *http.Request {
body := bytes.NewReader(data)
req, err := newTestStreamingRequest(http.MethodPut,
getPutObjectURL("", bucketName, object),
int64(len(data)), int64(len(data)), body)
if err != nil {
t.Fatalf("Failed to create streaming request: %v", err)
}
req.Header.Set("x-amz-checksum-crc32", crc)
req.Header.Set("x-amz-trailer", "x-amz-checksum-crc32")
req.Header.Set("x-amz-sdk-checksum-algorithm", "CRC32")
currTime := UTCNow()
signature, err := signStreamingRequest(req, credentials.AccessKey, credentials.SecretKey, currTime)
if err != nil {
t.Fatalf("Failed to sign streaming request: %v", err)
}
req, err = assembleStreamingChunks(req, body, int64(len(data)), credentials.SecretKey, signature, currTime)
if err != nil {
t.Fatalf("Failed to assemble streaming chunks: %v", err)
}
return req
}
// 1. Correct checksum: must succeed and echo the checksum on the response.
{
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, newChunkedJavaForm("chunked-ok", goodCRC))
if rec.Code != http.StatusOK {
t.Fatalf("%s: chunked+CRC32 PutObject: expected 200, got %d (%s)", instanceType, rec.Code, apiCode(rec))
}
if got := rec.Header().Get("x-amz-checksum-crc32"); got != goodCRC {
t.Fatalf("%s: response checksum echo = %q, want %q", instanceType, got, goodCRC)
}
// Read the stored checksum back via GetObject with checksum mode enabled.
greq, err := newTestSignedRequestV4(http.MethodGet, getGetObjectURL("", bucketName, "chunked-ok"),
0, nil, credentials.AccessKey, credentials.SecretKey, map[string]string{"x-amz-checksum-mode": "ENABLED"})
if err != nil {
t.Fatalf("Failed to create GET request: %v", err)
}
grec := httptest.NewRecorder()
apiRouter.ServeHTTP(grec, greq)
if grec.Code != http.StatusOK {
t.Fatalf("%s: GetObject: expected 200, got %d", instanceType, grec.Code)
}
if got := grec.Header().Get("x-amz-checksum-crc32"); got != goodCRC {
t.Fatalf("%s: stored checksum read back = %q, want %q", instanceType, got, goodCRC)
}
}
// 2. Wrong checksum over the same chunked form: must still be rejected.
{
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, newChunkedJavaForm("chunked-wrong", wrongCRC))
if rec.Code != http.StatusBadRequest {
t.Fatalf("%s: chunked+wrong CRC32: expected 400, got %d", instanceType, rec.Code)
}
if code := apiCode(rec); code != "XAmzContentChecksumMismatch" {
t.Fatalf("%s: chunked+wrong CRC32: want XAmzContentChecksumMismatch, got %q", instanceType, code)
}
}
}
+70 -1
View File
@@ -28,6 +28,7 @@ import (
"github.com/minio/minio/internal/amztime"
"github.com/minio/minio/internal/bucket/lifecycle"
"github.com/minio/minio/internal/crypto"
"github.com/minio/minio/internal/event"
"github.com/minio/minio/internal/hash"
xhttp "github.com/minio/minio/internal/http"
@@ -35,6 +36,44 @@ import (
var etagRegex = regexp.MustCompile("\"*?([^\"]*?)\"*?$")
// federatedLastModified carries the committed object/part time on the write
// response. HTTP Date is a response time and Last-Modified loses subsecond
// precision, so neither can supply CopyObject/UploadPartCopy's timestamp.
const federatedLastModified = "X-Minio-Last-Modified"
type federatedWriteTimeKey struct{}
// withFederatedWriteTime allocates one result per SDK operation. SDK retries
// run sequentially with this context; concurrent copies never share the result.
func withFederatedWriteTime(ctx context.Context) (context.Context, *time.Time) {
modified := new(time.Time)
return context.WithValue(ctx, federatedWriteTimeKey{}, modified), modified
}
type federatedWriteTransport struct {
http.RoundTripper
}
func (t federatedWriteTransport) RoundTrip(r *http.Request) (*http.Response, error) {
resp, err := t.RoundTripper.RoundTrip(r)
if modified, ok := r.Context().Value(federatedWriteTimeKey{}).(*time.Time); ok && r.Method == http.MethodPut {
// Replace on every attempt, including a success from an older target
// without the header. Never retain a failed attempt's timestamp or fail
// a committed write just because its timestamp is unavailable.
*modified = time.Time{}
if err == nil && resp != nil && resp.StatusCode == http.StatusOK {
*modified, _ = time.Parse(time.RFC3339Nano, resp.Header.Get(federatedLastModified))
}
}
return resp, err
}
func setFederatedWriteTime(w http.ResponseWriter, r *http.Request, modified time.Time) {
if isFederatedInternalRequest(r.UserAgent()) && !modified.IsZero() {
w.Header().Set(federatedLastModified, modified.UTC().Format(time.RFC3339Nano))
}
}
// Validates the preconditions for CopyObjectPart, returns true if CopyObjectPart
// operation should not proceed. Preconditions supported are:
//
@@ -193,7 +232,15 @@ func checkPreconditionsPUT(ctx context.Context, w http.ResponseWriter, r *http.R
etagMatch := opts.PreserveETag != "" && isETagEqual(objInfo.ETag, opts.PreserveETag)
vidMatch := opts.VersionID != "" && opts.VersionID == objInfo.VersionID
if etagMatch && vidMatch {
// A matching version and ETag normally mean the destination already holds
// this version, so the write is skipped. They do not establish that for an
// authenticated SSE-C replica write: the destination cannot decrypt or
// re-encrypt the body without the customer key, so it cannot verify the
// replica, and this retransmission is how such a replica is repaired or
// updated. The predicate is the incoming request's restored SSE-C metadata,
// not what the destination happens to hold.
ssecReplica := isReplicaTrusted(r.Context()) && crypto.SSEC.IsEncrypted(opts.UserDefined)
if etagMatch && vidMatch && !ssecReplica {
writeHeaders()
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(ErrPreconditionFailed), r.URL)
return true
@@ -340,6 +387,28 @@ func canonicalizeETag(etag string) string {
return etagRegex.ReplaceAllString(etag, "$1")
}
// deleteIfMatchPreconditionFailed reports whether the If-Match precondition on a
// DeleteObject request fails, in which case the delete must be refused with 412.
// It is pure and never writes to the ResponseWriter: DeleteObject may evaluate
// the precondition off the request goroutine (e.g. during multi-pool cleanup).
//
// - A delete marker (a non-live latest version) has no entity-tag to match, so
// any If-Match value, including "*", fails against it.
// - "*" matches any existing live object, so it only requires existence.
// - A concrete ETag is compared against the object's public ETag. For
// SSE-C/SSE-KMS objects the public ETag is derived from the stored suffix
// without the customer key (getDecryptedETag), so a satisfiable condition is
// never rejected merely because the caller did not supply the key.
func deleteIfMatchPreconditionFailed(h http.Header, ifMatch string, oi ObjectInfo) bool {
if oi.DeleteMarker {
return true
}
if strings.TrimSpace(ifMatch) == "*" {
return false
}
return !isETagEqual(getDecryptedETag(h, oi, false), ifMatch)
}
// isETagEqual return true if the canonical representations of two ETag strings
// are equal, false otherwise
func isETagEqual(left, right string) bool {
@@ -0,0 +1,146 @@
// Copyright (c) 2015-2025 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"encoding/xml"
"net/http"
"net/http/httptest"
"testing"
"github.com/dustin/go-humanize"
"github.com/minio/minio/internal/auth"
xhttp "github.com/minio/minio/internal/http"
)
// TestAPIDeleteObjectHandlerIfMatch verifies conditional DeleteObject behavior
// for the If-Match request header (AWS S3 conditional deletes):
// - a non-matching ETag must return 412 Precondition Failed and preserve the object,
// - a matching ETag (or "*") must delete the object and return 204,
// - a request without If-Match must be unaffected,
// - If-Match against a missing key must return 404 NoSuchKey (not the idempotent 204).
func TestAPIDeleteObjectHandlerIfMatch(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPIDeleteObjectHandlerIfMatch, endpoints: []string{"DeleteObject"}})
}
func testAPIDeleteObjectHandlerIfMatch(obj ObjectLayer, instanceType, bucketName string, apiRouter http.Handler,
credentials auth.Credentials, t *testing.T,
) {
// putObject (re)creates an object and returns its ETag.
putObject := func(object string) string {
data := generateBytesData(1 * humanize.MiByte)
oi, err := obj.PutObject(context.Background(), bucketName, object,
mustGetPutObjReader(t, bytes.NewReader(data), int64(len(data)), "", ""), ObjectOptions{})
if err != nil {
t.Fatalf("%s: failed to put object %q: %v", instanceType, object, err)
}
return oi.ETag
}
// exists reports whether the object is still present.
exists := func(object string) bool {
_, err := obj.GetObjectInfo(context.Background(), bucketName, object, ObjectOptions{})
return err == nil
}
// doDelete issues a signed DELETE, optionally with an If-Match header.
doDelete := func(object, ifMatch string) *httptest.ResponseRecorder {
var hdrs map[string]string
if ifMatch != "" {
hdrs = map[string]string{xhttp.IfMatch: ifMatch}
}
req, err := newTestSignedRequestV4(http.MethodDelete, getDeleteObjectURL("", bucketName, object),
0, nil, credentials.AccessKey, credentials.SecretKey, hdrs)
if err != nil {
t.Fatalf("%s: failed to create DELETE request: %v", instanceType, err)
}
rec := httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
return rec
}
// (1) If-Match with a wrong ETag: 412 Precondition Failed, object preserved.
t.Run("wrong-etag-412", func(t *testing.T) {
object := "precond-wrong-etag"
putObject(object)
rec := doDelete(object, `"non-matching-etag"`)
if rec.Code != http.StatusPreconditionFailed {
t.Fatalf("%s: expected %d, got %d (body: %s)", instanceType, http.StatusPreconditionFailed, rec.Code, rec.Body.String())
}
if !exists(object) {
t.Fatalf("%s: object must still exist after a failed conditional delete", instanceType)
}
})
// (2) If-Match with the correct ETag: 204 No Content, object removed.
t.Run("correct-etag-204", func(t *testing.T) {
object := "precond-correct-etag"
etag := putObject(object)
rec := doDelete(object, `"`+etag+`"`)
if rec.Code != http.StatusNoContent {
t.Fatalf("%s: expected %d, got %d (body: %s)", instanceType, http.StatusNoContent, rec.Code, rec.Body.String())
}
if exists(object) {
t.Fatalf("%s: object must be removed after a matching conditional delete", instanceType)
}
})
// (3) If-Match "*" on an existing object matches any ETag: 204, object removed.
t.Run("wildcard-204", func(t *testing.T) {
object := "precond-wildcard"
putObject(object)
rec := doDelete(object, "*")
if rec.Code != http.StatusNoContent {
t.Fatalf("%s: expected %d, got %d (body: %s)", instanceType, http.StatusNoContent, rec.Code, rec.Body.String())
}
if exists(object) {
t.Fatalf("%s: object must be removed after a wildcard conditional delete", instanceType)
}
})
// (4) No If-Match header: unchanged behavior, 204, object removed.
t.Run("no-ifmatch-204", func(t *testing.T) {
object := "precond-none"
putObject(object)
rec := doDelete(object, "")
if rec.Code != http.StatusNoContent {
t.Fatalf("%s: expected %d, got %d (body: %s)", instanceType, http.StatusNoContent, rec.Code, rec.Body.String())
}
if exists(object) {
t.Fatalf("%s: object must be removed after an unconditional delete", instanceType)
}
})
// (5) If-Match against a non-existent key: 404 NoSuchKey, not the idempotent 204.
t.Run("missing-key-404", func(t *testing.T) {
rec := doDelete("precond-missing-key", `"some-etag"`)
if rec.Code != http.StatusNotFound {
t.Fatalf("%s: expected %d, got %d (body: %s)", instanceType, http.StatusNotFound, rec.Code, rec.Body.String())
}
var apiErr APIErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &apiErr); err != nil {
t.Fatalf("%s: failed to parse error response: %v", instanceType, err)
}
if apiErr.Code != "NoSuchKey" {
t.Fatalf("%s: expected error code NoSuchKey, got %q", instanceType, apiErr.Code)
}
})
}
@@ -0,0 +1,64 @@
// Copyright (c) 2015-2025 MinIO, Inc.
//
// This file is part of MinIO Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"net/http"
"testing"
"github.com/minio/minio/internal/crypto"
)
// TestDeleteIfMatchPreconditionFailed exercises the pure If-Match evaluation for
// DeleteObject, including delete markers, the "*" wildcard, and SSE-C objects
// whose public ETag must be derived without the customer key.
func TestDeleteIfMatchPreconditionFailed(t *testing.T) {
const plainETag = "d41d8cd98f00b204e9800998ecf8427e" // 32-char MD5, returned as-is
// SSE-C object: the stored ETag is longer than 32 chars and its public form
// is the trailing 32 chars, derived by getDecryptedETag without any key.
ssecPublic := "11112222333344445555666677778888"
ssecStored := "abcdef0123456789abcdef0123456789" + ssecPublic // 64 chars
ssecMeta := map[string]string{crypto.MetaSealedKeySSEC: "test-sealed-key"}
testCases := []struct {
name string
ifMatch string
oi ObjectInfo
want bool // true => precondition failed (delete refused, 412)
}{
{"live matching etag", plainETag, ObjectInfo{ETag: plainETag}, false},
{"live matching quoted etag", `"` + plainETag + `"`, ObjectInfo{ETag: plainETag}, false},
{"live non-matching etag", "0badbeef0badbeef0badbeef0badbeef", ObjectInfo{ETag: plainETag}, true},
{"live wildcard", "*", ObjectInfo{ETag: plainETag}, false},
{"live wildcard padded", " * ", ObjectInfo{ETag: plainETag}, false},
{"delete marker concrete etag", plainETag, ObjectInfo{DeleteMarker: true}, true},
{"delete marker wildcard", "*", ObjectInfo{DeleteMarker: true}, true},
{"ssec matching public etag", ssecPublic, ObjectInfo{ETag: ssecStored, UserDefined: ssecMeta}, false},
{"ssec wildcard", "*", ObjectInfo{ETag: ssecStored, UserDefined: ssecMeta}, false},
{"ssec non-matching etag", "99998888777766665555444433332222", ObjectInfo{ETag: ssecStored, UserDefined: ssecMeta}, true},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
if got := deleteIfMatchPreconditionFailed(http.Header{}, tc.ifMatch, tc.oi); got != tc.want {
t.Errorf("deleteIfMatchPreconditionFailed(%q) = %v, want %v", tc.ifMatch, got, tc.want)
}
})
}
}
+341 -71
View File
@@ -62,6 +62,7 @@ import (
"github.com/minio/minio/internal/logger"
"github.com/minio/minio/internal/s3select"
"github.com/minio/mux"
"github.com/minio/sio"
"github.com/pgsty/silo-pkg/v3/policy"
)
@@ -575,6 +576,8 @@ func (api objectAPIHandlers) getObjectHandler(ctx context.Context, objectAPI Obj
return
}
globalAccessTracker.note(bucket, object)
// Notify object accessed via a GET request.
sendEvent(eventArgs{
EventName: event.ObjectAccessedGet,
@@ -687,6 +690,20 @@ func (api objectAPIHandlers) getObjectAttributesHandler(ctx context.Context, obj
objInfo.decryptPartsChecksums(r.Header)
if _, ok := opts.ObjectAttributes[xhttp.ObjectParts]; ok {
// Report each part's uploaded plaintext byte length. Parts are stored
// transformed, compressed and/or encrypted, so the stored size is not
// what AWS defines ObjectPart.Size to be.
_, isEncrypted := crypto.IsEncrypted(objInfo.UserDefined)
isCompressed := objInfo.IsCompressed()
// Only a part of an encrypted multipart object is a stream of its
// own. A legacy encrypted object without that marker is one
// continuous stream that the erasure writer split into storage
// fragments, so no fragment has a plaintext length to report and
// each keeps its stored size, as ObjectInfo.DecryptedSize and
// DecryptBlocksRequestR also treat it.
hasEncryptedParts := isEncrypted &&
(crypto.IsMultiPart(objInfo.UserDefined) || len(objInfo.Parts) == 1)
OA.ObjectParts = new(objectAttributesParts)
OA.ObjectParts.PartNumberMarker = opts.PartNumberMarker
@@ -701,9 +718,38 @@ func (api objectAPIHandlers) getObjectAttributesHandler(ctx context.Context, obj
}
if len(OA.ObjectParts.Parts) == opts.MaxParts {
// This page is full and at least one more part
// is still pending, so the listing is truncated.
OA.ObjectParts.IsTruncated = true
break
}
partSize := objInfo.Parts[i].Size
switch {
case isCompressed:
// ActualSize is recorded by the same code that compresses,
// so it is always present for a compressed part.
if objInfo.Parts[i].ActualSize >= 0 {
partSize = objInfo.Parts[i].ActualSize
}
case hasEncryptedParts:
// ActualSize cannot be trusted for encrypted parts: a
// replicated SSE-C part records the ciphertext length, and
// parts written before actualSize existed record 0. Derive
// the plaintext length from the ciphertext instead, exactly
// as ObjectInfo.DecryptedSize does. A part whose stored
// length is not a valid encrypted stream has no logical
// length to report, and DecryptObjectInfo above only
// validates the object as a whole in that case, because
// ObjectInfo.isMultipart gives up on the first bad part.
decrypted, err := sio.DecryptedSize(uint64(partSize))
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, errObjectTampered), r.URL)
return
}
partSize = int64(decrypted)
}
OA.ObjectParts.NextPartNumberMarker = v.Number
OA.ObjectParts.Parts = append(OA.ObjectParts.Parts, &objectAttributesPart{
ChecksumSHA1: objInfo.Parts[i].Checksums["SHA1"],
@@ -712,13 +758,16 @@ func (api objectAPIHandlers) getObjectAttributesHandler(ctx context.Context, obj
ChecksumCRC32C: objInfo.Parts[i].Checksums["CRC32C"],
ChecksumCRC64NVME: objInfo.Parts[i].Checksums["CRC64NVME"],
PartNumber: objInfo.Parts[i].Number,
Size: objInfo.Parts[i].Size,
Size: partSize,
})
}
}
if OA.ObjectParts.NextPartNumberMarker != partsLength {
OA.ObjectParts.IsTruncated = true
// Part numbers may be sparse, so they cannot be compared against
// the part count. NextPartNumberMarker only carries a continuation
// token for a truncated listing, as in ListObjectParts.
if !OA.ObjectParts.IsTruncated {
OA.ObjectParts.NextPartNumberMarker = 0
}
}
@@ -1156,12 +1205,20 @@ const federatedInternalAppName = "minio-federated"
// Applicable only in a federated deployment
var getRemoteInstanceClient = func(r *http.Request, host string) (*miniogo.Core, error) {
cred := getReqAccessCred(r, globalSite.Region())
transport := getRemoteInstanceTransport()
if transport == nil {
var err error
transport, err = miniogo.DefaultTransport(globalIsTLS)
if err != nil {
return nil, err
}
}
// In a federated deployment, all the instances share config files
// and hence expected to have same credentials.
core, err := miniogo.NewCore(host, &miniogo.Options{
Creds: credentials.NewStaticV4(cred.AccessKey, cred.SecretKey, ""),
Secure: globalIsTLS,
Transport: getRemoteInstanceTransport(),
Transport: federatedWriteTransport{transport},
})
if err != nil {
return nil, err
@@ -1170,6 +1227,43 @@ var getRemoteInstanceClient = func(r *http.Request, host string) (*miniogo.Core,
return core, nil
}
// federatedChecksumType maps a requested server-side checksum type to the
// minio-go checksum the federation proxy asks the remote deployment to compute
// on the forwarded PutObject. Returns ChecksumNone for an unset/unknown type.
func federatedChecksumType(t hash.ChecksumType) miniogo.ChecksumType {
switch t.Base() {
case hash.ChecksumCRC32:
return miniogo.ChecksumCRC32
case hash.ChecksumCRC32C:
return miniogo.ChecksumCRC32C
case hash.ChecksumSHA1:
return miniogo.ChecksumSHA1
case hash.ChecksumSHA256:
return miniogo.ChecksumSHA256
case hash.ChecksumCRC64NVME:
return miniogo.ChecksumCRC64NVME
}
return miniogo.ChecksumNone
}
// federatedChecksumValue returns the base64 checksum the remote deployment
// reported for the requested type on the forwarded PutObject.
func federatedChecksumValue(t hash.ChecksumType, info miniogo.UploadInfo) string {
switch t.Base() {
case hash.ChecksumCRC32:
return info.ChecksumCRC32
case hash.ChecksumCRC32C:
return info.ChecksumCRC32C
case hash.ChecksumSHA1:
return info.ChecksumSHA1
case hash.ChecksumSHA256:
return info.ChecksumSHA256
case hash.ChecksumCRC64NVME:
return info.ChecksumCRC64NVME
}
return ""
}
// Check if the destination bucket is on a remote site, this code only gets executed
// when federation is enabled, ie when globalDNSConfig is non 'nil'.
//
@@ -1318,11 +1412,27 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
}
allowReplicationMetadata := replicaTrusted
// Check if bucket encryption is enabled
sseConfig, _ := globalBucketSSEConfigSys.Get(dstBucket)
sseConfig.Apply(r.Header, sse.ApplyOptions{
AutoEncrypt: globalAutoEncryption,
})
// Federation only: the destination bucket lives on another deployment and
// the copy is forwarded to it as a PutObject. That remote write owns the
// destination's storage transformations, so this handler hands it the
// logical (decompressed, decrypted) bytes at their logical size and lets
// the remote compress and encrypt once. Encrypting here as well would
// forward ciphertext under the destination's own SSE option, which either
// fails the length check or, for SSE to SSE, has the remote encrypt the
// ciphertext a second time and store an unreadable object (#158).
remoteCallRequired := isRemoteCopyRequired(ctx, srcBucket, dstBucket, objectAPI)
// Apply the destination bucket's default encryption only when this
// deployment writes the destination. For a remote destination the proxy
// has no authority over that bucket's defaults: applying its own here
// would forward an explicit SSE header that the remote then honors in
// place of the destination's configuration (#167).
if !remoteCallRequired {
sseConfig, _ := globalBucketSSEConfigSys.Get(dstBucket)
sseConfig.Apply(r.Header, sse.ApplyOptions{
AutoEncrypt: globalAutoEncryption,
})
}
var srcOpts, dstOpts ObjectOptions
srcOpts, err = copySrcOpts(ctx, r, srcBucket, srcObject)
if err != nil {
@@ -1387,6 +1497,12 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
}
defer gr.Close()
srcInfo := gr.ObjInfo
// A trusted SSE-C replica reads raw ciphertext. Federation's ordinary PUT
// expects plaintext and does not forward the replica's sealed-key metadata.
if remoteCallRequired && replicaTrusted && crypto.SSEC.IsEncrypted(srcInfo.UserDefined) {
writeErrorResponse(ctx, w, toAPIError(ctx, NotImplemented{Message: "federated raw SSE-C replica CopyObject is not supported"}), r.URL)
return
}
// maximum Upload size for object in a single CopyObject operation.
if isMaxObjectSize(srcInfo.Size) {
@@ -1437,7 +1553,7 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
// Pass the decompressed stream to such calls.
isDstCompressed := isCompressible(r.Header, dstObject) &&
length > minCompressibleSize &&
!isRemoteCopyRequired(ctx, srcBucket, dstBucket, objectAPI)
!remoteCallRequired
if isDstCompressed {
compressMetadata = make(map[string]string, 2)
// Preserving the compression metadata.
@@ -1512,8 +1628,19 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
// data instead, the rotation has to go through the regular re-encrypting
// copy, or the destination ends up holding plaintext under metadata that
// claims the object is encrypted.
//
// A checksum algorithm the client asks for has to be computed over the object
// data, which an in-place rotation never reads, so leave the fast path and
// let the re-encrypting copy compute, store and report it. A replica-trusted
// request is not such a client: getOpts.ReplicationRequest leaves its source
// reader encrypted, so a rewrite would hash ciphertext, and a replica has to
// keep the checksum its source assigned.
// Without a requested destination checksum algorithm, an in-place SSE-C
// rotation preserves the stored checksum state, including absence; it does
// not add the default CRC64NVME.
canRotateKeyInPlace := !srcInfo.Legacy &&
!copyRewritesObjectData(srcInfo.metadataOnly, copySrcOpts, dstOpts)
!copyRewritesObjectData(srcInfo.metadataOnly, copySrcOpts, dstOpts) &&
(replicaTrusted || !hash.NewChecksumHeader(r.Header).IsSet())
// If src == dst and either
// - the object is encrypted using SSE-C and two different SSE-C keys are present
@@ -1555,6 +1682,9 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
var targetSize int64
switch {
case remoteCallRequired:
// The remote receives the logical bytes, see above.
targetSize = actualSize
case isDstCompressed:
targetSize = -1
case !isSourceEncrypted && !isTargetEncrypted:
@@ -1625,7 +1755,7 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
pReader.setChecksumReader(checksumReader)
}
if isTargetEncrypted {
if isTargetEncrypted && !remoteCallRequired {
var encReader io.Reader
kind, _ := crypto.IsRequested(r.Header)
encReader, objEncKey, err = newEncryptReader(ctx, srcInfo.Reader, kind, newKeyID, newKey, dstBucket, dstObject, encMetadata, kmsCtx)
@@ -1649,7 +1779,7 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
return
}
if isTargetEncrypted {
if isTargetEncrypted && !remoteCallRequired {
pReader, err = pReader.WithEncryption(srcInfo.Reader, &objEncKey)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
@@ -1712,35 +1842,28 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
// apply default bucket configuration/governance headers for dest side.
retentionMode, retentionDate, legalHold, s3Err := checkPutObjectLockAllowed(ctx, r, dstBucket, dstObject, getObjectInfo, retPerms, holdPerms, replicaTrusted)
if s3Err == ErrNone && retentionMode.Valid() {
if dstOpts.ReplicationRequest {
srcTimestamp := dstOpts.ReplicationSourceRetentionTimestamp
if storedLock.retentionIsOlderThan(srcTimestamp) {
srcInfo.UserDefined[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
srcInfo.UserDefined[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
srcInfo.UserDefined[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = srcTimestamp.UTC().Format(time.RFC3339Nano)
} else {
storedLock.restoreRetention(srcInfo.UserDefined)
}
} else {
srcInfo.UserDefined[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
srcInfo.UserDefined[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
srcInfo.UserDefined[ReservedMetadataPrefixLower+ObjectLockRetentionTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
}
if s3Err == ErrNone {
applyReplicatedObjectLock(srcInfo.UserDefined, storedLock, replicaTrusted,
retentionMode, retentionDate, legalHold,
dstOpts.ReplicationSourceRetentionTimestamp, dstOpts.ReplicationSourceLegalholdTimestamp)
dstOpts.ReplicaLockReconcile = replicaTrusted && dstOpts.VersionID != ""
if s3Err == ErrNone && legalHold.Status.Valid() {
if dstOpts.ReplicationRequest {
srcTimestamp := dstOpts.ReplicationSourceLegalholdTimestamp
if storedLock.legalHoldIsOlderThan(srcTimestamp) {
srcInfo.UserDefined[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
srcInfo.UserDefined[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = srcTimestamp.UTC().Format(time.RFC3339Nano)
} else {
storedLock.restoreLegalHold(srcInfo.UserDefined)
if replicaTrusted {
// An SSE-C key rotation snapshots every stored reserved key into
// encMetadata above, before this decision exists, and the merge that
// preserves the encryption headers would put the stored ordering
// timestamps back over it. For a trusted replica the decision just
// made is authoritative, so let the snapshot agree with it.
for _, key := range []string{
ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp,
ReservedMetadataPrefixLower + ObjectLockLegalHoldTimestamp,
} {
if value, ok := srcInfo.UserDefined[key]; ok {
encMetadata[key] = value
} else {
delete(encMetadata, key)
}
}
} else {
srcInfo.UserDefined[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
srcInfo.UserDefined[ReservedMetadataPrefixLower+ObjectLockLegalHoldTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
}
if s3Err != ErrNone {
@@ -1810,9 +1933,6 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
srcInfo.metadataOnly = false
}
// Federation only.
remoteCallRequired := isRemoteCopyRequired(ctx, srcBucket, dstBucket, objectAPI)
var objInfo ObjectInfo
var os *objSweeper
if remoteCallRequired {
@@ -1834,23 +1954,109 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
// Remove the metadata for remote calls.
delete(srcInfo.UserDefined, ReservedMetadataPrefix+"compression")
delete(srcInfo.UserDefined, ReservedMetadataPrefix+"actual-size")
// A plain federated PutObject must not carry any internal storage
// metadata. The remote rejects every reserved-prefix header as a class
// (containsReservedMetadata), so strip the whole class here rather than
// an enumerated subset: an inline source object also carries
// inline-data, and replication bookkeeping adds still more, so removing
// only compression/actual-size just defers the next rejected key. Match
// the remote's case-insensitive detection.
for k := range srcInfo.UserDefined {
if stringsHasPrefixFold(k, ReservedMetadataPrefix) {
delete(srcInfo.UserDefined, k)
}
}
// Forward a clone: request-only keys are removed or added below and
// srcInfo.UserDefined stays the resolved record for the response.
forwardedMeta := cloneMSS(srcInfo.UserDefined)
// minio-go prefixes any UserMetadata key it does not recognize with
// x-amz-meta-, and it recognizes the retention headers but not
// x-amz-object-lock-legal-hold, so a hold left in the map reaches the
// remote as user metadata and is silently dropped (#166). Carry it on
// the typed option instead. Retention stays in the map: the typed
// RetainUntilDate would truncate the date to whole seconds.
legalHoldKey := strings.ToLower(xhttp.AmzObjectLockLegalHold)
forwardedLegalHold := forwardedMeta[legalHoldKey]
delete(forwardedMeta, legalHoldKey)
opts := miniogo.PutObjectOptions{
UserMetadata: srcInfo.UserDefined,
UserMetadata: forwardedMeta,
ServerSideEncryption: dstOpts.ServerSideEncryption,
UserTags: tag.ToMap(),
LegalHold: miniogo.LegalHoldStatus(forwardedLegalHold),
}
remoteObjInfo, rerr := core.PutObject(ctx, dstBucket, dstObject, srcInfo.Reader,
srcInfo.Size, "", "", opts)
// The destination must carry the same checksum the local path would
// produce; the federated path has the remote compute, validate, persist
// and return it. Without this the federated copy silently returns an
// empty checksum (#99).
wantChecksumType := dstOpts.WantServerSideChecksumType
var checksumHeaderValue string
switch {
case wantChecksumType.IsSet():
if actualSize == 0 {
// minio-go streams no trailing checksum for an empty body, so the
// remote would compute none, the bind below would fail, and every
// empty-object federated copy would 500. Forward the empty-content
// digest as an ordinary checksum header instead, so the remote
// validates, persists and returns it (parity with the local path,
// e.g. CRC32 "AAAAAA==").
if empty := hash.NewChecksumFromData(wantChecksumType, nil); empty != nil {
checksumHeaderValue = empty.Encoded
}
} else {
// Stream a trailing checksum of the requested type so the remote
// computes and persists it over the copied bytes.
opts.Checksum = federatedChecksumType(wantChecksumType)
}
case dstOpts.WantChecksum != nil && dstOpts.WantChecksum.Type.IsSet():
// A full-object checksum inherited from a checksum-bearing source: the
// local path persists it without recomputation (only multipart
// composite sources are promoted to WantServerSideChecksumType above,
// so this is never composite). Forward the value as an ordinary
// checksum header so the remote validates and persists it, and bind
// the returned value below. WantChecksum.Encoded is always a plain
// digest, never a "-N" multipart form.
wantChecksumType = dstOpts.WantChecksum.Type.Base()
checksumHeaderValue = dstOpts.WantChecksum.Encoded
}
if checksumHeaderValue != "" {
opts.UserMetadata[wantChecksumType.Key()] = checksumHeaderValue
}
// For an ordinary federated copy, actualSize is the logical plaintext
// length to declare. srcInfo.Size is not a reliable wire length because
// the read path may already have adjusted it.
writeCtx, modified := withFederatedWriteTime(ctx)
remoteObjInfo, rerr := core.PutObject(writeCtx, dstBucket, dstObject, srcInfo.Reader,
actualSize, "", "", opts)
if rerr != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, rerr), r.URL)
return
}
objInfo.UserDefined = cloneMSS(opts.UserMetadata)
// Keep the resolved legal hold in response and event metadata; the
// request-only checksum header exists only in the forwarding map.
objInfo.UserDefined = cloneMSS(srcInfo.UserDefined)
// The response headers and the ObjectCreated:Copy event describe the
// object this handler wrote, so name it: without these the response
// carries no x-amz-version-id and the event has an empty key and zero
// size (#170). Size is the logical size the remote was handed, which is
// what it reports back.
objInfo.Bucket = dstBucket
objInfo.Name = dstObject
objInfo.Size = actualSize
objInfo.VersionID = remoteObjInfo.VersionID
objInfo.ETag = remoteObjInfo.ETag
objInfo.ModTime = remoteObjInfo.LastModified
objInfo.ModTime = *modified
// Bind the checksum the remote computed for this exact write. A single
// forwarded PutObject must yield a full-object digest, so reject a
// missing, malformed, or multipart-marked ("-N") value rather than
// mislabel the destination as composite.
if wantChecksumType.IsSet() {
cs := hash.NewChecksumWithType(wantChecksumType, federatedChecksumValue(wantChecksumType, remoteObjInfo))
if cs == nil || cs.Type.Is(hash.ChecksumMultipart) {
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(ErrInternalError), r.URL)
return
}
objInfo.Checksum = cs.AppendTo(nil, nil)
}
} else {
os = newObjSweeper(dstBucket, dstObject).WithVersioning(dstOpts.Versioned, dstOpts.VersionSuspended)
// Get appropriate object info to identify the remote object to delete
@@ -2087,11 +2293,18 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
delete(metadata, xhttp.AmzBucketReplicationStatus)
}
// Check if bucket encryption is enabled
sseConfig, _ := globalBucketSSEConfigSys.Get(bucket)
sseConfig.Apply(r.Header, sse.ApplyOptions{
AutoEncrypt: globalAutoEncryption,
})
// A validated raw SSE-C replica carries the source ciphertext and the
// source seal. Its bytes must be stored verbatim: default encryption would
// overwrite the SSE-C IV and compression would invalidate the ciphertext.
rawSSECReplica := isRawSSECReplica(r.Header, replicaTrusted)
if !rawSSECReplica {
// Check if bucket encryption is enabled
sseConfig, _ := globalBucketSSEConfigSys.Get(bucket)
sseConfig.Apply(r.Header, sse.ApplyOptions{
AutoEncrypt: globalAutoEncryption,
})
}
var reader io.Reader
reader = rd
@@ -2105,7 +2318,7 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
actualSize := size
var idxCb func() []byte
if isCompressible(r.Header, object) && size > minCompressibleSize {
if !rawSSECReplica && isCompressible(r.Header, object) && size > minCompressibleSize {
// Storing the compression metadata.
metadata[ReservedMetadataPrefix+"compression"] = compressionAlgorithmV2
metadata[ReservedMetadataPrefix+"actual-size"] = strconv.FormatInt(size, 10)
@@ -2170,9 +2383,21 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
r.Header.Get(xhttp.IfMatch) != "" ||
r.Header.Get(xhttp.IfNoneMatch) != "" {
opts.CheckPrecondFn = func(oi ObjectInfo) bool {
if _, err := DecryptObjectInfo(&oi, r); err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return true
// A pure raw SSE-C replica overwrite (no public precondition) fully
// replaces the stored object, so the destination must not first
// require the stored object to decrypt: a replica an older destination
// bug left as compress(ciphertext) or double-encrypted has an invalid
// decrypted length, and requiring it here blocks the retransmission
// that repairs it. The predicate is the incoming request's restored
// SSE-C metadata, the same one checkPreconditionsPUT uses to exempt the
// version/ETag duplicate. A conditional request (If-Match/If-None-Match)
// still needs the decrypted, client-visible ETag, so it keeps the check.
ssecReplica := isReplicaTrusted(ctx) && crypto.SSEC.IsEncrypted(opts.UserDefined)
if !ssecReplica || r.Header.Get(xhttp.IfMatch) != "" || r.Header.Get(xhttp.IfNoneMatch) != "" {
if _, err := DecryptObjectInfo(&oi, r); err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return true
}
}
return checkPreconditionsPUT(ctx, w, r, oi, opts)
}
@@ -2184,12 +2409,29 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
getObjectInfo := objectAPI.GetObjectInfo
retentionMode, retentionDate, legalHold, s3Err := checkPutObjectLockAllowed(ctx, r, bucket, object, getObjectInfo, retPerms, holdPerms, isReplicaTrusted(ctx))
if s3Err == ErrNone && retentionMode.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
}
if s3Err == ErrNone && legalHold.Status.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
if s3Err == ErrNone {
// A trusted replica write addressing a specific version can be a full
// retransmit over an existing version whose lock state is newer than the
// source snapshot (issue #120). Order the incoming update against what is
// stored so a stale value cannot overwrite it; a non-replica write (and a
// marker-only peer write) has no stored state to order against and takes
// the helper's ordinary-write branch.
var storedLock objectLockState
if isReplicaTrusted(ctx) && opts.VersionID != "" {
var lerr error
if storedLock, lerr = replicaStoredLock(ctx, getObjectInfo, bucket, object, opts.VersionID); lerr != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, lerr), r.URL)
return
}
}
applyReplicatedObjectLock(metadata, storedLock, isReplicaTrusted(ctx),
retentionMode, retentionDate, legalHold,
opts.ReplicationSourceRetentionTimestamp, opts.ReplicationSourceLegalholdTimestamp)
// The decision above orders against the version as read here; let the
// object layer re-run it against the version read under the write lock
// that guards the replacement, so a newer hold or retention committed in
// between is not rolled back, across all pools and encryption modes.
opts.ReplicaLockReconcile = isReplicaTrusted(ctx) && opts.VersionID != ""
}
if s3Err != ErrNone {
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(s3Err), r.URL)
@@ -2200,7 +2442,10 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
metadata[ReservedMetadataPrefixLower+ReplicationStatus] = dsc.PendingStatus()
}
var objectEncryptionKey crypto.ObjectKey
if crypto.Requested(r.Header) {
// rawSSECReplica also gates the branch itself, not just the default-SSE
// header that would normally reach it: a trusted peer that sends an
// explicit public SSE header alongside a source seal must not re-encrypt.
if !rawSSECReplica && crypto.Requested(r.Header) {
if crypto.SSECopy.IsRequested(r.Header) {
writeErrorResponse(ctx, w, toAPIError(ctx, errInvalidEncryptionParameters), r.URL)
return
@@ -2301,6 +2546,7 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
}
setPutObjHeaders(w, objInfo, false, r.Header)
setFederatedWriteTime(w, r, objInfo.ModTime)
// Notify object created event.
evt := eventArgs{
@@ -2846,6 +3092,20 @@ func (api objectAPIHandlers) DeleteObjectHandler(w http.ResponseWriter, r *http.
return
}
// If-Match conditional delete (see AWS S3 conditional deletes). Delete the
// object only if its current ETag matches the client-supplied value,
// otherwise the request is refused with 412 Precondition Failed and the
// object is left intact. The precondition is evaluated in
// erasureServerPools.DeleteObject, against the version that will actually be
// removed, while the delete lock is held, so the object cannot change
// between the ETag check and the delete.
if ifMatch := r.Header.Get(xhttp.IfMatch); ifMatch != "" {
opts.HasIfMatch = true
opts.CheckPrecondFn = func(oi ObjectInfo) bool {
return deleteIfMatchPreconditionFailed(r.Header, ifMatch, oi)
}
}
rcfg, _ := globalBucketObjectLockSys.Get(bucket)
if rcfg.LockEnabled && opts.DeletePrefix {
apiErr := toAPIError(ctx, errInvalidArgument)
@@ -2911,6 +3171,13 @@ func (api objectAPIHandlers) DeleteObjectHandler(w http.ResponseWriter, r *http.
return
}
if isErrObjectNotFound(err) || isErrVersionNotFound(err) {
if opts.HasIfMatch {
// A conditional (If-Match) delete cannot satisfy its
// precondition against a missing object, so surface the
// not-found error instead of the idempotent 204 response.
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
// Send an event when the object is not found
objInfo.Name = object
objInfo.VersionID = opts.VersionID
@@ -2957,7 +3224,7 @@ func (api objectAPIHandlers) DeleteObjectHandler(w http.ResponseWriter, r *http.
if objInfo.ReplicationStatus == replication.Pending || objInfo.VersionPurgeStatus == replication.VersionPurgePending {
dmVersionID := ""
versionID := ""
if objInfo.DeleteMarker {
if objInfo.DeleteMarker && objInfo.VersionPurgeStatus.Empty() {
dmVersionID = objInfo.VersionID
} else {
versionID = objInfo.VersionID
@@ -3432,9 +3699,6 @@ func (api objectAPIHandlers) PutObjectTaggingHandler(w http.ResponseWriter, r *h
}
tagsStr := tags.String()
// Set this such that authorization policies can be applied on the object tags.
r.Header.Set(xhttp.AmzObjectTagging, tagsStr)
logger.GetReqInfo(ctx).BucketName = bucket
logger.GetReqInfo(ctx).ObjectName = object
if s3Error := authenticateRequest(ctx, r, policy.PutObjectTaggingAction); s3Error != ErrNone {
@@ -3442,6 +3706,12 @@ func (api objectAPIHandlers) PutObjectTaggingHandler(w http.ResponseWriter, r *h
return
}
// Set this such that authorization policies can be applied on the object
// tags. This is derived from the request body, so it must be injected only
// after signature verification: the SigV4 verifier now rejects unsigned
// x-amz-* request headers, and this synthesized header is never signed.
r.Header.Set(xhttp.AmzObjectTagging, tagsStr)
opts, err := getOpts(ctx, r, bucket, object)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
+64
View File
@@ -0,0 +1,64 @@
package cmd
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
)
// tamperedObjectLayer returns a fixed error from the read paths so the handler
// response can be tested without reproducing the storage defect that produces
// an unreadable object. No object bytes are sent to the caller.
type tamperedObjectLayer struct {
ObjectLayer
err error
}
func (o *tamperedObjectLayer) GetObjectNInfo(context.Context, string, string, *HTTPRangeSpec, http.Header, ObjectOptions) (*GetObjectReader, error) {
return nil, o.err
}
func (o *tamperedObjectLayer) GetObjectInfo(context.Context, string, string, ObjectOptions) (ObjectInfo, error) {
return ObjectInfo{}, o.err
}
// TestObjectTamperedGETHEADStatus asserts that an object the server cannot
// decode is reported with a 5xx status, not a success status. Returning
// http.StatusPartialContent here let SDKs accept the XML error document as
// object content. See pgsty/silo#110.
func TestObjectTamperedGETHEADStatus(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, credentials auth.Credentials, t *testing.T) {
// An encrypted stream shorter than one complete encryption package is
// a real size-validation origin of errObjectTampered.
damaged := ObjectInfo{Size: 31, UserDefined: map[string]string{crypto.MetaAlgorithm: crypto.InsecureSealAlgorithm}}
_, err := damaged.DecryptedSize()
if err != errObjectTampered {
t.Fatalf("damaged-size error = %v, want errObjectTampered", err)
}
previous := newObjectLayerFn()
setObjectLayer(&tamperedObjectLayer{ObjectLayer: obj, err: err})
defer setObjectLayer(previous)
for _, method := range []string{http.MethodGet, http.MethodHead} {
req, err := newTestSignedRequestV4(method, getGetObjectURL("", bucket, "damaged-object"), 0, nil, credentials.AccessKey, credentials.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if method == http.MethodGet && !strings.Contains(rec.Body.String(), "<Code>XMinioObjectTampered</Code>") {
t.Fatalf("%s: GET did not reach the damaged-object response: %d %s", instanceType, rec.Code, rec.Body.String())
}
if method == http.MethodHead && rec.Header().Get(xMinIOErrCodeHeader) != "XMinioObjectTampered" {
t.Fatalf("%s: HEAD did not reach the damaged-object response: %d %v", instanceType, rec.Code, rec.Header())
}
if rec.Code != http.StatusInternalServerError {
t.Errorf("%s: %s of a damaged object returned HTTP %d, want %d", instanceType, method, rec.Code, http.StatusInternalServerError)
}
}
}})
}
+29 -1
View File
@@ -979,7 +979,10 @@ func testAPIGetObjectWithPartNumberHandler(obj ObjectLayer, instanceType, bucket
t.Fatalf("Object: %s Object Index %d: Unexpected err: %v", object, oindex, err)
}
rs := partNumberToRangeSpec(oinfo, partNumber)
rs, err := partNumberToRangeSpec(oinfo, partNumber)
if err != nil {
t.Fatalf("Object: %s Object Index %d: Unexpected err: %v", object, oindex, err)
}
size, err := oinfo.GetActualSize()
if err != nil {
t.Fatalf("Object: %s Object Index %d: Unexpected err: %v", object, oindex, err)
@@ -1796,6 +1799,12 @@ func testAPICopyObjectPartHandlerSanity(obj ObjectLayer, instanceType, bucketNam
req.Header.Set("X-Amz-Copy-Source", url.QueryEscape(pathJoin(bucketName, objectName)))
req.Header.Set("X-Amz-Copy-Source-Range", fmt.Sprintf("bytes=%d-%d", a, b))
// Re-sign so the copy-source x-amz-* headers are covered by the
// signature, as real clients do; the verifier rejects unsigned x-amz-*.
if err = signRequestV4(req, credentials.AccessKey, credentials.SecretKey); err != nil {
t.Fatalf("Test failed to re-sign HTTP request for copy object part: <ERROR> %v", err)
}
// Since `apiRouter` satisfies `http.Handler` it has a ServeHTTP to execute the logic of the handler.
// Call the ServeHTTP to execute the handler, `func (api objectAPIHandlers) CopyObjectHandler` handles the request.
a = globalMinPartSize + 1
@@ -2197,6 +2206,15 @@ func testAPICopyObjectPartHandler(obj ObjectLayer, instanceType, bucketName stri
}
}
// Re-sign so the copy-source x-amz-* headers set above are covered by
// the signature, as real clients do; the verifier rejects unsigned
// x-amz-* headers.
if testCase.accessKey != "" && testCase.secretKey != "" {
if err = signRequestV4(req, testCase.accessKey, testCase.secretKey); err != nil {
t.Fatalf("Test %d: Failed to re-sign HTTP request for copy Object: <ERROR> %v", i+1, err)
}
}
// Since `apiRouter` satisfies `http.Handler` it has a ServeHTTP to execute the logic of the handler.
// Call the ServeHTTP to execute the handler, `func (api objectAPIHandlers) CopyObjectHandler` handles the request.
apiRouter.ServeHTTP(rec, req)
@@ -2623,6 +2641,16 @@ func testAPICopyObjectHandler(obj ObjectLayer, instanceType, bucketName string,
if testCase.metadataGarbage {
req.Header.Set("X-Amz-Metadata-Directive", "Unknown")
}
// The x-amz-copy-source and related x-amz-* headers set above must be
// part of the SigV4 signature, exactly as real S3 clients send them.
// Re-sign now that they are present; the verifier rejects unsigned
// x-amz-* headers (an unsigned x-amz-copy-source could otherwise turn a
// PUT grant into a server-side copy).
if testCase.accessKey != "" && testCase.secretKey != "" {
if err = signRequestV4(req, testCase.accessKey, testCase.secretKey); err != nil {
t.Fatalf("Test %d: Failed to re-sign HTTP request for copy Object: <ERROR> %v", i, err)
}
}
// Since `apiRouter` satisfies `http.Handler` it has a ServeHTTP to execute the logic of the handler.
// Call the ServeHTTP to execute the handler, `func (api objectAPIHandlers) CopyObjectHandler` handles the request.
apiRouter.ServeHTTP(rec, req)
+140
View File
@@ -0,0 +1,140 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-only
package cmd
import (
"fmt"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/minio/minio/internal/auth"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
xhttp "github.com/minio/minio/internal/http"
)
func setTestBucketDefaultRetention(t *testing.T, bucket, mode string) {
t.Helper()
enableBucketObjectLock(t, bucket)
if mode == "" {
return
}
meta, err := globalBucketMetadataSys.Get(bucket)
if err != nil {
t.Fatal(err)
}
meta.ObjectLockConfigXML = fmt.Appendf(nil, `<ObjectLockConfiguration xmlns="http://s3.amazonaws.com/doc/2006-03-01/"><ObjectLockEnabled>Enabled</ObjectLockEnabled><Rule><DefaultRetention><Mode>%s</Mode><Days>2</Days></DefaultRetention></Rule></ObjectLockConfiguration>`, mode)
if err := meta.parseAllConfigs(t.Context(), newObjectLayerFn()); err != nil {
t.Fatal(err)
}
globalBucketMetadataSys.Set(bucket, meta)
}
func TestAPIObjectLockDefaultRetentionWithLegalHold(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
endpoints: []string{"NewMultipart", "PutObjectPart", "CompleteMultipart", "CopyObject", "PutObject", "DeleteObject"},
objAPITest: func(obj ObjectLayer, instanceType, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
data := []byte("default retention and legal hold are independent")
putCopyChecksumSource(t, router, cred, bucket, "source", data, nil)
for _, mode := range []string{"GOVERNANCE", "COMPLIANCE", ""} {
setTestBucketDefaultRetention(t, bucket, mode)
for _, hold := range []string{"ON", "OFF"} {
for _, operation := range []string{"put", "copy", "multipart"} {
t.Run(instanceType+"/"+mode+"/"+hold+"/"+operation, func(t *testing.T) {
object := mode + "-" + hold + "-" + operation
headers := map[string]string{xhttp.AmzObjectLockLegalHold: hold}
// The stored retention date carries millisecond precision.
before := UTCNow().Truncate(time.Millisecond)
switch operation {
case "put":
putCopyChecksumSource(t, router, cred, bucket, object, data, headers)
case "copy":
rec := federatedCopyRequest(t, router, cred, bucket, "source", bucket, object, headers)
if rec.Code != http.StatusOK {
t.Fatalf("copy: %d %s", rec.Code, rec.Body.String())
}
case "multipart":
putFederationMultipartSource(t, router, cred, bucket, object, [][]byte{data}, headers)
}
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
ret := objectlock.GetObjectRetentionMeta(info.UserDefined)
if string(ret.Mode) != mode {
t.Errorf("stored retention mode = %q, want %q", ret.Mode, mode)
}
if mode != "" && (ret.RetainUntilDate.Before(before.Add(48*time.Hour)) || ret.RetainUntilDate.After(UTCNow().Add(48*time.Hour))) {
t.Errorf("default retention date = %s, want write time + 2 days", ret.RetainUntilDate)
}
if got := objectlock.GetObjectLegalHoldMeta(info.UserDefined).Status; string(got) != hold {
t.Errorf("stored legal hold = %q, want %q", got, hold)
}
if mode == "COMPLIANCE" && hold == "OFF" {
req, err := newTestSignedRequestV4(http.MethodDelete, getPutObjectURL("", bucket, object)+"?versionId="+info.VersionID, 0, nil, cred.AccessKey, cred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != http.StatusBadRequest || !strings.Contains(rec.Body.String(), "InvalidRequest") {
t.Errorf("version DELETE must enforce retention: %d %s", rec.Code, rec.Body.String())
}
}
})
}
}
}
},
})
}
func TestObjectLockDefaultRetentionBoundaries(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{
t: t,
objAPITest: func(obj ObjectLayer, instanceType, bucket string, _ http.Handler, _ auth.Credentials, t *testing.T) {
setTestBucketDefaultRetention(t, bucket, "COMPLIANCE")
for _, tc := range []struct {
name string
replica, marker, explicit bool
retentionErr, holdErr APIErrorCode
wantMode objectlock.RetMode
wantErr APIErrorCode
}{
{name: "trusted replica", replica: true},
{name: "marker-only ordinary write", marker: true, wantMode: objectlock.RetCompliance},
{name: "explicit retention", explicit: true, wantMode: objectlock.RetGovernance},
{name: "retention permission denied", retentionErr: ErrAccessDenied, wantErr: ErrAccessDenied},
{name: "legal hold permission denied", holdErr: ErrAccessDenied, wantErr: ErrAccessDenied},
} {
t.Run(instanceType+"/"+tc.name, func(t *testing.T) {
r := httptest.NewRequest(http.MethodPut, "http://minio.local/"+bucket+"/object", nil)
r.Header.Set(xhttp.AmzObjectLockLegalHold, "OFF")
if tc.marker {
r.Header.Set(xhttp.MinIOSourceReplicationRequest, "true")
}
until := UTCNow().Add(24 * time.Hour).Truncate(time.Second).Add(789 * time.Millisecond)
if tc.explicit {
r.Header.Set(xhttp.AmzObjectLockMode, "GOVERNANCE")
r.Header.Set(xhttp.AmzObjectLockRetainUntilDate, until.Format(time.RFC3339Nano))
}
mode, date, _, code := checkPutObjectLockAllowed(t.Context(), r, bucket, "object", obj.GetObjectInfo, tc.retentionErr, tc.holdErr, tc.replica)
if code != tc.wantErr || mode != tc.wantMode {
t.Fatalf("got mode %s error %s, want mode %s error %s", mode, niceError(code), tc.wantMode, niceError(tc.wantErr))
}
if tc.explicit && !date.Equal(until) {
t.Errorf("explicit retention lost precision: %s != %s", date, until)
}
if tc.replica && !date.IsZero() {
t.Errorf("replica acquired a destination default: %s", date)
}
})
}
},
})
}
@@ -30,6 +30,7 @@ import (
"strings"
"sync"
"testing"
"time"
miniogo "github.com/minio/minio-go/v7"
miniocredentials "github.com/minio/minio-go/v7/pkg/credentials"
@@ -51,7 +52,7 @@ func TestAPIFederatedUploadPartChecksumResponse(t *testing.T) {
})
}
func testAPIFederatedUploadPartChecksumResponse(_ ObjectLayer, instanceType, bucketName string,
func testAPIFederatedUploadPartChecksumResponse(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
algorithms := []struct {
@@ -101,6 +102,18 @@ func testAPIFederatedUploadPartChecksumResponse(_ ObjectLayer, instanceType, buc
if got := rec.Header().Get(xhttp.AmzChecksumType); got != "" {
t.Fatalf("%s: UploadPart returned checksum type %q", instanceType, got)
}
modified := rec.Header().Get(federatedLastModified)
if userAgent.want {
parts, err := objectAPI.ListObjectParts(t.Context(), bucketName, object, uploadID, 0, 100, ObjectOptions{})
if err != nil || len(parts.Parts) != 1 {
t.Fatalf("stored part: %v, %v", parts, err)
}
if modified != parts.Parts[0].LastModified.UTC().Format(time.RFC3339Nano) {
t.Errorf("response time %q does not match committed part time %s", modified, parts.Parts[0].LastModified)
}
} else if modified != "" {
t.Errorf("ordinary UploadPart exposed federation timestamp %q", modified)
}
})
}
}
@@ -158,7 +171,7 @@ func TestAPIFederatedUploadPartChecksumConcurrentOverwrite(t *testing.T) {
})
}
func testAPIFederatedUploadPartChecksumConcurrentOverwrite(_ ObjectLayer, instanceType, bucketName string,
func testAPIFederatedUploadPartChecksumConcurrentOverwrite(objectAPI ObjectLayer, instanceType, bucketName string,
apiRouter http.Handler, credentials auth.Credentials, t *testing.T,
) {
object := "federation/concurrent-overwrite"
@@ -194,6 +207,10 @@ func testAPIFederatedUploadPartChecksumConcurrentOverwrite(_ ObjectLayer, instan
}
close(start)
wg.Wait()
parts, err := objectAPI.ListObjectParts(t.Context(), bucketName, object, uploadID, 0, 100, ObjectOptions{})
if err != nil || len(parts.Parts) != 1 {
t.Fatalf("stored part: %v, %v", parts, err)
}
for i, rec := range recorders {
if rec.Code != http.StatusOK {
@@ -213,6 +230,13 @@ func testAPIFederatedUploadPartChecksumConcurrentOverwrite(_ ObjectLayer, instan
if want := hex.EncodeToString(md5sum[:]); canonicalizeETag(etags[0]) != want {
t.Fatalf("%s: concurrent request %d ETag %q, want %q", instanceType, i, etags[0], want)
}
modified, err := time.Parse(time.RFC3339Nano, rec.Header().Get(federatedLastModified))
if err != nil || modified.IsZero() {
t.Fatalf("concurrent request %d returned invalid time %v: %v", i, modified, err)
}
if canonicalizeETag(etags[0]) == parts.Parts[0].ETag && !modified.Equal(parts.Parts[0].LastModified) {
t.Errorf("surviving part has time %s, its writer returned %s", parts.Parts[0].LastModified, modified)
}
}
}
@@ -352,6 +376,9 @@ func testAPIFederatedCopyObjectPartChecksum(objectAPI ObjectLayer, instanceType,
if len(parts.Parts) != 1 {
t.Fatalf("%s: ListParts returned %d parts, want 1", instanceType, len(parts.Parts))
}
if response.LastModified != parts.Parts[0].LastModified {
t.Errorf("CopyPartResult time = %q, want stored part time %q", response.LastModified, parts.Parts[0].LastModified)
}
if got := partChecksum(algorithm.typ, parts.Parts[0]); got != want {
t.Fatalf("%s: persisted part %s is %q, want %q", instanceType, algorithm.typ.String(), got, want)
}
+53 -30
View File
@@ -34,7 +34,6 @@ import (
"github.com/minio/minio-go/v7"
"github.com/minio/minio-go/v7/pkg/encrypt"
"github.com/minio/minio-go/v7/pkg/tags"
"github.com/minio/minio/internal/amztime"
sse "github.com/minio/minio/internal/bucket/encryption"
objectlock "github.com/minio/minio/internal/bucket/object/lock"
"github.com/minio/minio/internal/bucket/replication"
@@ -60,8 +59,8 @@ import (
//
// This is only a response-shape hint. User-Agent is not authenticated and must
// never gate authorization, object visibility, or request validation. It is
// safe here because the only effect is returning the checksum of the body the
// caller was already authorized to upload.
// safe here because the only effect is returning the checksum and modification
// time of the body the caller was already authorized to upload.
func isFederatedInternalRequest(userAgent string) bool {
for _, product := range strings.Fields(userAgent) {
name, version, ok := strings.Cut(product, "/")
@@ -180,11 +179,17 @@ func (api objectAPIHandlers) NewMultipartUploadHandler(w http.ResponseWriter, r
ctx, r = applyReplicationTrust(ctx, r, trustedReplication, replicaTrusted)
}
// Check if bucket encryption is enabled
sseConfig, _ := globalBucketSSEConfigSys.Get(bucket)
sseConfig.Apply(r.Header, sse.ApplyOptions{
AutoEncrypt: globalAutoEncryption,
})
// A validated raw SSE-C replica upload carries source ciphertext in every
// part; the destination must not add its own encryption or compression.
rawSSECReplica := isRawSSECReplica(r.Header, replicaTrusted)
if !rawSSECReplica {
// Check if bucket encryption is enabled
sseConfig, _ := globalBucketSSEConfigSys.Get(bucket)
sseConfig.Apply(r.Header, sse.ApplyOptions{
AutoEncrypt: globalAutoEncryption,
})
}
// Validate the storage class header if present. Query values retain the
// existing compatibility path, including its historical validation behavior.
@@ -213,19 +218,7 @@ func (api objectAPIHandlers) NewMultipartUploadHandler(w http.ResponseWriter, r
return
}
ssecRepHeaders := []string{
"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm",
"X-Minio-Replication-Server-Side-Encryption-Sealed-Key",
"X-Minio-Replication-Server-Side-Encryption-Iv",
}
ssecRep := false
for _, header := range ssecRepHeaders {
if val := r.Header.Get(header); val != "" {
ssecRep = true
break
}
}
if !ssecRep || !replicaTrusted {
if !rawSSECReplica {
if err = setEncryptionMetadata(r, bucket, object, encMetadata); err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -266,12 +259,32 @@ func (api objectAPIHandlers) NewMultipartUploadHandler(w http.ResponseWriter, r
getObjectInfo := objectAPI.GetObjectInfo
retentionMode, retentionDate, legalHold, s3Err := checkPutObjectLockAllowed(ctx, r, bucket, object, getObjectInfo, retPerms, holdPerms, replicaTrusted)
if s3Err == ErrNone && retentionMode.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockMode)] = string(retentionMode)
metadata[strings.ToLower(xhttp.AmzObjectLockRetainUntilDate)] = amztime.ISO8601Format(retentionDate.UTC())
}
if s3Err == ErrNone && legalHold.Status.Valid() {
metadata[strings.ToLower(xhttp.AmzObjectLockLegalHold)] = string(legalHold.Status)
if s3Err == ErrNone {
// A trusted replica NewMultipartUpload addressing a specific version can
// be a full retransmit over an existing version whose lock state is newer
// than the source snapshot (issue #120). Order the incoming update against
// what is stored so a stale value cannot overwrite it. opts is built below
// (its ServerSideEncryption depends on the encMetadata merge that has not
// happened yet), so read the replica ordering inputs the way
// putOptsFromHeaders will; a malformed timestamp fails the request when
// opts is built, so a parse error here is left as a zero time.
var (
storedLock objectLockState
srcRetentionTS, srcLegalholdTS time.Time
)
if replicaTrusted {
srcRetentionTS, _ = time.Parse(time.RFC3339, strings.TrimSpace(r.Header.Get(xhttp.MinIOSourceObjectRetentionTimestamp)))
srcLegalholdTS, _ = time.Parse(time.RFC3339, strings.TrimSpace(r.Header.Get(xhttp.MinIOSourceObjectLegalHoldTimestamp)))
if versionID := strings.TrimSpace(r.Form.Get(xhttp.VersionID)); versionID != "" {
var lerr error
if storedLock, lerr = replicaStoredLock(ctx, getObjectInfo, bucket, object, versionID); lerr != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, lerr), r.URL)
return
}
}
}
applyReplicatedObjectLock(metadata, storedLock, replicaTrusted,
retentionMode, retentionDate, legalHold, srcRetentionTS, srcLegalholdTS)
}
if s3Err != ErrNone {
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(s3Err), r.URL)
@@ -289,7 +302,7 @@ func (api objectAPIHandlers) NewMultipartUploadHandler(w http.ResponseWriter, r
// Ensure that metadata does not contain sensitive information
crypto.RemoveSensitiveEntries(metadata)
if isCompressible(r.Header, object) {
if !rawSSECReplica && isCompressible(r.Header, object) {
// Storing the compression metadata.
metadata[ReservedMetadataPrefix+"compression"] = compressionAlgorithmV2
}
@@ -559,7 +572,8 @@ func (api objectAPIHandlers) CopyObjectPartHandler(w http.ResponseWriter, r *htt
SSE: dstOpts.ServerSideEncryption,
}
partInfo, err := core.PutObjectPart(ctx, dstBucket, dstObject, uploadID, partID, gr, length, popts)
writeCtx, modified := withFederatedWriteTime(ctx)
partInfo, err := core.PutObjectPart(writeCtx, dstBucket, dstObject, uploadID, partID, gr, length, popts)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
@@ -567,7 +581,7 @@ func (api objectAPIHandlers) CopyObjectPartHandler(w http.ResponseWriter, r *htt
response := generateCopyObjectPartResponse(PartInfo{
ETag: partInfo.ETag,
LastModified: partInfo.LastModified,
LastModified: *modified,
ChecksumCRC32: partInfo.ChecksumCRC32,
ChecksumCRC32C: partInfo.ChecksumCRC32C,
ChecksumSHA1: partInfo.ChecksumSHA1,
@@ -1069,6 +1083,7 @@ func (api objectAPIHandlers) PutObjectPartHandler(w http.ResponseWriter, r *http
// checksum cannot be mixed with a concurrent overwrite.
hash.AddChecksumHeader(w, partChecksumMap(partInfo))
}
setFederatedWriteTime(w, r, partInfo.LastModified)
writeSuccessResponseHeadersOnly(w)
}
@@ -1172,6 +1187,14 @@ func (api objectAPIHandlers) CompleteMultipartUploadHandler(w http.ResponseWrite
}
opts.Versioned = versioned
opts.VersionSuspended = suspended
// A replicated multipart completion carries the internal replication marker
// (the sender does not re-assert REPLICA status on Complete, so this is keyed
// on trusted replication, matching completeMultipartOpts). The object layer
// re-orders the Object Lock it carries against the destination version read
// under the write lock, but only for an SSE-C upload -- the scope this issue
// enables -- so a marker-only non-SSE-C completion keeps ordinary write
// semantics (issue #120).
opts.ReplicaLockReconcile = trustedReplication
// First, we compute the ETag of the multipart object.
// The ETag of a multi-part object is always:
+5
View File
@@ -165,6 +165,11 @@ func testAPIZeroByteSSECAuthenticatesKey(obj ObjectLayer, instanceType, bucketNa
t.Fatal(err)
}
req.Header.Set(xhttp.AmzCopySource, SlashSeparator+pathJoin(bucketName, object))
// Re-sign so x-amz-copy-source is covered by the signature, as real S3
// clients send it; the verifier rejects unsigned x-amz-* headers.
if err = signRequestV4(req, credentials.AccessKey, credentials.SecretKey); err != nil {
t.Fatalf("%s: failed to re-sign UploadPartCopy request: %v", instanceType, err)
}
rec = httptest.NewRecorder()
apiRouter.ServeHTTP(rec, req)
if rec.Code != http.StatusForbidden {

Some files were not shown because too many files have changed in this diff Show More